Risk identification methods and terminal devices
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-24
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]本申请实施例提供了一种风险识别方法和终端设备,以解决现有的视频监控方法在上述问题下降低了对隐蔽风险的识别准确率的问题
本申请实施例提供的一种风险识别方法,通过对收集到的语音信息进行声源定位,得到定位结果;基于定位结果对语音信息进行声源分离,得到不同人员各自对应的语音数据;不同人员包括已知人员;对已知人员的语音数据进行分析,得到已知人员在不同维度的生理特征值;基于不同维度的生理特征值进行风险分析,得到风险计算结果。本申请先通过声源定位和声源分离实现多人语音解耦,可以降低多人同时说话导致的语音识别混淆、干扰,从混杂的多人语音中自动拆分出已知人员的独立语音数据,提高了后续分析的准确性;之后从多个维度综合分析人员生理状态,避免了单一维度误判,提高了对潜在风险状态的识别精度,从而可以在视觉失效(如盲区、暗光)或行为隐蔽(如低语威胁)的场景下准确识别隐蔽胁迫。
Smart Images

Figure CN122575406A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, and in particular relates to a risk identification method and terminal device. Background Technology
[0002] Currently, existing technologies for identifying theft and coercion primarily rely on video surveillance. However, existing video surveillance systems have significant limitations in practical applications. Firstly, there are blind spots and obstruction issues: cameras have limited fields of view and are easily blocked. Secondly, there are lighting limitations: at night or in dimly lit spaces, insufficient light degrades image quality, making it difficult to identify details. Finally, there is difficulty in identifying covert coercion: non-violent methods such as whispered threats and concealed weapon control may force individuals to maintain a normal posture, making them appear visually normal. Therefore, existing video surveillance methods, due to these problems, reduce the accuracy of identifying covert risks. Summary of the Invention
[0003] This application provides a risk identification method and terminal device to address the problem that existing video surveillance methods have reduced the accuracy of identifying hidden risks under the aforementioned conditions.
[0004] In a first aspect, embodiments of this application provide a risk identification method, including: The collected speech information is used to locate the sound source, and the location result is obtained; Based on the localization results, the speech information is separated into sound sources to obtain the speech data corresponding to different people; different people include known people. By analyzing the voice data of known individuals, we can obtain their physiological characteristic values in different dimensions. Risk analysis is performed based on physiological characteristic values of different dimensions to obtain risk calculation results.
[0005] Optionally, "different personnel" may also include other personnel; risk analysis is performed based on physiological characteristic values of different dimensions to obtain risk calculation results, including: Speech recognition is performed on the voice data of other people to obtain the corresponding voice content of other people; Risk analysis is performed based on voice content and physiological feature values of different dimensions to obtain risk calculation results.
[0006] Optionally, risk analysis can be performed based on speech content and physiological feature values of different dimensions to obtain risk calculation results, including: Risk analysis is performed based on location results, voice content, and physiological feature values of different dimensions to obtain risk calculation results.
[0007] Optionally, risk analysis can be performed based on the location results, voice content, and physiological feature values of different dimensions to obtain risk calculation results, including: The stress risk score is obtained by weighting and summing the location results, physiological feature values of different dimensions, and voice content. The total weight of the physiological feature values of different dimensions is greater than the weight of the voice content, and the weight of the voice content is greater than the weight of the location results. If the stress risk score is greater than the risk threshold, the risk calculation result is determined to indicate the existence of hidden stress.
[0008] Optionally, physiological characteristic values in different dimensions include microfibrillation characteristic values, respiratory tachycardia, and baseline deviation values; the voice data of known individuals are analyzed to obtain physiological characteristic values of known individuals in different dimensions, including: Microvibration detection is performed on the speech data of known individuals using nonlinear signal processing tools to obtain microvibration feature values; The speech data of known individuals is divided into speech segments and non-speech segments; Determine the degree of tachypnea based on non-speech segments and speech segments; Based on the speech data of known individuals and the standard speech baseline of known individuals, the baseline deviation value is calculated.
[0009] Optionally, tachypnea includes frequency perturbation, amplitude perturbation, and harmonic-to-noise ratio; based on microfreak characteristics, tachypnea, and a known standard speech baseline, a baseline deviation is calculated, including: The physiological feature vector is obtained by combining the micro-vibration feature value, frequency perturbation value, amplitude perturbation value, and harmonic-to-noise ratio. The baseline deviation value is obtained by calculating Mahalanobis distance based on physiological feature vectors and standard speech baseline.
[0010] Optionally, after obtaining the voice data corresponding to each individual, the method may also include: Feature extraction is performed on the speech data of different individuals to obtain the initial spectral features of each individual; The initial spectral features of different individuals are input into a trained deep spectral mapping network for enhancement processing to obtain the target spectral features of different individuals. Based on the target spectral characteristics of different individuals, the enhanced speech data of each individual is obtained. Accordingly, the voice data of known individuals are analyzed to obtain their physiological characteristic values in different dimensions, including: By analyzing the enhanced speech data of known individuals, we can obtain their physiological characteristic values in different dimensions.
[0011] Optionally, a microphone array is installed inside the vehicle; the collected voice information is used to locate the sound source, and the location results are obtained, including: The speech information is initially located based on the phase transform weighted generalized cross-correlation algorithm and the microphone array to obtain the vocal area; Based on the multi-signal classification algorithm and the vocal region, the speech information is re-localized to obtain the vocal position; The location of the sound is mapped to the target area where different people are located, and the location results are obtained.
[0012] Optionally, the localization results include localization sub-results for different individuals; based on the localization results, sound source separation is performed on the speech information to obtain the speech data corresponding to each individual, including: Based on voice information and localization results for different individuals, a minimum variance distortion-free response beamformer is constructed to output target domain voice signals for different individuals. Independent component analysis was used to separate the sound sources of the target domain speech signals of different individuals, thereby obtaining the speech data corresponding to each individual.
[0013] Secondly, embodiments of this application provide a risk identification device, including: The sound source localization unit is used to locate the sound source of the collected speech information and obtain the localization result; The sound source separation unit is used to separate the sound sources of speech information based on the localization results, and obtain the speech data corresponding to different people; different people include known people. The speech analysis unit is used to analyze the speech data of known individuals to obtain the physiological characteristic values of the known individuals in different dimensions; The risk analysis unit is used to perform risk analysis based on physiological characteristic values of different dimensions and obtain risk calculation results.
[0014] Thirdly, embodiments of this application provide a terminal device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the risk identification method as described in any one of the first aspects above.
[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the risk identification method as described in any one of the first aspects above.
[0016] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, enables the terminal device to execute the risk identification method described in any one of the first aspects.
[0017] The beneficial effects of the embodiments in this application compared with the prior art are: This application provides a risk identification method that involves: 1) Localizing the collected speech information to obtain a localization result; 2) Separating the speech information based on the localization result to obtain speech data corresponding to different individuals, including known individuals; 3) Analyzing the speech data of known individuals to obtain their physiological characteristic values in different dimensions; and 4) Performing risk analysis based on these physiological characteristic values to obtain risk calculation results. This application first achieves multi-person speech decoupling through sound source localization and separation, reducing speech recognition confusion and interference caused by multiple people speaking simultaneously. It automatically extracts independent speech data of known individuals from mixed multi-person speech, improving the accuracy of subsequent analysis. Then, it comprehensively analyzes the physiological state of individuals from multiple dimensions, avoiding misjudgment based on a single dimension and improving the accuracy of identifying potential risk states. This allows for accurate identification of concealed coercion in scenarios with visual impairment (such as blind spots or low light) or covert behavior (such as whispered threats). Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the implementation of a risk identification method provided in an embodiment of this application; Figure 2 This is a flowchart illustrating the specific implementation of step S101 in a risk identification method provided in an embodiment of this application; Figure 3 This is a flowchart illustrating the specific implementation of step S102 in a risk identification method provided in an embodiment of this application; Figure 4 This is a flowchart illustrating the specific implementation of step S103 in a risk identification method provided in an embodiment of this application; Figure 5 This is a flowchart illustrating the implementation of a risk identification method provided in another embodiment of this application; Figure 6 This is a flowchart illustrating the implementation of a risk identification method provided in another embodiment of this application; Figure 7 This is a schematic diagram of the overall process of an intent recognition system provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a risk identification device provided in an embodiment of this application; Figure 9This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation
[0020] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0021] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0022] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0023] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0024] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0025] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0026] Please see Figure 1 , Figure 1 This is a flowchart illustrating the implementation of a risk identification method according to an embodiment of this application. In this embodiment, the entity executing the risk identification method is a terminal device.
[0027] The following section will use the application of risk identification methods to vehicles as an example to explain in detail the specific implementation process of the risk identification method. Based on this, the aforementioned terminal device can be an in-vehicle terminal.
[0028] like Figure 1 As shown, a risk identification method provided in one embodiment of this application may include S101~S105, which are described in detail below: In S101, the collected speech information is used to locate the sound source and obtain the location result.
[0029] It should be noted that a microphone array can be deployed inside the vehicle in order to collect all voice information inside the vehicle in real time.
[0030] In some possible embodiments, the above-mentioned microphone array can be deployed as follows: at least three non-collinear microphones forming a planar array, or four or eight or more non-coplanar microphones forming a spatial array.
[0031] In this embodiment, after the terminal device acquires the in-vehicle voice information through the aforementioned microphone array, it can preprocess the voice information. The voice information may include one voice signal acquired by each microphone in the microphone array.
[0032] The above preprocessing includes echo cancellation and automatic gain control.
[0033] In practical applications, Acoustic Echo Cancellation (AEC) is used to eliminate echo interference, while Automatic Gain Control (AGC) is used to stabilize volume.
[0034] It should be noted that the specific implementation process of the above echo cancellation and automatic gain control can be found in existing echo cancellation and automatic gain control technologies in the field of audio communication, and will not be elaborated here.
[0035] In this embodiment, after receiving the preprocessed speech information, the terminal device can calculate the azimuth, elevation, and distance coordinates of each speech signal relative to the microphone array based on the Time Difference of Arrival (TDOA), Intensity Difference of Arrival (IDOA), and Generalized Cross-Correlation with Phase Transform (GCC-PHAT) algorithms. Then, the terminal device can cluster speech signals from the same time and spatial location to form independent sound source location clusters, thereby obtaining the final localization result.
[0036] The location results may include location sub-results for different individuals. Different individuals include, but are not limited to, known individuals and other individuals.
[0037] Known personnel are those who may pose hidden risks within the vehicle and whose risks require identification. Other personnel are those who may exert covert coercion on known personnel, i.e., those who pose a risk to known personnel.
[0038] Among them, the known personnel may be the driver.
[0039] Each location sub-result includes the location range and coordinates of the corresponding person in the target area.
[0040] The aforementioned positions include, but are not limited to: the driver's seat, the front passenger seat, the seat behind the driver, and the seat behind the front passenger.
[0041] In practical applications, the time difference of arrival is used to describe the time difference between the arrival of a sound from the same sound source at two microphones.
[0042] The arrival intensity difference is used to describe the energy difference / amplitude difference / sound pressure level difference of sound emitted from the same sound source reaching two microphones.
[0043] Specifically, after receiving the preprocessed speech information, the terminal device can calculate the cross-power spectrum between the speech signals acquired by any two microphones in the microphone array using Generalized Cross-Correlation Phase Transform (GCC-PHAT). Then, it performs PHAT weighted normalization on this cross-power spectrum and performs an inverse transform to obtain the generalized cross-correlation function. Next, the terminal device can determine the cross-correlation peak point in this generalized cross-correlation function as the Time Difference of Arrival (TDOA), and, combined with the speed of sound, convert the TDOA into the distance difference between the two microphones and the sound source, thus establishing multiple sets of hyperboloid equations. The spatial coordinates of the sound source are then obtained through least-squares iteration. Simultaneously, the terminal device can calculate the short-time average energy of each of the two speech signals based on the Difference of Arrival (IDOA) algorithm, and calculate the IDOA based on the short-time average energy of each of the two speech signals to assist in distance constraints. Finally, the terminal device can use the center of the microphone array as the origin, calculate the horizontal azimuth and vertical elevation angles of each voice signal relative to the microphone array based on the spatial coordinates of each sound source, and cluster voice signals at the same time and in the same spatial location to form independent sound source location clusters, so as to obtain the final positioning result.
[0044] In one embodiment of this application, when a microphone array is installed inside the vehicle, the terminal device can specifically utilize, as shown in the example below. Figure 2 Steps S201 to S203 shown implement step S101, as detailed below: In S201, the speech information is initially located based on the phase transform weighted generalized cross-correlation algorithm and the microphone array to obtain the sound production area.
[0045] In this embodiment, the terminal device can calculate the spatial response power of the microphone array using the Steered Response Power with Phase Transform (SRP-PHAT) algorithm, and combine it with the microphone array coordinates to perform preliminary localization of the speech information, so as to lock the sound source area of the speech information.
[0046] In some possible embodiments, the spatial response power at spatial point x in the target region can be specifically calculated according to the following formula: ; in, Spatial points representing the target area Spatial response power (dimensionless) at a given location. Indicates the first in the microphone array The spectrum of the speech signal collected by the microphone. Indicates the first in the microphone array The conjugate of the spectrum of the speech signal collected by the microphone. Representing a spatial point to microphone The time delay difference.
[0047] Based on this, the terminal device can calculate the response power matrix according to the above formula. This response power matrix includes the spatial response power of any point within the vehicle.
[0048] In this embodiment, after obtaining the response power matrix, the terminal device can traverse various locations inside the vehicle, extract the directional range corresponding to the power peak in the response power matrix, and combine it with the microphone array coordinates to eliminate directional ranges that exceed the target area, thereby obtaining a three-dimensional spatial region, namely the sound emission area.
[0049] In S202, the speech information is re-localized based on the multi-signal classification algorithm and the speech region to obtain the speech location.
[0050] In this embodiment, after obtaining the sound-emitting area, the terminal device can accurately locate the sound source within that area using the Multiple Signal Classification (MUSIC) algorithm to obtain the specific sound source location (such as three-dimensional coordinates).
[0051] Specifically, the terminal device can construct the received signal matrix of the microphone array based on the voice information and calculate the covariance matrix of the received signal matrix; then, the terminal device can perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors; then, the terminal device can divide the signal subspace and noise subspace according to the magnitude of each eigenvalue.
[0052] For example, the terminal device can select the feature vectors corresponding to the first K largest feature values (K is the dimension of the signal subspace) to form the signal subspace, and the feature vectors corresponding to the remaining NK feature values to form the noise subspace. Here, N represents the number of microphones in the microphone array.
[0053] Subsequently, the terminal device can construct an array steering vector point by point within the initially located sound-emitting area, according to the preset azimuth and elevation search ranges inside the vehicle, and calculate the MUSIC spatial spectrum. The azimuth and elevation angles corresponding to the peak values of the spatial spectrum are the azimuth parameters of the sound source.
[0054] In some possible embodiments, the peak value of the spatial spectrum can be calculated using the following formula: ; in, Indicates azimuth angle and pitch angle The MUSIC spatial spectrum (dimensionless) at the location. Indicates the array steering vector. This represents the conjugate transpose of the array steering vector. Represents the noise subspace matrix. This represents the conjugate transpose of the noise subspace matrix.
[0055] In this embodiment, the terminal device can combine the distance estimation from the origin of the microphone array to the sound-emitting area and the directional parameters of the sound source to calculate the three-dimensional coordinates of the sound source (in the microphone array coordinate system). These three-dimensional coordinates are the sound-emitting position obtained by repositioning.
[0056] It should be noted that the above-mentioned voice positions may include the voice positions of known individuals and the voice positions of other individuals.
[0057] In S203, the location of the sound is mapped to the target area where different people are located to obtain the positioning result.
[0058] It should be noted that the target area for different people can be inside the vehicle.
[0059] In this embodiment, after obtaining the voice positions of different people, the terminal device can perform coordinate system transformation on the voice positions of different people, since the voice positions of different people are three-dimensional coordinates in the microphone array coordinate system, so as to map the voice positions of different people to the target area coordinate system, thereby obtaining the final positioning result.
[0060] Combining steps S201-S203, this embodiment uses SRP-PHAT for initial localization, thus maintaining strong robustness in target areas with strong reverberation and complex noise (such as inside a vehicle). It can quickly determine the sound-emitting area, avoiding blind searching throughout the entire space, reducing computational complexity, and improving real-time localization. Subsequently, based on the sound-emitting area constraints, MUSIC is used for fine localization, which significantly reduces the MUSIC traversal space, reduces false peak interference, lowers computational load, and retains the high angular resolution advantage of MUSIC. This enables accurate differentiation of multiple sound sources at close range, resulting in higher localization accuracy and stronger stability.
[0061] In S102, the speech information is separated into sound sources based on the positioning results to obtain the speech data corresponding to different people; different people include known people and other people.
[0062] In this embodiment of the application, the terminal device can divide the voice information into multiple independent sound source channels according to the spatial location in the positioning result. Then, the terminal device can perform sound source separation on the voice sub-information output by each independent sound source channel based on the time domain end-to-end separation algorithm to obtain pure voice data corresponding to each person.
[0063] In practical applications, the above-mentioned end-to-end temporal separation algorithms include, but are not limited to: Time-domain Audio Separation Network (TasNet) and Convolutional Time-domain Audio Separation Network (Conv-TasNet).
[0064] It should be noted that the specific implementation process of the above-mentioned time-domain end-to-end separation algorithms can be found in their respective existing technologies, and will not be elaborated here.
[0065] In one embodiment of this application, when the location result includes location sub-results for different individuals, the terminal device can specifically use methods such as... Figure 3 Steps S301 to S302 shown implement step S102, as detailed below: In S301, based on voice information and the localization results of different people, a minimum variance distortion-free response beamformer is constructed to output the target domain voice signal of different people.
[0066] In this embodiment, the terminal device can construct a minimum variance distortionless response (MVDR) beamformer based on the positioning results of different personnel, perform directional enhancement on the mixed speech signal, i.e., speech information, extract the target domain speech signal of each person, and suppress the interference of other people's speech and environmental noise.
[0067] Specifically, for each person's location sub-result, the terminal device can initialize its corresponding MVDR beamformer parameters to construct a minimum variance distortion-free response beamformer for each person. Simultaneously, the terminal device can also calculate the covariance matrix corresponding to the voice information.
[0068] Then, the terminal device can calculate the beamformer weight vector corresponding to different personnel according to the following formula: ; in, This represents the beamformer weight vector corresponding to any person. This represents the covariance matrix corresponding to the speech information. The array steering vector representing the target direction corresponding to the positioning sub-result of any person. This represents the conjugate transpose of the array steering vector described above.
[0069] In this embodiment, the terminal device can input voice information into the MVDR beamformer corresponding to different people, and then use the weight vector of each beamformer to perform weighted summation on the multi-microphone signals to obtain the target domain voice signals of different people.
[0070] In S302, independent component analysis is used to separate the source of the target domain speech signal of different people to obtain the speech data corresponding to each person.
[0071] In this embodiment, for the target domain speech signal of each person, the terminal device can further separate the residual aliasing signal through the independent component analysis (FastICA) algorithm to obtain pure single-person speech data, that is, the speech data corresponding to each person.
[0072] Specifically, the process of sound source separation using the FastICA (Fast Independent Component Analysis) algorithm is as follows: ; in, This represents the mixed signal vector corresponding to the target domain speech information of any person. Represents the separation matrix. This represents the separated independent component vectors, i.e., the speech data corresponding to any given person. Represents a non-Gaussianity measure (such as kurtosis or negative entropy).
[0073] Combining steps S301 to S302, this embodiment combines beamforming and blind source separation in a cascaded process to improve the separability and recognizability of human voices in noisy backgrounds, making it easier to capture and verify low-volume threatening speech.
[0074] In S103, the voice data of the known person is analyzed to obtain the physiological characteristic values of the known person in different dimensions.
[0075] In this embodiment, the terminal device can extract features from the voice data of a known person to obtain physiological feature values of the known person in different dimensions. These physiological feature values include microfreak feature values, respiratory rate, and baseline deviation values.
[0076] It should be noted that the baseline deviation value is used to describe the deviation between the current speech characteristics of a known person and the standard speech baseline of the known person. The standard speech baseline describes the speech characteristics of the known person in a normal state.
[0077] Current speech characteristics include, but are not limited to, the micro-tremor characteristics and respiratory rate of known individuals.
[0078] In some possible embodiments, the terminal device can calculate the amplitude difference between any adjacent frames in the voice data of a known person, and calculate the micro-tremor feature value of the known person based on each amplitude difference.
[0079] In other possible embodiments, the terminal device can calculate the respiratory rate (number of breaths per minute), the mean respiratory interval, and the standard deviation of the respiratory interval based on respiratory segments in the voice data of a known person, and then perform a weighted summation of the respiratory rate, the mean respiratory interval, and the standard deviation to obtain the respiratory tachycardia of the known person. The respiratory segments can be determined based on silence frames in the voice data of the known person.
[0080] In one embodiment of this application, when different dimensions of physiological characteristic values include microfibrillation characteristic values, tachypnea, and baseline deviation values, the terminal device can specifically achieve the following: Figure 4 Steps S401 to S404 shown implement step S103, as detailed below: In S401, a nonlinear signal processing tool is used to detect microvibrations in the speech data of known individuals, and microvibration feature values are obtained.
[0081] It should be noted that the nonlinear signal processing tool can be the Teager energy operator (TEO).
[0082] In this embodiment, the terminal device can perform bandpass filtering on the voice data of known personnel to obtain a micro-vibration frequency band of 8-12Hz; then, the terminal device can use the Teager energy operator to perform micro-vibration detection on any sampling point of the discrete voice signal corresponding to the micro-vibration frequency band to obtain the Teager energy value corresponding to each sampling point.
[0083] In some possible embodiments, the Teager energy operator is as follows: ; in, This represents the Teager energy operator. The first discrete speech signal represents the... One sampling point.
[0084] In this embodiment, after obtaining the Teager energy value corresponding to each sampling point, the terminal device can calculate the micro-vibration characteristic value according to the following formula: ; in, Represents the microtremor characteristic value (dimensionless, range of values: ), Indicates the number of signal sampling points. The first discrete speech signal represents the... One sampling point.
[0085] In S402, the speech data of known personnel is divided into speech segments and non-speech segments.
[0086] In S403, the degree of respiratory tachycardia is determined based on non-speech segments and speech segments.
[0087] In this embodiment, since the degree of respiratory urgency can be determined by non-speech segments and speech segments, the terminal device can divide the speech data of known individuals into speech segments and non-speech segments.
[0088] Afterwards, the terminal device can obtain each speech cycle of the known person's speech data through non-speech segments and speech segments, and calculate the frequency perturbation value (Jitter) and amplitude perturbation value (Shimmer) based on each speech cycle.
[0089] Among them, the frequency perturbation value is used to describe the small fluctuations in pitch between each speech cycle of the speech signal.
[0090] The amplitude perturbation value is used to measure the minute fluctuations in amplitude between each voice cycle of the second-transmitted voice signal.
[0091] In some possible embodiments, the frequency perturbation value can be calculated according to the following formula: ; in, Indicates the first The duration of a speech cycle (in seconds), where K represents the total number of cycles.
[0092] In some other possible embodiments, the amplitude perturbation value can be calculated according to the following formula: ; in, Indicates the first The amplitude of a speech cycle, where K represents the total number of cycles.
[0093] In this embodiment, the terminal device can calculate the His Noise Reduction (HNR) based on the speech segment.
[0094] In practical applications, the harmonic-to-noise ratio (HNR) refers to the ratio of harmonic energy to noise energy.
[0095] It should be noted that the specific harmonic noise ratio mentioned above can be obtained by referring to existing harmonic noise ratio calculation methods, and will not be elaborated here.
[0096] In this embodiment, the terminal device can normalize and splice the above-mentioned frequency perturbation value, amplitude perturbation value and harmonic noise ratio to obtain a breathing feature vector, and determine the breathing feature vector as the breathing tachycardia.
[0097] In S404, the baseline deviation value is calculated based on the speech data of the known personnel and the standard speech baseline of the known personnel.
[0098] In this embodiment, after obtaining the voice data of a known person, the terminal device can compare the voice data with the standard voice baseline of the known person to determine the baseline deviation value.
[0099] In some possible embodiments, the terminal device can input the aforementioned voice data and the standard voice baseline of a known person into a trained baseline deviation determination model for processing to obtain a baseline deviation value.
[0100] It should be noted that the baseline deviation determination model described above can be trained from a pre-built neural network model.
[0101] Specifically, the baseline deviation determination model can be obtained by training a pre-built neural network model based on a preset sample set. Each sample data in the preset sample set includes first sample information (which includes the sample speech data and the corresponding standard speech baseline) and a baseline deviation value corresponding to the first sample information. When training the pre-built neural network model, the first sample information in each sample is used as the input to the neural network model, and the baseline deviation value corresponding to the first sample information in each sample is used as the output of the neural network model. Through training, the neural network model can learn the correspondence between all possible first sample information and baseline deviation values, and the trained neural network model becomes the baseline deviation determination model.
[0102] In one embodiment of this application, when the respiratory tachycardia includes frequency perturbation value, amplitude perturbation value, and harmonic-to-noise ratio, the terminal device can specifically implement step S404 according to the following steps, detailed below: The physiological feature vector is obtained by combining the micro-vibration feature value, frequency perturbation value, amplitude perturbation value, and harmonic-to-noise ratio. The baseline deviation value is obtained by calculating Mahalanobis distance based on physiological feature vectors and standard speech baseline.
[0103] In this embodiment, after calculating the micro-vibration feature value and respiratory rate based on the voice data of a known person, the terminal device can normalize and concatenate the aforementioned micro-vibration feature value and respiratory rate to obtain the physiological feature vector of the known person. ,in, This represents the harmonic-to-noise ratio. The terminal device can then calculate the Mahalanobis distance between this physiological feature vector and the standard speech baseline, and determine this Mahalanobis distance as the baseline deviation value. The standard speech baseline includes, but is not limited to, the standard microfreak feature values and standard respiratory tachycardia of known individuals.
[0104] In some possible embodiments, the Mahalanobis distance described above can be calculated using the following formula: ; in, This represents the baseline deviation value; the larger the value, the more abnormal the physiological state. Represents physiological feature vectors, Indicates the standard speech baseline. This represents the characteristic covariance matrix.
[0105] Combining steps S401-S404, this embodiment selects three dimensions: micro-tremor feature value, respiratory rate, and baseline deviation value. This can cover the physiological state details in known individuals' speech, capturing not only the tension / stress state reflected by vocal micro-tremor, but also the physiological stress response reflected by respiratory rate, and quantifying the difference between the current state and the normal state through the baseline deviation value. This multi-dimensional complementarity avoids the one-sidedness of a single physiological feature and improves the comprehensiveness of physiological state recognition. In addition, this embodiment uses fine-grained cues (micro-tremor feature value, respiratory rate, and baseline deviation value) that are more difficult to subjectively control to characterize tension and fear states. This allows for the formation of a discriminative basis even when the semantic content is not obvious or the behavior is deliberately hidden, thereby enhancing the ability to resist spoofing.
[0106] In S104, risk analysis is performed based on physiological characteristic values of different dimensions to obtain risk calculation results.
[0107] In this embodiment, after obtaining physiological feature values of different dimensions, the terminal device can perform a weighted summation of these values to obtain a physiological fear score. The terminal device can then compare this physiological fear score with a risk threshold. The risk threshold can be determined according to actual needs and is not limited here. For example, the risk threshold could be 0.7.
[0108] When a terminal device detects that the physiological fear score is greater than the risk threshold, it can determine that the risk calculation result is concealed coercion.
[0109] In some possible embodiments, the terminal device may output an alarm message when it detects that the risk calculation result is a concealed threat.
[0110] In one embodiment of this application, when different personnel include other personnel, in order to further improve the accuracy of determining the risk calculation results, the terminal device can specifically implement step S104 through the following steps, detailed below: Speech recognition is performed on the voice data of other people to obtain their voice content; Risk analysis is performed based on voice content and physiological feature values of different dimensions to obtain risk calculation results.
[0111] In this embodiment, the terminal device can input the voice data of other people into a speech recognition model for processing to obtain the text content corresponding to the voice data. The speech recognition model can be trained using a Transformer model or a Conformer model.
[0112] It should be noted that the encoder of the speech recognition model is used to extract and encode features from speech frames, and the decoder maps the encoded features to the corresponding text sequence, outputs the preliminary text results in real time, and then splices the preliminary text results of each consecutive speech frame to eliminate text redundancy between frames in order to obtain complete text content.
[0113] In this embodiment, after obtaining the text content, the terminal device can perform keyword extraction, semantic slot extraction, and tone recognition on the text content to obtain the keyword sequence and semantic intent tag corresponding to the text content, that is, to obtain the voice content of other people. Among them, the semantic intent tag is obtained by integrating the semantic information obtained from semantic slot extraction and the tone recognition result, including but not limited to: instruction type, threat level, behavioral intent, target, and tone category.
[0114] Specifically, the terminal device can use a TF-IDF algorithm combined with string matching to perform keyword retrieval on the aforementioned text content, that is, to match it with each keyword in a pre-built keyword dictionary. Afterwards, the matched keywords can be deduplicated and sorted by frequency of occurrence to obtain a keyword sequence. Simultaneously, the terminal device can also parse the text content using semantic analysis algorithms to obtain semantic information including instruction type, threat level, behavioral intent, and target. Furthermore, the terminal device can extract tone-related features based on the sentence structure, keyword categories, and punctuation marks (such as exclamation marks and question marks), such as commanding sentences ("Don't move!"), threatening sentences ("Otherwise, I won't be polite!"), persuasive sentences ("Cooperate with me, and I'll let you go"), reassuring sentences ("Don't be afraid, I won't hurt you"), and coercive sentences ("You must listen to me"). Then, the terminal device can determine the tone recognition result based on these tone features. The tone recognition result includes, but is not limited to: commanding, threatening, persuasive, reassuring, coercive, and neutral.
[0115] In this embodiment, after obtaining the voice content of other people, the terminal device can input the physiological feature values and voice content of different dimensions into the trained risk decision model for risk analysis to obtain the risk calculation result.
[0116] It should be noted that the risk decision-making model can be obtained by training a pre-built first deep learning model based on a first preset sample set. Each sample data in the first preset sample set includes second sample information (including physiological feature values and speech content in different dimensions) and the corresponding sample risk calculation result. When training the pre-built first deep learning model, the second sample information in each sample is used as the input to the first deep learning model, and the sample risk calculation result corresponding to the second sample information in each sample is used as the output of the first deep learning model. Through training, the first deep learning model can learn the correspondence between all possible second sample information and sample risk calculation results, and the trained first deep learning model becomes the risk decision-making model.
[0117] In another embodiment of this application, after obtaining the voice content of other persons, in order to further improve the accuracy of determining the risk calculation results, the terminal device can specifically implement step S104 through the following steps, detailed below: Risk analysis is performed based on location results, voice content, and physiological feature values of different dimensions to obtain risk calculation results.
[0118] In this embodiment, after obtaining the voice content of other people, the terminal device can input the positioning results, physiological feature values of different dimensions, and voice content into the trained risk calculation model for risk analysis to obtain the risk calculation results.
[0119] It should be noted that the risk calculation model can be obtained by training a pre-built second deep learning model based on a second preset sample set. Each sample data in the second preset sample set includes third sample information (including localization results, physiological feature values of different dimensions, and speech content) and the corresponding sample risk calculation result. When training the pre-built second deep learning model, the third sample information in each sample is used as the input to the second deep learning model, and the sample risk calculation result corresponding to the third sample information in each sample is used as the output of the second deep learning model. Through training, the second deep learning model can learn the correspondence between all possible third sample information and sample risk calculation results, and the trained second deep learning model becomes the risk calculation model.
[0120] In another embodiment of this application, the terminal device can specifically be implemented through, as follows: Figure 5The steps S501-S502 shown are implemented based on the positioning results, voice content, and physiological feature values of different dimensions to perform risk analysis and obtain risk calculation results, which are detailed below: In S501, a stress risk score is obtained by weighted summation based on the positioning results, physiological feature values of different dimensions, and voice content. The total weight corresponding to the physiological feature values of different dimensions is greater than the weight corresponding to the voice content, and the weight corresponding to the voice content is greater than the weight corresponding to the positioning results.
[0121] In this embodiment, the terminal device can determine the positional relationship between known individuals and other individuals based on the positioning results, and quantify this positional relationship based on preset position quantization rules to obtain a spatial position risk value. Simultaneously, the terminal device can quantify the speech content using preset speech quantization rules to obtain a semantic coercion score. The preset position quantization rules and preset speech quantization rules can be set according to actual needs and are not limited here.
[0122] Then, the terminal device can perform a weighted summation of the spatial location risk value, physiological characteristic values of different dimensions, and semantic risk value to obtain a stress risk score.
[0123] It should be noted that, since known individuals under covert coercion will exhibit obvious physiological stress responses, such as rapid breathing and trembling voices, and other individuals under covert coercion will exhibit covert threats and control-like speech, such as whispered commands and veiled threats, the terminal device can determine that the total weight of physiological feature values of different dimensions is greater than the weight corresponding to the semantic coercion score, and the weight of the semantic coercion score is greater than the weight corresponding to the spatial location risk value.
[0124] In some possible embodiments, when the location result includes location sub-results for different individuals, the terminal device can specifically implement step S501 through the following steps, detailed below: Based on the location sub-results of different personnel, the spatial location risk value is calculated; Determine semantic coercion scores based on speech content; The physiological fear score is obtained by weighted summation of physiological characteristic values from different dimensions. The stress risk score is obtained by weighting and summing the spatial location risk value, semantic stress score, and physiological fear score.
[0125] In this embodiment, the terminal device can determine the positional relationship between known personnel and other personnel based on the positioning sub-results of different personnel, and quantify the positional relationship based on the preset position quantification rules to obtain the spatial position risk value.
[0126] The terminal device can extract keyword sequences and semantic recognition results from the speech content, and quantify the keyword sequences and semantic recognition results according to the preset speech quantization rules to obtain a semantic stress score.
[0127] Terminal devices can normalize physiological feature values in different dimensions, and then perform weighted summation of the normalized physiological feature values in different dimensions to obtain a physiological fear score.
[0128] Then, the terminal device can normalize the above spatial location risk value, semantic coercion score, and physiological fear score, and then perform a weighted summation of the normalized spatial location risk value, semantic coercion score, and physiological fear score to obtain the coercion risk score.
[0129] In some possible embodiments, the terminal device can calculate the coercion risk score according to the following formula: ; in, The stress risk score is dimensionless and ranges from 100 to 100. ), Represents physiological fear score (dimensionless, range of values: ), Represents semantic coercion score (dimensionless, range of values: ), Spatial location risk value (dimensionless, range: ), This represents the Sigmoid activation function. This indicates the weight corresponding to the physiological fear score (e.g., 0.5). This represents the weight corresponding to the semantic coercion score (e.g., 0.3). The weight corresponding to the spatial location risk value (e.g., 0.2).
[0130] It should be noted that the weighting coefficients satisfy... .
[0131] This embodiment combines three core pieces of information—location results, multi-dimensional physiological feature values, and voice content—to construct a three-dimensional stress risk assessment system that integrates spatial location, semantic content, and physiological state. This system can comprehensively cover the key features of stress scenarios, avoid the one-sidedness of single-dimensional assessments, and improve the comprehensiveness and reliability of stress risk judgment.
[0132] In S502, in response to a stress risk score greater than a risk threshold, the risk calculation result is determined to indicate the existence of hidden stress.
[0133] In this embodiment, after obtaining the coercion risk score, the terminal device can compare the coercion risk score with the risk threshold.
[0134] When a terminal device detects that the stress risk score is greater than the risk threshold, it can determine that the risk calculation result is a hidden stress.
[0135] Combining steps S501-S502, this embodiment integrates three dimensions of information—spatial (location results), physiological (multi-dimensional feature values), and semantic (voice content)—for coercion risk analysis. This allows for comprehensive capture of the characteristic signals of concealed coercion (such as spatial confinement, physiological stress, and cryptic SOS pleas), avoiding missed detections from a single dimension and further improving the accuracy of concealed coercion intent identification. Simultaneously, by using weighted summation to synergistically fuse the features of each dimension, the weights of each dimension can be flexibly adjusted according to different scenarios, adapting to the concealed coercion identification needs in various complex environments (such as vehicle-mounted systems and security systems). This balances recognition accuracy with scenario adaptability, improving system robustness.
[0136] As can be seen from the above, the risk identification method provided in this application provides a method for identifying risks by performing sound source localization on collected voice information to obtain localization results; performing sound source separation on the voice information based on the localization results to obtain voice data corresponding to different individuals; these different individuals include known individuals; analyzing the voice data of known individuals to obtain physiological feature values of known individuals in different dimensions; and performing risk analysis based on the physiological feature values in different dimensions to obtain risk calculation results. This application first achieves decoupling of multi-person voices through sound source localization and sound source separation, which can reduce voice recognition confusion and interference caused by multiple people speaking at the same time, and automatically separates the independent voice data of known individuals from the mixed multi-person voices, improving the accuracy of subsequent analysis; then, it comprehensively analyzes the physiological state of individuals from multiple dimensions, avoiding misjudgment from a single dimension, and improving the accuracy of identifying potential risk states, thereby accurately identifying concealed coercion in scenarios with visual failure (such as blind spots, low light) or concealed behavior (such as whispered threats).
[0137] Please see Figure 6 , Figure 6 This is a flowchart illustrating the implementation of a risk identification method provided in another embodiment of this application. Compared to... Figure 1 In a corresponding embodiment, this embodiment may further include S601~S603 after S102, as detailed below: In S601, feature extraction is performed on the speech data of different people to obtain the initial spectral features of different people.
[0138] In this embodiment, the terminal device can convert the speech data of different people into their respective frequency domain signals using a Short-Time Fourier Transform (STFT) to obtain the complex spectrum (including amplitude spectrum and phase spectrum) corresponding to each frame of speech of different people. Then, the terminal device can extract the core features of the complex spectrum corresponding to different people and concatenate them by channel to form the initial spectral features. The core features may include: log-magnitude spectrum features, phase spectrum features, and Mel spectrum features.
[0139] In practical applications, amplitude spectrum is used to characterize the relative magnitude of sound energy at different frequencies, while phase spectrum features are used to characterize the temporal position and superposition sequence of sound waves at each frequency.
[0140] Mel spectrum characteristics are obtained by weighting and filtering the amplitude / power spectrum through a Mel filter bank.
[0141] In some possible embodiments, the terminal device can time-align the initial spectral features corresponding to all voice frames of different people, and then splice them together in the order of the voice frames to form a feature sequence.
[0142] In S602, the initial spectral features of different individuals are input into a trained deep spectral mapping network for enhancement processing to obtain the target spectral features of different individuals.
[0143] In this embodiment, after obtaining the initial spectral features of different individuals, the terminal device can sequentially input the initial spectral features of different individuals into the trained Deep Spectrum Mapping Net for enhancement processing to obtain the target spectral features of different individuals.
[0144] It should be noted that the aforementioned deep spectrum mapping network can adopt the U-Net architecture, which includes three main modules: an encoder, a decoder, and skip connections. The encoder can consist of four layers of a convolutional neural network (CNN) to progressively extract deep semantic information from the initial spectral features and suppress noise features. The decoder can consist of four layers of a deconvolutional network to map the deep features back to the original spectral dimension, restoring speech details. Skip connections are used to fuse the shallow features from each layer of the encoder with the deep features from the corresponding layers of the decoder, improving the accuracy of spectrum reconstruction.
[0145] Specifically, for any individual's initial spectral features, the terminal device can convert these features into tensors (e.g., dimensions of [batch size, number of frames, feature dimension, number of channels]) according to the network's requirements and input them into the aforementioned deep spectrum mapping network. The encoder of the deep spectrum mapping network can perform convolution and pooling operations on the input initial spectral features to extract deep, effective features and filter out noise-related features. Then, through skip connections in the deep spectrum mapping network, the deep, effective features extracted by the encoder can be fed into the decoder. The decoder can gradually recover the spectral dimension through deconvolution operations, while simultaneously fusing deep and shallow features to output the initially enhanced spectral features. Subsequently, the deep spectrum mapping network can optimize the output initially enhanced spectral features using a spectral smoothing algorithm (e.g., Gaussian smoothing) to eliminate spectral jitter and correct spectral distortion that occurs during the enhancement process. Simultaneously, threshold filtering is used to remove abnormal spectral feature points, resulting in the output of the optimized target spectral features corresponding to the initial spectral features of any individual.
[0146] It should be noted that the dimensions and number of frames of the target spectral features are consistent with the input initial spectral features, including the enhanced amplitude spectrum, phase spectrum and Mel spectrum features, and retaining the original semantic and physiological information of the speech.
[0147] In some possible embodiments, the terminal device may obtain the target spectral characteristics of any person in the following ways: ; in, Let represent the target spectral characteristics of any person, and let represent the initial spectral characteristics of any person. Represents a deep spectrum mapping network. This represents the network parameters of the deep spectrum mapping network. The low-frequency spectral amplitude represents the initial spectral characteristics of any individual. The low-frequency spectral phase represents the initial spectral characteristics of any individual.
[0148] In S603, based on the target spectral characteristics of different people, the enhanced speech data of each person is obtained.
[0149] In this embodiment, after obtaining the target spectral features of different people, the terminal device can inversely convert the target spectral features of different people into time-domain speech signals to obtain the enhanced speech data of each person.
[0150] Specifically, for any given person, the terminal device can perform an inverse short-time fourier transform (ISTFT) on the amplitude and phase spectra of the target spectral features of that person to convert the frequency domain target spectral features into individual time-domain speech frames. Then, the terminal device can splice the time-domain speech frames obtained from the inverse transform and use an overlapping addition method to eliminate inter-frame splicing traces to obtain a complete time-domain speech signal.
[0151] It should be noted that the parameters of the ISTFT above need to be consistent with the STFT parameters used to convert the speech data of different people into frequency domain signals, so as to ensure that the speech signal after inverse conversion has the same duration as the original speech.
[0152] In one embodiment of this application, after step S603, the terminal device can implement step S103 through the enhanced voice data of the known person, that is, analyze the enhanced voice data of the known person to obtain the physiological characteristic values of the known person in different dimensions.
[0153] In another embodiment of this application, after step S603, the terminal device can perform speech recognition on the speech data of other people through the enhanced speech data of other people to obtain the speech content corresponding to other people, that is, perform speech recognition on the enhanced speech data of other people to obtain the speech content of other people.
[0154] As can be seen from the above, the risk identification method provided in this embodiment, after obtaining the speech data corresponding to different individuals, can extract features from the speech data of different individuals to obtain the initial spectral features of different individuals; input the initial spectral features of different individuals into a trained deep spectral mapping network for enhancement processing to obtain the target spectral features of different individuals; and based on the target spectral features of different individuals, obtain the enhanced speech data of each individual. Since the whispers commonly seen in coercive scenarios usually lack the fundamental frequency and harmonic structure of vocal cord vibration, this embodiment uses a spectral mapping network to learn the mapping relationship from the "whisper spectrum" to the "normal speech spectrum" in order to restore the missing low-frequency formant structure and make the whisper content easier to identify.
[0155] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0156] Please see Figure 7 , Figure 7 This is a schematic diagram of the overall process of an intent recognition system provided in an embodiment of this application. It should be noted that the intent recognition system can achieve the above-described... Figures 1-6 Any corresponding implementation example.
[0157] like Figure 7 As shown, the intent recognition system may specifically include: a signal acquisition layer 71, a sound field reconstruction layer 72, a separation enhancement layer 73, a feature analysis layer 74, and a fusion discrimination layer 75. Specifically, the signal acquisition layer 71 is connected to the sound field reconstruction layer 72, the sound field reconstruction layer 72 is connected to both the separation enhancement layer 73 and the fusion discrimination layer 75, the separation enhancement layer 73 is connected to the feature analysis layer 74, and the feature analysis layer 74 is connected to the fusion discrimination layer 75.
[0158] In this embodiment, the signal acquisition layer 71 can simultaneously acquire multi-channel audio (such as voice information) through a microphone array disposed in the target area, and preprocess the acquired multi-channel audio. The preprocessing may include echo cancellation and automatic gain control.
[0159] After obtaining the preprocessed multi-channel audio, the sound field reconstruction layer 72 can construct a 3D sound source distribution map of the target area through spatial scanning of the phase transformation weighted generalized cross-correlation algorithm and fine localization of the multi-signal classification algorithm, track the sound source location in real time, and perform regional mapping on the sound source location to obtain the localization result.
[0160] The separation enhancement layer 73 can pick up the target domain speech signals of different people based on beamforming and perform blind source separation using independent component analysis to obtain speech data of different people. Then, the speech data of different people is enhanced by a deep spectrum mapping network to enhance the whispering features and obtain the enhanced speech data of different people.
[0161] The feature analysis layer 74 can perform microvibration detection, respiratory analysis, and baseline deviation calculation on the enhanced speech data of known individuals to obtain physiological feature values such as Teager energy microvibration, respiratory tachycardia, and baseline deviation values corresponding to the known individuals.
[0162] Meanwhile, the feature analysis layer 74 can also perform semantic keyword detection on the enhanced speech data of other people.
[0163] Subsequently, the feature analysis layer 74 can calculate the physiological fear score in the physiological risk dimension based on the physiological feature values of different dimensions, calculate the semantic coercion score in the semantic risk dimension based on the results of semantic keyword detection, and calculate the spatial location risk value in the spatial risk dimension based on the location sub-results of other people.
[0164] The fusion discrimination layer 75 can perform multi-dimensional fusion, i.e. weighted fusion, on the scores corresponding to the semantic risk dimension, physiological risk dimension, and spatial risk dimension to obtain the final coercion risk score. The coercion risk score is then compared with the risk threshold to determine whether the risk calculation result is a hidden coercion.
[0165] Corresponding to the risk identification method described in the above embodiments, Figure 8 A schematic diagram of a risk identification device provided in an embodiment of this application is shown. For ease of explanation, only the parts relevant to the embodiment of this application are shown. (Refer to...) Figure 8 The risk identification device 800 includes: a sound source localization unit 81, a sound source separation unit 82, a voice analysis unit 83, and a risk analysis unit 85. Wherein: The sound source localization unit 81 is used to locate the sound source of the collected speech information and obtain the localization result.
[0166] The sound source separation unit 82 is used to separate the sound sources of the speech information based on the positioning results, and obtain the speech data corresponding to different people; different people include known people.
[0167] The speech analysis unit 83 is used to analyze the speech data of known persons to obtain the physiological characteristic values of the known persons in different dimensions.
[0168] Risk analysis unit 84 is used to perform risk analysis based on physiological characteristic values of different dimensions and obtain risk calculation results.
[0169] In one embodiment of this application, "different persons" also includes other persons; the risk analysis unit 84 specifically includes: a speech recognition unit and a risk analysis subunit. Wherein: The speech recognition unit is used to perform speech recognition on the speech data of other people to obtain the corresponding speech content of other people.
[0170] The risk analysis subunit is used to perform risk analysis based on speech content and physiological feature values of different dimensions to obtain risk calculation results.
[0171] In one embodiment of this application, the risk analysis unit 84 specifically includes a risk calculation unit.
[0172] The risk calculation unit is used to perform risk analysis based on the location results, voice content, and physiological feature values of different dimensions to obtain the risk calculation results.
[0173] In one embodiment of this application, the risk calculation unit specifically includes: a summation unit and a risk determination unit. Wherein: The summation unit is used to perform weighted summation based on the localization results, physiological feature values of different dimensions, and speech content to obtain a stress risk score; the total weight corresponding to the physiological feature values of different dimensions is greater than the weight corresponding to the speech content, and the weight corresponding to the speech content is greater than the weight corresponding to the localization results.
[0174] The risk determination unit is used to determine that a hidden threat exists in response to a stress risk score that is greater than a risk threshold.
[0175] In one embodiment of this application, physiological characteristic values of different dimensions include microfracture characteristic values, respiratory tachycardia, and baseline deviation values; the speech analysis unit 83 specifically includes: a microfracture detection unit, a segmentation unit, a respiratory tachycardia determination unit, and a deviation calculation unit. Wherein: The microvibration detection unit is used to detect microvibrations in the speech data of known individuals based on nonlinear signal processing tools, and to obtain microvibration feature values.
[0176] The segmentation unit is used to divide the speech data of known personnel into speech segments and non-speech segments.
[0177] The respiratory tachymetry determination unit is used to determine respiratory tachymetry based on non-speech segments and speech segments.
[0178] The deviation calculation unit is used to calculate the baseline deviation value based on the speech data of known personnel and the standard speech baseline of known personnel.
[0179] In one embodiment of this application, the respiratory tachycardia includes frequency perturbation value, amplitude perturbation value, and harmonic-to-noise ratio; the deviation calculation unit specifically includes: a combination unit and a distance calculation unit. Wherein: The combination unit is used to combine micro-vibration feature values, frequency perturbation values, amplitude perturbation values, and harmonic-to-noise ratio to obtain physiological feature vectors.
[0180] The distance calculation unit is used to calculate Mahalanobis distance based on physiological feature vectors and standard speech baselines to obtain baseline deviation values.
[0181] In one embodiment of this application, the risk identification device 800 further includes: a feature extraction unit, an enhancement unit, and a voice data determination unit; correspondingly, the voice analysis unit 83 specifically includes a data analysis unit. Wherein: The feature extraction unit is used to extract features from the speech data of different people to obtain the initial spectral features of different people.
[0182] The enhancement unit is used to input the initial spectral features of different people into a trained deep spectral mapping network for enhancement processing, so as to obtain the target spectral features of different people.
[0183] The speech data determination unit is used to obtain the enhanced speech data of different individuals based on the target spectral characteristics of different individuals.
[0184] The data analysis unit is used to analyze the enhanced speech data of known individuals to obtain physiological characteristic values of the known individuals in different dimensions.
[0185] In one embodiment of this application, a microphone array is installed inside the vehicle; the sound source localization unit 81 specifically includes: a preliminary localization unit, a secondary localization unit, and a position mapping unit. Wherein: The preliminary localization unit is used to perform preliminary localization of speech information based on the phase transform weighted generalized cross-correlation algorithm and microphone array to obtain the sound production area.
[0186] The relocalization unit is used to relocalize speech information based on a multi-signal classification algorithm and the speech region to obtain the speech location.
[0187] The location mapping unit is used to map the sound source location to the target area where different people are located, so as to obtain the location result.
[0188] In one embodiment of this application, the positioning result includes positioning sub-results for different personnel; the sound source separation unit 82 specifically includes: a signal output unit and a separation sub-unit. Wherein: The signal output unit is used to construct a minimum variance distortion-free response beamformer based on voice information and the localization results of different people, and output the target domain voice signals of different people.
[0189] The separation subunit is used to separate the target domain speech signals of different people through independent component analysis, so as to obtain the speech data corresponding to each person.
[0190] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0191] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0192] Figure 9 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Figure 9 As shown, the terminal device 9 of this embodiment includes: at least one processor 90 ( Figure 9 (Only one is shown in the diagram), memory 91, and computer program 92 stored in said memory 91 and executable on said at least one processor 90, wherein said processor 90 executes said computer program 92 to implement the steps in any of the above-described risk identification method embodiments.
[0193] The terminal device 9 may include, but is not limited to, a processor 90 and a memory 91. Those skilled in the art will understand that... Figure 9 This is merely an example of terminal device 9 and does not constitute a limitation on terminal device 9. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0194] The processor 90 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0195] In some embodiments, the memory 91 may be an internal storage unit of the terminal device 9, such as the RAM of the terminal device 9. In other embodiments, the memory 91 may be an external storage device of the terminal device 9, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal device 9. Furthermore, the memory 91 may include both internal and external storage units of the terminal device 9. The memory 91 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 91 can also be used to temporarily store data that has been output or will be output.
[0196] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0197] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.
[0198] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0199] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0200] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A risk identification method, characterized in that, include: The collected speech information is used to locate the sound source, and the location result is obtained; Based on the positioning results, the voice information is separated into sound sources to obtain voice data corresponding to different people; the different people include known people. The voice data of the known individuals are analyzed to obtain their physiological characteristic values in different dimensions. Risk analysis is performed based on the different dimensions of physiological characteristics to obtain risk calculation results.
2. The risk identification method as described in claim 1, characterized in that, The "different personnel" also include other personnel; the risk analysis based on the different dimensions of physiological characteristic values, to obtain the risk calculation results, includes: Speech recognition is performed on the speech data of the other persons to obtain the speech content corresponding to the other persons; Risk analysis is performed based on the voice content and the physiological feature values of different dimensions to obtain the risk calculation results.
3. The risk identification method as described in claim 2, characterized in that, The risk analysis based on the speech content and the physiological feature values of different dimensions, to obtain the risk calculation result, includes: Risk analysis is performed based on the positioning results, the voice content, and the physiological feature values of different dimensions to obtain the risk calculation results.
4. The risk identification method as described in claim 3, characterized in that, The risk calculation result obtained by performing risk analysis based on the positioning result, the voice content, and the physiological feature values of different dimensions includes: A stress risk score is obtained by weighting and summing the location results, the physiological feature values of different dimensions, and the voice content; the total weight corresponding to the physiological feature values of different dimensions is greater than the weight corresponding to the voice content, and the weight corresponding to the voice content is greater than the weight corresponding to the location results. In response to the stress risk score being greater than a risk threshold, the risk calculation result is determined to indicate the existence of hidden stress.
5. The risk identification method according to any one of claims 1-4, characterized in that, The different dimensions of physiological characteristics include microfibrillation characteristics, respiratory tachycardia, and baseline deviation. The analysis of the voice data of the known individuals to obtain their physiological characteristic values in different dimensions includes: The speech data of the known individuals are subjected to micro-vibration detection using nonlinear signal processing tools to obtain the micro-vibration feature values. The voice data of the known individuals is divided into voice segments and non-voice segments; The respiratory tachycardia is determined based on the non-speech segment and the speech segment; The baseline deviation value is calculated based on the microfracture feature value, the respiratory tachycardia, and the standard speech baseline of the known person.
6. The risk identification method as described in claim 5, characterized in that, The respiratory tachycardia includes frequency perturbation value, amplitude perturbation value, and harmonic-to-noise ratio; The baseline deviation value is calculated based on the microfibrillation feature value, the respiratory tachycardia, and the standard speech baseline of the known person, including: The physiological feature vector is obtained by combining the micro-vibration feature value, the frequency perturbation value, the amplitude perturbation value, and the harmonic-to-noise ratio. The baseline deviation value is obtained by calculating Mahalanobis distance based on the physiological feature vector and the standard speech baseline.
7. The risk identification method according to any one of claims 1-4, characterized in that, After obtaining the voice data corresponding to each individual, the method further includes: Feature extraction is performed on the speech data of the different individuals to obtain the initial spectral features of the different individuals; The initial spectral features of the different individuals are input into a trained deep spectral mapping network for enhancement processing to obtain the target spectral features of the different individuals. Based on the target spectral characteristics of the different individuals, the enhanced speech data of each individual is obtained; Accordingly, the analysis of the voice data of the known person to obtain the physiological characteristic values of the known person in different dimensions includes: The enhanced speech data of the known individuals is analyzed to obtain their physiological characteristic values in different dimensions.
8. The risk identification method according to any one of claims 1-4, characterized in that, The voice information is collected through a microphone array; the process of locating the sound source from the collected voice information to obtain the location result includes: The speech information is initially located based on the phase transform weighted generalized cross-correlation algorithm and the microphone array to obtain the sound production area; The speech information is re-localized based on the multiple signal classification algorithm and the speech region to obtain the speech location; The location of the sound is mapped to the target area where the different people are located to obtain the positioning result.
9. The risk identification method according to any one of claims 1-4, characterized in that, The positioning result includes positioning sub-results for different individuals; the step of performing sound source separation on the speech information based on the positioning result to obtain speech data corresponding to each individual includes: Based on the voice information and the location sub-results of the different people, a minimum variance distortion-free response beamformer is constructed to output the target domain voice signals of the different people. Independent component analysis was used to separate the sound sources of the target domain speech signals of the different individuals, thereby obtaining the speech data corresponding to each individual.
10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the risk identification method as described in any one of claims 1 to 9.