Voice signal extraction and recognition method, device, electronic device and storage medium

By real-time monitoring and analysis of fuzzy audio signals, deviated audio signals are extracted and combined with emotion recognition models, the problem of difficult to identify events in public places in the prior art is solved, precise judgment of the nature of the event and the location of occurrence is achieved, and the efficiency of security management is improved.

CN119724246BActive Publication Date: 2025-06-06SHANGHAI HUITONGSHENG INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510239987.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-06
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The prior art is difficult to identify events in public places based on fuzzy audio signals, especially in the case of noise interference and noisy environment, and it is difficult to accurately determine whether disturbances have occurred.

Method used

By monitoring the sound signals of the target environment in real time, using the sliding window method to obtain the audio signal and compare it with the preset audio, extract the deviation audio signal, input the recognition model for type recognition and sound source positioning, and analyze the various characteristics of the audio signal in combination with the emotional recognition model to determine the nature and location of the event.

Benefits of technology

It can extract key features from complex and fuzzy audio signals, reduce the risk of misjudgment, accurately judge the nature and location of events, improve the efficiency of identifying the nature of events, and provide an effective reference for security management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724246B_ABST
    Figure CN119724246B_ABST
Patent Text Reader

Abstract

The present invention provides a speech signal extraction and recognition method, device, electronic device and storage medium, and belongs to the field of audio processing technology. The speech signal extraction and recognition method of the present invention, through layer-by-layer analysis of audio signals, from deviation detection to complex emotion analysis and positioning, ensures that even when the audio signal is complex or fuzzy, it can still gradually focus on and extract key features to reduce the risk of misjudgment. By comprehensively analyzing multiple features of the audio signal, the nature and location of the event can be accurately inferred. The audio signal is used for positioning and emotion color analysis, which increases the efficiency of identifying the nature of the event. Clues can be found from complex and fuzzy audio signals to determine the type of event, thereby providing a reference for security management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio processing technology, and in particular to a method, device, electronic equipment and storage medium for extracting and recognizing a speech signal. Background Art

[0002] In some public places, such as office buildings, hospitals, airports and other crowded places, security monitoring systems are particularly important to ensure public safety and order. These places usually require effective management and monitoring of large crowds in order to promptly detect and respond to various possible disruptive behaviors. However, although modern security systems have made significant progress in image monitoring, there are still many difficulties and limitations in relying solely on image monitoring technology to identify and judge such events as loud noises and disputes.

[0003] Image surveillance systems usually rely on video signals to detect and identify abnormal behaviors. However, the advantages of image surveillance are mainly reflected in the ability to identify objects, track people's activities, and detect crowd density. However, for some behaviors triggered by sound, image systems often cannot provide direct clues. For example, two or more people may have a dispute or quarrel for some reason, but if the manifestation of these behaviors is not obvious, such as no obvious physical conflict or large movements, the image system may find it difficult to detect them in time.

[0004] Therefore, in this case, the introduction of audio monitoring becomes an effective supplementary means. Audio signals can provide important supplementary information for security systems, especially in scenes that cannot be captured visually. Audio can promptly reflect whether an abnormality has occurred through changes in sound. For example, in some public places, the intensity of the sound, changes in frequency, tone and emotion of speech may all be signals of disputes or conflicts. Audio signals can help the monitoring system capture these changes, thereby identifying potential threats or disruptive behaviors in a timely manner.

[0005] However, the application of audio signals is not without challenges. Unlike images, audio signals are characterized by a certain degree of ambiguity, and it is often difficult to directly obtain accurate content information. There are usually a large number of noise sources in public places, such as the sound of people talking, background music, traffic noise, etc., which can seriously interfere with the recognition of audio signals. Especially when the sound source is complex or the environment is noisy, it is often difficult to recognize the specific content of each word spoken by the audio signal, and for the sake of privacy protection, it is also a big challenge to obtain the specific content of the person's voice information.

[0006] Therefore, although audio signals provide surveillance systems with more identification means, their ambiguity and complexity make it quite difficult to determine whether an incident that disrupts the public environment has occurred based solely on audio. Summary of the invention

[0007] The present invention provides a speech signal extraction and recognition method, device, electronic device and storage medium, which are used to solve the defect in the prior art that it is difficult to identify events in an environment based on fuzzy audio, and achieve the effect of finding clues from complex and fuzzy audio signals to determine the type of event.

[0008] The present invention provides a method for extracting and recognizing a speech signal, comprising:

[0009] Performing real-time monitoring of sound signals in a target environment, acquiring a first audio signal based on a sliding window method, and comparing the first audio signal with a preset audio signal corresponding to the target environment;

[0010] In the case where there is a target deviation between the first audio signal and the preset audio, extracting a second audio signal including the audio signal with the deviation from the first audio signal;

[0011] Inputting the second audio signal into a recognition model to identify the type of the second audio signal;

[0012] Identifying a sound source position of the second audio signal, and determining that the sound source position of the second audio signal matches the target environment;

[0013] In a case where it is determined based on the type of the second audio signal that the second audio signal contains human voice, identifying the position of a person object making a sound in the second audio signal and the emotional color of the sound corresponding to each person object;

[0014] Based on the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object, it is determined that a target event occurs at a target position of the target environment.

[0015] According to a speech signal extraction and recognition method provided by the present invention, the identifying the position of the person object making the sound in the second audio signal and the emotional color of the sound corresponding to each person object includes:

[0016] Performing initial emotion recognition on the second audio signal to obtain initial emotion labels of different frequency bands in the second audio signal;

[0017] Based on the initial emotion label, the sound source position of the second audio signal and the position of the microphone group arranged at the sound source position, the second audio signal is separated into sound sources, audio sub-signals corresponding to the various human objects emitting sounds are extracted, and the sound positions of the various human objects corresponding to the various audio sub-signals are determined;

[0018] Each audio sub-signal is input into the emotion recognition model to obtain the emotional color of the sound corresponding to each person object corresponding to each audio sub-signal.

[0019] According to a speech signal extraction and recognition method provided by the present invention, based on the initial emotion label, the sound source position of the second audio signal and the position of the microphone group arranged at the sound source position, the second audio signal is separated from the sound source to extract the audio sub-signals corresponding to the various human objects that make sounds, including:

[0020] Determine the positions of the respective sound sources in the second audio signal by a beamforming algorithm based on the sound source position of the second audio signal, the positions of the microphone group arranged at the sound source position, and the time difference between the sound reaching different microphones of the microphone group;

[0021] Based on the positions of each sound source in the second audio signal and the correlation of the emotional color levels of the initial emotional labels corresponding to different frequency bands in the second audio signal, audio sub-signals corresponding to each person object at the position of the sound source are extracted from the second audio signal.

[0022] According to a speech signal extraction and recognition method provided by the present invention, based on the positions of each sound source in the second audio signal and the correlation of the emotional color levels of the initial emotional labels corresponding to different frequency bands in the second audio signal, extracting the audio sub-signals corresponding to each person object at the position of the sound source from the second audio signal includes:

[0023] Based on the position of each sound source in the second audio signal, determine the audio signal at the position of each sound source;

[0024] Based on the correlation of the emotional color levels of the initial emotional labels corresponding to different frequency bands in the second audio signal, the audio source signal after the audio signal at each sound source position is decomposed is evaluated to obtain the audio sub-signal corresponding to each person object.

[0025] According to a speech signal extraction and recognition method provided by the present invention, the method of determining the occurrence of a target event at a target position of the target environment based on the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object includes:

[0026] Based on the level of the emotional color of the voice corresponding to each person object, determining that there are two adjacent target person objects whose level evaluation results of the emotional color meet the preset conditions;

[0027] Based on the positions of the two target personnel objects, determining that the distance between the two target personnel objects is less than the target distance;

[0028] Determine that a target event occurs at a target location of the target environment.

[0029] According to a speech signal extraction and recognition method provided by the present invention, the target event includes at least one of quarreling, arguing and fighting.

[0030] According to a speech signal extraction and recognition method provided by the present invention, the target deviation is determined based on at least one of frequency, intensity and timbre.

[0031] The present invention also provides a speech signal extraction and recognition device, comprising:

[0032] An acquisition module, used to monitor the sound signal in the target environment in real time, acquire the first audio signal based on a sliding window method, and compare it with the preset audio corresponding to the target environment;

[0033] an extraction module, configured to extract a second audio signal including an audio signal with a deviation from the first audio signal when there is a target deviation between the first audio signal and the preset audio;

[0034] A first recognition module, configured to input the second audio signal into a recognition model to identify a type of the second audio signal;

[0035] A second recognition module, used to identify the sound source position of the second audio signal, and determine whether the sound source position of the second audio signal matches the target environment;

[0036] A third recognition module is used to identify the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object in the second audio signal when it is determined that the second audio signal contains human voice based on the type of the second audio signal;

[0037] The processing module is used to determine the occurrence of a target event at a target position of the target environment based on the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object.

[0038] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the speech signal extraction and recognition method described above is implemented.

[0039] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the speech signal extraction and recognition method as described in any one of the above is implemented.

[0040] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the speech signal extraction and recognition method described above is implemented.

[0041] The speech signal extraction and recognition method, device, electronic device and storage medium provided by the present invention analyze the audio signal layer by layer, from deviation detection to complex emotion analysis and positioning, to ensure that even when the audio signal is complex or ambiguous, it can still gradually focus on and extract key features to reduce the risk of misjudgment. By comprehensively analyzing multiple features of the audio signal, the nature and location of the event can be accurately inferred. The audio signal is used for positioning and emotion color analysis, which increases the efficiency of identifying the nature of the event. Clues can be found from complex and ambiguous audio signals to determine the type of event, thereby providing a reference for security management. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 This is one of the flow charts of the speech signal extraction and recognition method provided by the present invention;

[0044] Figure 2 This is the second flow chart of the speech signal extraction and recognition method provided by the present invention;

[0045] Figure 3 It is a structural schematic diagram of the speech signal extraction and recognition device provided by the present invention;

[0046] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0048] Combine the following Figure 1-Figure 4 The invention describes the speech signal extraction and recognition method, device, electronic device and storage medium.

[0049] like Figure 1As shown, the speech signal extraction and recognition method of the embodiment of the present invention mainly includes step 110, step 120, step 130, step 140, step 150 and step 160.

[0050] Step 110 , real-time monitoring of sound signals in the target environment, obtaining a first audio signal based on a sliding window method, and comparing the first audio signal with a preset audio signal corresponding to the target environment.

[0051] The target environment refers to the actual environment that the monitoring system focuses on, such as the public area of ​​a floor of an office building and an area of ​​an airport.

[0052] The sliding window method is a technique commonly used in signal processing. It sets a window of fixed time length to slide continuously in the time series and extracts signal features in each window for analysis. In this step, the sliding window method is used to extract audio clips of a certain length from real-time sound signals.

[0053] The preset audio is a predefined standard audio signal, which may be a typical sound pattern in the target environment, such as regular noise of a crowd or traffic sound.

[0054] The sliding window method includes the fixed-length sliding window method and the variable-length sliding window method. The fixed-length sliding window method can use a fixed-length time window to gradually analyze the characteristics of the audio. Each time the window slides a small step, new features are extracted for comparison and analysis.

[0055] The variable-length sliding window method is different from the fixed-length sliding window method in that the variable-length sliding window allows the length of the window to change as the content and characteristics of the audio signal change. That is, the window size is not pre-set, but is adaptively adjusted according to the characteristics of the real-time audio.

[0056] In the process of real-time audio monitoring, the main task is to identify abnormal or suspicious audio events, such as sudden high-decibel sounds, noise of a specific frequency, abnormal sound patterns, etc., and then identify special events. By comparing with the preset normal audio template, the sound that deviates from the normal pattern can be quickly detected.

[0057] In some embodiments, features such as frequency spectrum, time-frequency features, voice activity detection of speech, etc. can be extracted from the audio signal to help analyze the content and changes of the audio signal. The audio features monitored in real time are compared with the preset normal audio features to determine whether there are any abnormalities.

[0058] It should be noted that the features in the audio signal may change dramatically in a short period of time, such as a sudden loud scream or explosion. At this time, using a fixed-length sliding window may not be able to capture these changes in time, while a variable-length sliding window can automatically adjust the analysis time window according to the changes in the audio features. For example, when an emergency occurs, the length of the sliding window can be quickly increased to include more sound data for in-depth analysis.

[0059] In the stable part of the audio signal, the sliding window length can be small to quickly process the audio stream; while in the area where the audio signal changes greatly, the sliding window length can be increased to more comprehensively analyze and capture these changes. Certain special sounds, such as violent collisions and sharp screams, may contain complex frequency patterns in a short period of time, and using a longer window helps capture more contextual information.

[0060] It is understandable that the environmental sound is monitored in real time by the sliding window method, and a short-term first audio signal is obtained. The first audio signal is compared with the preset audio to identify abnormalities or deviations in the sound pattern. The preset audio provides a benchmark to help quickly determine whether the current audio is consistent with the normal audio pattern of the target environment.

[0061] It is understandable that through real-time monitoring and comparison, deviations in audio signals can be effectively captured, laying the foundation for subsequent event identification. The sliding window method processes real-time signals with high time sensitivity, which helps to promptly detect abnormal audio in the environment.

[0062] Step 120: When there is a target deviation between the first audio signal and the preset audio, extract a second audio signal including the audio signal with the deviation from the first audio signal.

[0063] Deviation audio signals are audio signals that are significantly different from preset audio signals, including unexpected noise, abnormal sound sources, or audio manifestations of sudden events.

[0064] The second audio signal is an audio signal segment containing a deviation extracted from the first audio signal.

[0065] When there is a significant deviation between the first audio signal and the preset audio, the deviation audio parts can be focused on to extract the audio segments containing the abnormal signals, namely the second audio signal.

[0066] The target deviation is determined based on at least one of frequency, intensity, and timbre.

[0067] Frequency is a fundamental parameter in audio signals that determines the pitch. Target deviations can be identified based on frequency differences. For example, if the audio signal in a certain frequency range deviates from the frequency pattern of the preset audio, this could be a deviation.

[0068] Intensity refers to the loudness of the audio. The deviation audio signal may be different in loudness from the preset audio, for example, some parts are louder or quieter than the preset audio, and the deviation can also be identified by the change in intensity.

[0069] Timbre is the texture of sound and determines the color of the sound. Even if two audio clips have the same frequency and intensity, their timbres may be different, for example, due to different sound sources or environmental changes. Changes in timbre can also be a sign of target deviation.

[0070] In practical applications, the audio signal may be subjected to frequency analysis, intensity analysis, and timbre analysis, and the difference from the preset audio may be detected.

[0071] Therefore, the target deviation is determined based on at least one of frequency, intensity, and timbre, i.e., when analyzing the audio, one or more of these features are checked to see if they are significantly different from the preset audio. If a significant deviation is found in frequency, intensity, or timbre, it can be determined that the audio signal contains a deviation, thereby extracting a second audio segment containing an abnormal signal. For example, in some cases, the target deviation may be determined solely based on the frequency change of the sound; in other cases, changes in timbre or intensity may also need to be considered.

[0072] It can be understood that by extracting the part containing the deviation from the audio signal, the potential event signal can be effectively screened out from the noise.

[0073] Step 130: Input the second audio signal into a recognition model to identify the type of the second audio signal.

[0074] Recognition models are usually models trained based on machine learning, deep learning and other technologies. They are used to analyze audio features and determine the type of audio signals, such as human voices, animal sounds, and mechanical sounds.

[0075] The audio signal type refers to the classification result of the audio signal, such as whether it belongs to human voice, environmental noise or mechanical equipment.

[0076] The second audio signal can be input into a trained recognition model, which analyzes the signal and identifies the type of the signal. At this time, it can be quickly and easily identified whether the signal comes from human conversation, quarrel, object collision, or other environmental noise.

[0077] In some embodiments, sound classification can be performed based on deep learning. The second audio signal after noise analysis can be input into a pre-trained deep neural network such as ResNet or LSTM for event classification. The model can distinguish different types of sounds, including people talking, fighting, and glass breaking.

[0078] In other embodiments, audio event calibration can be identified by setting a threshold. According to different types of sounds, corresponding intensity thresholds are set. For example, sounds exceeding a certain decibel number and having certain frequency characteristics are set as human conversation sounds. When the intensity of the sound event exceeds the threshold, it is automatically marked as potential noise.

[0079] In this embodiment, the event can be preliminarily identified by identifying the audio type, and then the influence of human voice and other environmental noise on the judgment of the event can be distinguished.

[0080] Inputting audio signals into the recognition model can automatically complete audio classification and feature extraction, reducing the workload of manual judgment.

[0081] Step 140: Identify the sound source position of the second audio signal and determine whether the sound source position of the second audio signal matches the target environment.

[0082] The sound source position refers to the approximate location where the audio signal is emitted, which may be determined by a multi-microphone array or sound source localization technology.

[0083] Target environment matching refers to whether the source of the sound is consistent with the location of the target environment scene.

[0084] The source position of the second audio signal can be determined by sound source localization technology and compared with the specific position in the target environment. If the second audio signal comes from a specific area of ​​the target environment that needs to be monitored, the audio needs to be identified and analyzed.

[0085] It can be understood that, through the sound source localization in this step, it is possible to preliminarily determine whether the origin of the audio signal is in the target environment of the monitoring area.

[0086] Step 150: When it is determined based on the type of the second audio signal that the second audio signal contains human voice, identify the position of the person object making the sound in the second audio signal and the emotional color of the sound corresponding to each person object.

[0087] Emotional color refers to the emotional state in the audio signal, such as anger, anxiety, surprise, happiness, etc. Person objects refer to the specific people who make sounds in the audio signal. The system needs to identify the source of these sounds through audio analysis and perform emotional analysis.

[0088] After confirming that the audio signal belongs to a human voice, the emotional color of the voice can be further analyzed. The emotional state of the speaker can be judged through the voice's pitch, speaking speed, volume and other characteristics. At the same time, the location of each sound emitter can be located to provide more contextual information for event analysis.

[0089] It is understandable that sentiment analysis of audio signals can help determine whether the event corresponding to the audio is a target event.

[0090] Step 160 , based on the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object, determine that a target event occurs at a target position of the target environment.

[0091] The target location refers to the specific location where an event may occur, inferred from audio analysis.

[0092] A target event refers to an abnormal event that is identified and may require response measures to be taken. The target event may include at least one of quarrels, disputes, and fights.

[0093] The specific location and nature of the incident can be inferred by combining multiple elements in the second audio signal, such as the location of the person and the emotional color, with the scene information of the target environment. If the emotional color is anger or panic, it can be judged as an emergency situation such as violent conflict in words or actions, and further attention needs to be paid to the incident in this situation.

[0094] The speech signal extraction and recognition method provided by the embodiment of the present invention analyzes the audio signal layer by layer, from deviation detection to complex emotion analysis and positioning, to ensure that even when the audio signal is complex or ambiguous, it can still gradually focus on and extract key features to reduce the risk of misjudgment. By comprehensively analyzing the various features of the audio signal, the nature and location of the event can be accurately inferred. The audio signal is used for positioning and emotion color analysis, which increases the efficiency of identifying the nature of the event. Clues can be found from complex and ambiguous audio signals to determine the type of event, thereby providing a reference for security management.

[0095] In some embodiments, Figure 2 As shown, identifying the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object includes step 210, step 220 and step 230.

[0096] Step 210: Perform initial emotion recognition on the second audio signal to obtain initial emotion labels of different frequency bands in the second audio signal.

[0097] First, the second audio signal is subjected to initial emotion recognition, and an initial emotion label is assigned to each sound signal.

[0098] At this time, emotion recognition can be based on the pitch, intonation, speech speed, timbre and other features in the audio to simply judge the speaker's emotional state, such as happiness, anger, sadness, anxiety, etc. In this embodiment, the initial emotion recognition is performed to simply obtain the difference between different frequency bands, that is, the initial emotion label is used as a classification label to identify different sound sources.

[0099] The second audio signal may contain information of multiple frequency bands, and the audio signal of each frequency band may represent different sound sources or emotional expressions. Emotion recognition can assign initial emotional labels to these different frequency bands. For example, the sound in a certain frequency band can reflect the emotional state of a person.

[0100] Step 220, based on the initial emotion label, the sound source position of the second audio signal and the position of the microphone group arranged at the sound source position, the second audio signal is separated from the sound source, the audio sub-signals corresponding to each person object that makes the sound are extracted, and the sound position of each person object corresponding to each audio sub-signal is determined.

[0101] Sound source separation refers to separating multiple sound sources in an audio signal, such as the voices of different people, from a mixed audio signal. Spatial information such as the arrangement of microphone groups and features in the audio signal can be used to identify and extract sounds from different sources.

[0102] In some scenarios, microphone groups are arranged in a specific way, and the position of each sound source can be inferred by the sound intensity and delay received by the microphone group. If the microphone groups are arranged in different positions in an area, the position of each speaker can be inferred by the time difference between the sound reaching different microphones. Microphone array sound source localization technology determines the spatial position of the sound source by analyzing the differences in sound wave signals received by multiple microphones and combining geometric models and signal processing algorithms. Its core principle is to construct a mathematical model using the differences in the propagation paths of sound waves reaching different microphones. For example, the hyperbolic localization method infers the trajectory of the sound source through the time difference, while beamforming technology enhances the sound source in a specific direction through weighted signals.

[0103] In some embodiments, the hyperbolic positioning method uses two microphones as the focus, the delay difference corresponds to the intersection of the hyperbolas, and the three-microphone array can be positioned by the intersection of two hyperbolas. Finally, the optimal position is solved for the delay equations of multiple microphones by the least squares method to obtain the sound source position.

[0104] At the same time, sound source separation will also combine the previous emotional label information, because different emotional states may affect the characteristics of the sound. For example, angry tones and sounds may be higher or harsher, while calm sounds may be softer. Through emotional labels, each sound source can be separated and identified more accurately.

[0105] Step 230: input each audio sub-signal into an emotion recognition model to obtain the emotional color of the sound corresponding to each person object corresponding to each audio sub-signal.

[0106] After the sound source is separated, the voice of each speaker, that is, the voice of each person object, is extracted separately to form an audio sub-signal. These audio sub-signals will have different emotional labels, indicating the emotional state of each speaker.

[0107] In the process of sound source separation, the position of each sound source, i.e., human object, in the target environment can be determined through the microphone group and positioning technology. That is, the sound position of each sound source can be inferred through the time difference between the sound signal reaching different microphones, or the relative position relationship between the microphone groups.

[0108] It should be noted that the focus of emotion recognition is on non-linguistic features, not the language content itself. Therefore, even if the content in the audio is vague and not every word in the speech is understood, as long as these acoustic features can be obtained, emotion recognition can still make relatively accurate judgments. For example, even if you can't hear clearly what someone is saying, if their pitch rises sharply, their speech speed increases, and their volume increases, you can still judge that they may be angry or excited.

[0109] The task of speech recognition is to convert speech into text and understand what is being said, which requires accurately capturing the content of each word. The task of emotion recognition is to extract emotional clues from the sound and understand the speaker's emotions. This relies on acoustic features and does not require the content of each word. It is suitable for event recognition in scenarios where the acquired speech signal is ambiguous.

[0110] It is understandable that emotional color can be identified through sound wave features, such as pitch, volume, and speech speed. Especially in cases where emotions are more obvious, such as anger and happiness, these features can be effectively judged through emotion recognition. In scenes where there are quarrels, disputes, and fights, the features used to identify emotional color, such as pitch, volume, and speech speed, are very obvious, and the recognition accuracy is higher, which is conducive to better identification of target events.

[0111] In order to identify more accurate emotional colors and make reference for event recognition, an emotion recognition model can be used to obtain more accurate emotion recognition results. The emotion recognition model is a sentiment analysis tool based on audio signals, which extracts audio features and uses machine learning or deep learning methods to classify emotions.

[0112] The steps of training the emotion recognition model include data preparation, feature extraction, model selection, training and optimization, etc. For example, a convolutional neural network (CNN) can be used for training.

[0113] The emotion recognition model includes a feature extraction module and a classification module. The feature extraction module converts the audio signal into a numerical form that can represent the emotional characteristics of the speech. The classification module classifies the emotions based on the extracted features and predicts the emotions of the audio clips.

[0114] Emotion recognition needs to be trained based on labeled speech datasets. Common datasets include RAVDESS and TESS. You can use labeled emotion audio datasets to train emotion recognition models. The goal of the model is to classify emotions as accurately as possible by learning the mapping relationship between audio features and emotion labels.

[0115] In some embodiments, based on the initial emotion label, the sound source position of the second audio signal and the position of the microphone group arranged at the sound source position, the second audio signal is separated from the sound source to extract audio sub-signals corresponding to each person object emitting the sound, including: determining the position of each sound source in the second audio signal through a beamforming algorithm based on the sound source position of the second audio signal, the position of the microphone group arranged at the sound source position and the time difference between the sound reaching different microphones in the microphone group; extracting audio sub-signals corresponding to each person object at the position of the sound source from the second audio signal based on the position of each sound source in the second audio signal and the correlation of the emotion color levels of the initial emotion labels corresponding to different frequency bands in the second audio signal.

[0116] It is understandable that the second audio signal has undergone preliminary emotion analysis, and an initial emotion label is assigned to each audio signal. When separating the sound source, the spatial position of the sound source can be used for analysis. The sound source position here refers to the specific position of the speaker in the target environment, which is usually inferred by the arrangement of the microphone group and the difference in the arrival time of the sound signal received by the microphone.

[0117] Specifically, the position of each sound source can be inferred by calculating the time difference (ie, sound wave propagation time) from different sound sources to different microphones.

[0118] The beamforming algorithm is used to determine the location of each sound source in the second audio signal. By using the time difference between the sound reaching different microphones, the beamforming algorithm can help locate the direction of the sound source, thereby determining the location of each speaker. Beamforming is a signal processing technology used to focus the signal received from multiple microphones to a specific direction.

[0119] It should be noted that the levels of emotional colors and the relationship between adjacent levels need to be defined. Emotional colors are divided into multiple levels, and there is a certain correlation between the levels of each emotional color. For example, some levels between happy and neutral are adjacent, while other emotions have different correlations.

[0120] For example, happiness can be divided into three levels: happiness level 1 (slight happiness), happiness level 2 (moderate happiness), and happiness level 3 (extreme happiness).

[0121] Neutrality can be divided into three levels: Neutral 1 (mild neutral), Neutral 2 (moderate neutral) and Neutral 3 (extremely neutral).

[0122] Anger can be divided into three levels: anger 1 (mild anger), anger 2 (moderate anger) and anger 3 (extreme anger).

[0123] The lowest level of happiness (Happy 1) is adjacent to the highest level of neutrality (Neutral 3), so there is a high correlation.

[0124] There will be a smaller correlation between the levels of anger and happiness, for example, level 3 anger will be more different from level 3 happiness.

[0125] In the process of sound source separation, the initial emotion label can be used as additional guiding information to help the separation algorithm optimize the sound source according to the emotion category and level. In this embodiment, it is necessary to combine the information of emotional color with the separation of audio signals so that the emotional layering can be better reproduced in the separated audio signals.

[0126] In some embodiments, based on the positions of each sound source in the second audio signal and the correlation between the emotional color levels of the initial emotional labels corresponding to different frequency bands in the second audio signal, audio sub-signals corresponding to each person object at the position of the sound source are extracted from the second audio signal, including the following process.

[0127] The audio signal at each sound source position can be determined based on the position of each sound source in the second audio signal. The audio source signal after decomposition of the audio signal at each sound source position is evaluated based on the correlation of the emotional color levels of the initial emotional labels corresponding to different frequency bands in the second audio signal to obtain the audio sub-signal corresponding to each person object.

[0128] Use blind source separation techniques such as non-negative matrix factorization (NMF) or independent component analysis (ICA) to separate the mixed signal in the entire audio signal into different sound source signals. At this time, it is necessary to combine the location of the sound source and the initial emotional label to optimize the separation effect. For example, if the initial emotional label corresponding to a certain frequency band is happy, then the sound source audio signal with the initial emotional label of happy can be extracted from the entire audio and is more likely to be determined as the same sound source. However, the frequency bands with the initial emotional labels of happy and sad are unlikely to be the same sound source.

[0129] According to the correlation between emotional labels and frequency bands, emotional label constraints can be introduced into the sound source separation algorithm. For example, in the process of non-negative matrix factorization (NMF) or independent component analysis (ICA), the spectrum separation process can be adjusted according to the emotional label so that signal components with different emotional colors can be accurately separated.

[0130] In the NMF or ICA model, different spectral bases, i.e., different basis matrices, can be established for each emotion label, and then the audio signal can be weighted according to the initial emotion label so as to extract the audio component corresponding to the emotion from the audio signal. Specifically, the level corresponding to the initial emotion label can be assigned a distance value, and then the similarity of the distance values ​​can be calculated for audio signals in different frequency bands, and the similarity between the audio signals corresponding to the two initial emotion labels can be determined according to the size of the similarity.

[0131] It can be understood that the distance values ​​between different levels of the same emotion are very small, and the distance values ​​between similar emotions are very small. For example, the distance value between Happy Level 1 (slightly happy) and Happy Level 2 (moderately happy) can be 1. The distance value between Neutral Level 1 (mildly neutral) and Happy Level 1 (slightly happy) can be 2, and the distance value between Neutral Level 3 (extremely neutral) and Angry Level 1 (slightly angry) can also be 2. In this way, there is a neutral level between happiness and anger, the distance is far, and the similarity is low.

[0132] In the separation process, we can also introduce prior knowledge of emotional color as a regularization term to ensure that the emotional color of each source signal is consistent with the label. For example, happy emotional signals may need to have stronger weights in the high frequency band, while sad emotional signals should be more concentrated in the low frequency band.

[0133] It can be understood that by introducing the differentiated constraints of the emotional color level, the extracted audio signal for each audio source is made more consistent with the characteristics of the emotion, and thus more consistent with the audio source corresponding to the independent person object.

[0134] In some embodiments, determining the occurrence of a target event at a target location of a target environment based on the location of a person object making a sound in the second audio signal and the emotional color of the sound corresponding to each person object includes the following process.

[0135] Based on the level of the emotional color of the voice corresponding to each person object, it can be determined that the level evaluation results of the emotional color of two adjacent target person objects meet the preset conditions.

[0136] According to the emotional color evaluation results of each person in the audio signal, it is determined whether the emotional color levels of two target person objects meet certain preset conditions. Generally, the preset conditions may include: the emotional colors of the two people are negative emotions (such as anger or sadness).

[0137] If two people's emotional colors are angry and intense, or sad and intense, this indicates that the two people are emotionally intense and may be in conflict.

[0138] Based on the positions of the two target personnel objects, determining that the distance between the two target personnel objects is less than the target distance;

[0139] In addition to the matching of emotional colors, the distance between the target person objects is also an important factor.

[0140] Based on the spatial position of the target person, the distance between the two people is calculated. If the distance between the two target people is less than a certain target distance threshold, such as a critical value of 1 meter or 2 meters, it means that the two people are very close and may be interacting or arguing fiercely. If the two people are far apart, it means that they may be on the phone separately, etc., and there is no need to pay attention.

[0141] In this case, it can be determined that the target event occurs at the target location of the target environment. When the distance and emotional color conditions are met, it can be determined that the target event occurs at the target location of the target environment.

[0142] The speech signal extraction and recognition device provided by the present invention is described below. The speech signal extraction and recognition device described below and the speech signal extraction and recognition method described above can be referred to each other.

[0143] like Figure 3 As shown, the speech signal extraction and recognition device according to the embodiment of the present invention mainly includes an acquisition module 310 , an extraction module 320 , a first recognition module 330 , a second recognition module 340 , a third recognition module 350 and a processing module 360 ​​.

[0144] The acquisition module 310 is used to monitor the sound signal in the target environment in real time, acquire the first audio signal based on the sliding window method, and compare it with the preset audio corresponding to the target environment;

[0145] The extraction module 320 is used to extract the second audio signal including the audio signal with the deviation from the first audio signal when there is a target deviation between the first audio signal and the preset audio;

[0146] The first recognition module 330 is used to input the second audio signal into the recognition model to identify the type of the second audio signal;

[0147] The second identification module 340 is used to identify the sound source position of the second audio signal and determine whether the sound source position of the second audio signal matches the target environment;

[0148] The third recognition module 350 is used to identify the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object when it is determined that the second audio signal contains human voice based on the type of the second audio signal;

[0149] The processing module 360 ​​is used to determine that a target event occurs at a target position of a target environment based on the position of a person object that makes a sound in the second audio signal and the emotional color of the sound corresponding to each person object.

[0150] The speech signal extraction and recognition device provided in the embodiment of the present invention analyzes the audio signal layer by layer, from deviation detection to complex emotion analysis and positioning, to ensure that even when the audio signal is complex or ambiguous, it can still gradually focus on and extract key features to reduce the risk of misjudgment. By comprehensively analyzing the various features of the audio signal, the nature and location of the event can be accurately inferred. The audio signal is used for positioning and emotion color analysis, which increases the efficiency of identifying the nature of the event. Clues can be found from complex and ambiguous audio signals to determine the type of event, thereby providing a reference for security management.

[0151] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device may include: a processor (processor) 410, a communication interface (Communications Interface) 420, a memory (memory) 430 and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call the logic instructions in the memory 430 to execute the speech signal extraction and recognition method, which includes: real-time monitoring of the sound signal in the target environment, obtaining the first audio signal based on the sliding window method, and comparing it with the preset audio corresponding to the target environment; when there is a target deviation between the first audio signal and the preset audio, extracting the second audio signal containing the deviated audio signal from the first audio signal; inputting the second audio signal into the recognition model to identify the type of the second audio signal; identifying the sound source position of the second audio signal, and determining that the sound source position of the second audio signal matches the target environment; when it is determined that the second audio signal contains human voice based on the type of the second audio signal, identifying the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object; based on the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object, determining that a target event occurs at the target position of the target environment.

[0152] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.

[0153] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech signal extraction and recognition method provided by the above-mentioned methods, which includes: real-time monitoring of sound signals in a target environment, obtaining a first audio signal based on a sliding window method, and comparing it with a preset audio corresponding to the target environment; when there is a target deviation between the first audio signal and the preset audio, extracting a second audio signal containing the deviated audio signal from the first audio signal; inputting the second audio signal into a recognition model to identify the type of the second audio signal; identifying the sound source position of the second audio signal, and determining that the sound source position of the second audio signal matches the target environment; when it is determined that the second audio signal contains human voice based on the type of the second audio signal, identifying the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object; based on the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object, determining that a target event occurs at a target position of the target environment.

[0154] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the speech signal extraction and recognition method provided by the above-mentioned methods, the method comprising: real-time monitoring of sound signals in a target environment, obtaining a first audio signal based on a sliding window method, and comparing it with a preset audio corresponding to the target environment; when there is a target deviation between the first audio signal and the preset audio, extracting a second audio signal containing the deviated audio signal from the first audio signal; inputting the second audio signal into a recognition model to identify the type of the second audio signal; identifying the sound source position of the second audio signal, and determining that the sound source position of the second audio signal matches the target environment; when it is determined that the second audio signal contains human voice based on the type of the second audio signal, identifying the position of the person object emitting the sound in the second audio signal and the emotional color of the sound corresponding to each person object; based on the position of the person object emitting the sound in the second audio signal and the emotional color of the sound corresponding to each person object, determining that a target event occurs at a target position of the target environment.

[0155] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0156] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A speech signal extraction and recognition method, characterized in that: include: Performing real-time monitoring of sound signals in a target environment, acquiring a first audio signal based on a sliding window method, and comparing the first audio signal with a preset audio signal corresponding to the target environment; In the case where there is a target deviation between the first audio signal and the preset audio, extracting a second audio signal including the deviated audio signal from the first audio signal; the target deviation is determined based on at least one of frequency, intensity and timbre; Inputting the second audio signal into a recognition model to identify the type of the second audio signal; Identifying a sound source position of the second audio signal, and determining that the sound source position of the second audio signal matches the target environment; In a case where it is determined based on the type of the second audio signal that the second audio signal contains human voice, identifying the position of a person object making a sound in the second audio signal and the emotional color of the sound corresponding to each person object; Based on the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object, determining that a target event occurs at a target position of the target environment; the target event includes at least one of quarreling, arguing, and fighting; The step of determining the occurrence of a target event at a target position of the target environment based on the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object includes: Based on the level of the emotional color of the voice corresponding to each person object, determining that there are two adjacent target person objects whose level evaluation results of the emotional color meet the preset conditions; Based on the positions of the two target personnel objects, determining that the distance between the two target personnel objects is less than the target distance; Determine that a target event occurs at a target location of the target environment.

2. The method for extracting and recognizing speech signals according to claim 1, characterized in that: The identifying the position of the person object making the sound in the second audio signal and the emotional color of the sound corresponding to each person object includes: Performing initial emotion recognition on the second audio signal to obtain initial emotion labels of different frequency bands in the second audio signal; Based on the initial emotion label, the sound source position of the second audio signal and the position of the microphone group arranged at the sound source position, the second audio signal is separated into sound sources, audio sub-signals corresponding to the various human objects emitting sounds are extracted, and the sound positions of the various human objects corresponding to the various audio sub-signals are determined; Each audio sub-signal is input into the emotion recognition model to obtain the emotional color of the sound corresponding to each person object corresponding to each audio sub-signal.

3. The method for extracting and recognizing speech signals according to claim 2, characterized in that: The method of performing sound source separation on the second audio signal based on the initial emotion label, the sound source position of the second audio signal, and the position of the microphone group arranged at the sound source position, and extracting audio sub-signals corresponding to each person object that emits the sound, includes: Determine the positions of the respective sound sources in the second audio signal by a beamforming algorithm based on the sound source position of the second audio signal, the positions of the microphone group arranged at the sound source position, and the time difference between the sound reaching different microphones of the microphone group; Based on the positions of each sound source in the second audio signal and the correlation of the emotional color levels of the initial emotional labels corresponding to different frequency bands in the second audio signal, audio sub-signals corresponding to each person object at the position of the sound source are extracted from the second audio signal.

4. The method for extracting and recognizing speech signals according to claim 3, characterized in that: The extracting, from the second audio signal, audio sub-signals corresponding to the respective person objects at the positions of the sound sources based on the positions of the respective sound sources in the second audio signal and the correlation of the emotional color levels of the initial emotional labels corresponding to different frequency bands in the second audio signal comprises: Based on the position of each sound source in the second audio signal, determine the audio signal at the position of each sound source; Based on the correlation of the emotional color levels of the initial emotional labels corresponding to different frequency bands in the second audio signal, the audio source signal after the audio signal at each sound source position is decomposed is evaluated to obtain the audio sub-signal corresponding to each person object.

5. A speech signal extraction and recognition device, characterized in that: The device uses the speech signal extraction and recognition method according to any one of claims 1 to 4 to extract the audio signal to identify the target event, and the device includes: An acquisition module, used to monitor the sound signal in the target environment in real time, acquire the first audio signal based on a sliding window method, and compare it with the preset audio corresponding to the target environment; an extraction module, configured to extract a second audio signal including an audio signal with a deviation from the first audio signal when there is a target deviation between the first audio signal and the preset audio; A first recognition module, configured to input the second audio signal into a recognition model to identify a type of the second audio signal; A second recognition module, used to identify the sound source position of the second audio signal, and determine whether the sound source position of the second audio signal matches the target environment; A third recognition module is used to identify the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object in the second audio signal when it is determined that the second audio signal contains human voice based on the type of the second audio signal; The processing module is used to determine the occurrence of a target event at a target position of the target environment based on the position of the person object that makes the sound in the second audio signal and the emotional color of the sound corresponding to each person object.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the program, the speech signal extraction and recognition method according to any one of claims 1 to 4 is implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech signal extraction and recognition method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Security monitoring method, device, robot and storage medium

    CN111601074A

  • Event occurrence probability determination method, storage medium and electronic device

    CN113903003A