Emergency call method and system based on voice recognition triggering, storage medium and processor

By constructing a corpus related to user language type and using the LSTM-HMM acoustic recognition model for training, the problem of difficulty in identifying distress signals in the elderly's language type usage habits is solved, and accurate recognition and rapid rescue of emergency distress signals are achieved.

CN120071906APending Publication Date: 2025-05-30GUANGXI LVFA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510176906.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The language type usage habits of middle-aged and elderly people in the prior art are difficult to accurately identify distress signals, resulting in a decrease in recognition accuracy during emergency response and delaying the best time for help.

Method used

By constructing a corpus related to user language types and using an acoustic recognition model based on LSTM-HMM for training, the characteristics and rules of call signal in different languages ​​can be learned, so as to achieve accurate identification of user emergency call signal.

Benefits of technology

It improves the accuracy of identification of emergency call signals, expands the scope of application of emergency call systems, and can quickly start the emergency rescue process without manual intervention, greatly shortens the rescue response time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071906A_ABST
    Figure CN120071906A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice recognition, and discloses an emergency call method and system based on voice recognition triggering, a storage medium and a processor, and the method comprises the steps: obtaining voice signals of a user and an accompanying chat robot; preprocessing the voice signal; extracting characteristic parameters of the processed voice signals; inputting the extracted characteristic parameters into an acoustic recognition model trained based on LSTM-HMM, and recognizing an emergency call signal of the user; and starting an emergency rescue process according to the identified emergency call signal of the user. According to the invention, based on the language habit used by the user, the corpus in which the language type of the user is related to the distress signal is constructed for training, so that the method has the capability of preparing to identify the language distress signal of the language type used by the user, and the identification accuracy of the emergency distress signal is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and particularly relates to an emergency call method, system, storage medium and processor triggered by speech recognition. Background Art

[0002] With the development of technology in the era, there have emerged many products for the companionship and protection of the elderly. For example, a chatbot is a product specifically for the companionship and protection of the elderly, and through the chatbot, conversations with the elderly can be realized. When the elderly are in danger and need first aid, the chatbot can recognize the distress signal of the elderly.

[0003] Although the chatbot can recognize the distress signal of the elderly, since the elderly may be accustomed to using a specific dialect, this may lead to a decrease in the accuracy of the distress voice signal recognition system and delay the best time for calling for help. Summary of the Invention

[0004] Aiming at the problem that it is difficult to accurately recognize the distress signal according to the language type usage habit of the elderly in the prior art, the present invention provides an emergency call method, system, storage medium and processor triggered by speech recognition, which can learn according to the language type usage habit of the user, accurately recognize the voice distress signal of the user, and realize emergency call. The specific technical solutions are as follows:

[0005] An emergency call method triggered by speech recognition includes:

[0006] Obtain the voice signals of the user and the chatbot;

[0007] Preprocess the voice signals;

[0008] Extract the characteristic parameters of the preprocessed voice signals;

[0009] Input the extracted characteristic parameters into an acoustic recognition model trained based on LSTM-HMM to recognize the emergency call signal of the user;

[0010] According to the recognized emergency call signal of the user, start the emergency rescue process;

[0011] The training process of the acoustic recognition model trained based on LSTM-HMM includes:

[0012] Construct a distress signal corpus related to the language type used by the user and the distress signal;

[0013] After preprocessing the distress signal corpus, extract the characteristic parameters;

[0014] The extracted feature parameters are input into the LSTM-HMM model for learning and training to obtain an acoustic recognition model capable of recognizing the language type used by the user for calling for help.

[0015] Preferably, the preprocessing of the speech signal includes:

[0016] Performing pre-emphasis processing on the speech signal through a preset pre-emphasis coefficient;

[0017] Performing frame segmentation on the pre-emphasized speech signal to divide it into speech frames of a preset time;

[0018] Using a window to smooth the speech frames;

[0019] Performing endpoint detection on the windowed speech frames through a double-threshold method.

[0020] Preferably, the feature parameters include Mel-frequency cepstral coefficients and spectrogram features.

[0021] Preferably, the extraction of the feature parameters of the processed speech signal includes:

[0022] Performing a fast Fourier transform on the processed speech signal to obtain the spectrum of each frame and calculating the power spectrum of each frame;

[0023] Mapping the spectrum to the Mel-frequency scale and performing weighted averaging on the power spectrum through a preset filter bank to obtain an output value;

[0024] Taking the logarithm of each value output by the filter bank and performing a discrete cosine transform to obtain Mel-frequency cepstral coefficients.

[0025] Preferably, the extraction of the feature parameters of the processed speech signal further includes:

[0026] Taking the logarithm of the power spectrum to obtain a logarithmic power spectrum;

[0027] Arranging the logarithmic power spectra in chronological order to form a spectrogram.

[0028] Preferably, the construction of the corpus of distress signal language types used by the user includes:

[0029] Obtaining speech distress data associated with distress information in the language type used by the user;

[0030] Using a preset common language to annotate the corpus text corresponding to the speech distress data;

[0031] Updating the pronunciation, initials, finals, and tones corresponding to the corpus text into a dictionary to obtain a corpus of distress signals.

[0032] An emergency call system triggered by voice, which is applied to the aforementioned emergency call method triggered by voice recognition, includes:

[0033] A data acquisition unit for acquiring voice signals between a user and a chatbot;

[0034] A data processing unit for preprocessing the voice signals;

[0035] A feature extraction unit for extracting feature parameters of the processed voice signals;

[0036] A distress signal recognition unit for inputting the extracted feature parameters into an acoustic recognition model trained based on LSTM-HMM to recognize the user's emergency distress signal;

[0037] A rescue activation unit for activating an emergency rescue process according to the recognized user's emergency distress signal;

[0038] The distress signal recognition unit includes:

[0039] A corpus construction module for constructing a distress signal corpus related to the language type used by the user and the distress signal;

[0040] A data processing module for preprocessing the distress signal corpus and then extracting feature parameters;

[0041] A model learning and training module for inputting the extracted feature parameters into an LSTM-HMM model for learning and training to obtain an acoustic recognition model capable of recognizing the user's distress in the language type used.

[0042] A computer-readable storage medium, which includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the aforementioned emergency call method triggered by voice recognition.

[0043] A processor for running a program. When the program runs, it executes the aforementioned emergency call method triggered by voice recognition.

[0044] Compared with the prior art, the beneficial effects of the present invention are:

[0045] The emergency call method triggered by voice recognition according to the present invention is trained by constructing a corpus related to the call signal for the user language type based on the user's language habits. The LSTM-HMM acoustic recognition model can learn the characteristics and rules of the call signal in different languages, so as to have the ability to accurately recognize the language call signal of the language type used by the user. This not only improves the recognition accuracy of the emergency call signal, but also expands the applicable range of the emergency call system. At the same time, once the emergency call signal of the user is recognized, the system can quickly start the emergency rescue process without manual intervention, greatly shortening the rescue response time. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0047] Figure 1 It is a flowchart of the emergency call method triggered by voice recognition according to the present invention.

[0048] Figure 2 It is a schematic diagram of the emergency call system triggered by voice recognition according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0051] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0052] It should also be further understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0053] For the following embodiments, please refer to Figure 1 and Figure 2 .

[0054] An embodiment of the present application provides an emergency call method triggered by voice recognition, including:

[0055] Step S1, obtaining a voice signal of a user and a chatbot;

[0056] By equipping the chatbot with a microphone, the voice signal of the user can be collected through the microphone. At the same time, a stable audio input interface and audio acquisition software are set in the chatbot, which can convert the analog audio signal collected by the microphone into a digital audio signal for storage and processing.

[0057] When the user interacts with the chatbot, the audio acquisition software of the chatbot starts to collect the voice signal of the user in real time. During the collection process, the software will perform digital conversion on the voice according to a certain preset frequency (16 kHz is adopted) and quantization bits (16 bits are adopted), and convert the continuous voice waveform into a discrete digital sequence.

[0058] The collected voice signal will be temporarily stored in the local storage device of the chatbot, or transmitted to the backend server for storage in real time through the network. The stored voice signal will be marked with information such as the collection time and user identification.

[0059] Step S2, preprocessing the voice signal;

[0060] The preprocessing of the voice signal includes:

[0061] Performing pre-emphasis processing on the voice signal through a preset pre-emphasis coefficient;

[0062] Performing frame segmentation on the pre-emphasized voice signal, and segmenting it into voice frames of a preset time;

[0063] Performing smoothing processing on the voice frames by using a window;

[0064] Performing endpoint detection on the windowed voice frames through a double-threshold method.

[0065] Preferably, the feature parameters include Mel frequency cepstral coefficients and spectrogram features.

[0066] Through preprocessing, the quality of the speech signal and the performance of the speech recognition system are improved. Pre-emphasis enhances the high-frequency information, framing and windowing optimize the short-term stationarity of the signal and the effect of spectral analysis, while endpoint detection removes invalid data and focuses on processing valid speech. This provides high-quality input data for subsequent feature extraction and model training.

[0067] Step S3: Extract the feature parameters of the processed speech signal;

[0068] The feature parameters include Mel Frequency Cepstral Coefficients (MFCC) and spectrogram features;

[0069] The preprocessed speech signal is framed and windowed, and then the Fast Fourier Transform (FFT) is performed on each frame of the speech signal to obtain the spectrum of the speech signal. Then, the spectrum is mapped to the Mel frequency scale through a Mel filter bank, and the output energy of each filter is calculated. Finally, the discrete cosine transform (DCT) is performed on the energy spectrum to obtain the MFCC feature parameters. The MFCC feature parameters can effectively represent the spectral characteristics of the speech signal and capture key information in the speech, such as timbre, pitch, etc. By performing the Fourier transform on the preprocessed speech signal, the spectrum of the speech signal within different time windows is calculated, and then the spectra are arranged in chronological order to form a spectrogram. The spectrogram can visually display the distribution of the speech signal in time and frequency and reflects the dynamic change characteristics of the speech signal. Multiple feature parameters can be extracted from the spectrogram, and these feature parameters can enrich the feature representation of the speech signal and improve the accuracy of speech recognition.

[0070] Step S4: Input the extracted feature parameters into the acoustic recognition model trained based on LSTM-HMM to recognize the user's emergency call signal;

[0071] The extracted Mel Frequency Cepstral Coefficients (MFCC) and spectrogram features are organized into an input vector in a certain format and then input into the acoustic recognition model trained based on LSTM-HMM. During the model loading process, it is necessary to ensure that the parameters and structure of the model are consistent with those during training so that accurate speech recognition can be performed. For example, if 13-dimensional MFCC feature parameters are used during training, then the same-dimensional feature parameters need to be used during recognition and input in the same order and format.

[0072] LSTM-HMM Model Recognition Process: The LSTM-HMM model combines the advantages of the LSTM network and the Hidden Markov Model (HMM), and can effectively process the temporal characteristics and statistical characteristics of speech signals. During the recognition process, the LSTM network first processes the input sequence of feature parameters, extracts the long-term dependencies and context information in the speech signal, and then transmits the processed information to the HMM. The HMM models and recognizes the speech signal based on the pre-trained state transition probabilities and observation probabilities, and outputs the corresponding speech text or semantic information. For example, when the input sequence of feature parameters corresponds to an emergency call signal, the LSTM-HMM model can accurately recognize the signal and output the corresponding call semantics.

[0073] Step S5: According to the recognized emergency call signal of the user, initiate the emergency rescue process;

[0074] Once the emergency call signal of the user is recognized, the system immediately triggers the emergency rescue process. The triggering method can be automatic or initiated after a certain confirmation mechanism.

[0075] The process of initiating the automatic emergency rescue process is as follows:

[0076] Send a distress message to the preset emergency contacts and rescue agencies;

[0077] Among them, the rescue information includes basic information such as the user's name, age, gender, current location, and emergency contacts, as well as the detailed content of the emergency call signal, the call time, the call location, etc. The rescue information can be pushed through various communication methods, such as text messages, phone calls, emails, instant messaging tools, etc., to ensure that rescue personnel and agencies can receive the information in a timely manner and take corresponding rescue measures.

[0078] The process of initiating the emergency rescue process with a confirmation mechanism is as follows:

[0079] Send a distress signal to the preset emergency contacts and rescue agencies:

[0080] The emergency contacts are relatives. Connect to the training robot through devices such as the mobile terminal of the emergency contact, turn on the camera of the chat robot to view the user's status (the elderly), and confirm whether to send a rescue request to the preset rescue center. Initiate the rescue process after obtaining the user's confirmation. Or the system can automatically send a rescue request to the preset rescue center after recognizing the emergency call signal. Initiate the rescue process after obtaining the user's confirmation.

[0081] The training process of the trained acoustic recognition model based on LSTM-HMM includes:

[0082] Construct a corpus of distress call signals related to the language type used by the user, specifically including:

[0083] Obtain the voice call-for-help data associated with the call-for-help information in the language type used by the user;

[0084] The language types used by the user may include dialects, accented Mandarin, etc. By collecting the user's voice call-for-help data associated with the call-for-help information, at the same time, screening a large amount of collected voice data and only retaining the voices directly related to the call-for-help information. For example, removing some ordinary conversation voices that are in this language but do not contain the intention of calling for help.

[0085] Annotate the corpus text corresponding to the voice call-for-help data using a preset common language;

[0086] Perform text semantic annotation on the user's language type using the vernacular of Mandarin to form the corresponding relationship between the voice call-for-help and the semantic text.

[0087] Update the pronunciation, initials, finals, and tones corresponding to the corpus text to the dictionary to obtain the call-for-help signal corpus.

[0088] Add the pronunciation information of the annotated corpus text to the dictionary. If the same word already exists in the dictionary but has a different pronunciation, it needs to be updated or supplemented according to the annotation result.

[0089] Update of initials, finals, and tones: Organize the initials, finals, and tone information in the corpus text respectively and update them to the corresponding parts of the dictionary. Integrate the updated dictionary with the original voice call-for-help data and the corresponding annotated corpus text to form a complete call-for-help signal corpus. This corpus not only contains rich call-for-help voice data but also has detailed pronunciation annotation information, which is convenient for subsequent model training and recognition.

[0090] Preprocess the call-for-help signal corpus and extract feature parameters;

[0091] Input the extracted feature parameters into the LSTM-HMM model for learning and training to obtain an acoustic recognition model capable of recognizing the language type used by the user for calling for help.

[0092] By collecting multi-channel and high-quality voice call-for-help data, it provides rich and real samples for model training. These data can reflect the call-for-help characteristics of different scenarios and different populations, enabling the trained model to have stronger generalization ability and be able to more accurately recognize various types of emergency call-for-help signals. Using a preset common language for annotation unifies the representation method of pronunciation, avoids the differences brought by different dialects or pronunciation habits, enables the model to more accurately capture the essential features of the voice during the learning process, and improves the model's ability to understand and recognize the voice.

[0093] In this embodiment, taking the Zhuang language of the Zhuang ethnic group as an example, the training process of the trained acoustic recognition model based on LSTM-HMM is described, including:

[0094] In this embodiment, taking the Zhuang language of the Zhuang ethnic group as an example, the training process of the trained acoustic recognition model based on LSTM-HMM is specifically described as follows:

[0095] The first step is to construct a corpus of Zhuang language distress signals.

[0096] Collect Zhuang language distress signal corpora through various channels, including on-site recording of the Zhuang language distress voice samples of chatbot users, collecting distress voice samples in Zhuang ethnic group-inhabited areas, and screening out voice data related to distress from existing Zhuang language voice databases. The collected Zhuang language distress corpora are detailedly annotated in vernacular Chinese, including the text content of the voice, speaker information, etc. The annotation work can be jointly completed by professional Zhuang language linguists and speech recognition experts.

[0097] The second step is to preprocess the Zhuang language distress signal corpus and extract feature parameters.

[0098] The process of preprocessing the Zhuang language distress voice signal is the same as that in step S2. After preprocessing the voice signal, the annotated Zhuang language distress corpus text is normalized to obtain the voice-text semantic correspondence of the Zhuang language distress signal. For example, the following table shows the voice-text semantic correspondence of the Zhuang language distress signal.

[0099] Serial number Zhuang language Semantic annotation 1 Goucaiq(goucaiq) Help 2 Goucaiqdaemz(goucaiqdaemz) Help! 3 Save me Goucaiqbiq(goucaiqbiq) 4 Goucaiqbiqdaemz(goucaiqbiqdaemz) Save me!

[0100] The extraction of feature parameters is the same as that in step S for extracting Mel Frequency Cepstral Coefficients (MFCC). MFCC feature parameters can effectively represent the spectral characteristics of the Zhuang language voice signal and capture key information in the Zhuang language voice, such as timbre, pitch, etc. In addition to MFCC features, other feature parameters of the Zhuang language distress signal are also extracted, including fundamental frequency, formant, and duration; the fundamental frequency and formant are closely related to the pitch and timbre characteristics of the Zhuang language; the duration feature can represent the duration of the Zhuang language distress signal. The extraction of these feature parameters can provide more comprehensive voice information for the model and help improve the recognition accuracy of the Zhuang language distress signal.

[0101] The third step is to input the extracted feature parameters into the LSTM-HMM model for learning and training.

[0102] Initialize the acoustic recognition model based on LSTM-HMM, and set the parameters and structure of the model. The LSTM part includes an input layer, multiple LSTM hidden layers, and an output layer, which are used to process the temporal features of Zhuang language distress signals; the HMM part is used to model the statistical characteristics of Zhuang language distress signals, including state transition probabilities and observation probabilities, etc. Reasonably set the parameters of the model according to the characteristics of Zhuang language distress signals and the scale of the corpus.

[0103] Input the extracted feature parameters of Zhuang language distress signals into the LSTM-HMM model for learning and training. Adopt the method of supervised learning, use the labeled Zhuang language distress corpus as training data, and through optimization methods such as the backpropagation algorithm and the expectation maximization (EM) algorithm, continuously adjust the parameters of the model so that the model can accurately learn the characteristics and rules of Zhuang language distress signals.

[0104] Through the above steps, an acoustic recognition model trained based on LSTM-HMM with the ability to recognize Zhuang language distress signals can be obtained. This model can accurately recognize the emergency distress signals issued by Zhuang language users, provide effective technical support for emergency rescue work in Zhuang areas, and has important application value and social significance.

[0105] Construct a corpus related to the user's language type and distress signals, which can accommodate the emergency distress signals issued by users, so that the model can be exposed to possible distress expression methods during the training process. Focus on the collection of distress signals to ensure that the data in the corpus is highly relevant to emergency distress scenarios, which helps the model to more accurately recognize distress signals in emergency situations and reduce the possibility of misrecognition and missed recognition.

[0106] The emergency distress method based on voice recognition trigger of the present invention is trained by constructing a corpus related to the user's language type and distress signals based on the user's language habits. The LSTM-HMM acoustic recognition model can learn the characteristics and rules of distress signals in different languages, so as to have the ability to accurately recognize the language distress signals of the language type used by the user, which not only improves the recognition accuracy of emergency distress signals, but also expands the applicable range of the emergency distress system. At the same time, once the emergency distress signal of the user is recognized, the system can quickly initiate the emergency rescue process without manual intervention, greatly shortening the rescue response time.

[0107] Specifically, in a preferred embodiment of the present application, the feature parameters of the processed speech signal include:

[0108] Perform a fast Fourier transform on the processed speech signal to obtain the spectrum of each frame, and calculate the power spectrum of each frame;

[0109] P[k] = |X[k]|^2

[0110] Among them, P[k] represents the power spectral value at the discrete frequency index k; X[k] represents the complex spectral value at the k-th frequency point.

[0111] Map the spectrum to the Mel frequency scale, and perform weighted averaging on the power spectrum through a preset filter bank to obtain an output value;

[0112] The spectrum can be weighted and averaged through a set of triangular filters. Among them, the relationship between the Mel frequency and the ordinary frequency f is as follows:

[0113] m = 2595 * log10(1 + f / 700)

[0114] In the formula, m represents the Mel frequency;

[0115] Each filter in a set of triangular filters is a triangle, and its shape is defined by three points: the left boundary, the center frequency, and the right boundary. The response of the filter is maximum at the center frequency and gradually decreases towards both sides.

[0116] Take the logarithm of each value output by the filter bank and perform discrete cosine transform to obtain the Mel frequency cepstral coefficients. The specific process is as follows:

[0117] Take the logarithm of each value output by the filter bank to simulate the human ear's perception of sound intensity. The formula is as follows:

[0118] L[m] = log(F[m])

[0119] In the formula, F[m] represents the output value of the filter bank; L[m] represents the value obtained after taking the logarithm of the filter bank output value F[m].

[0120] Perform discrete cosine transform (DCT) on the logarithmic Mel spectrum to obtain the Mel frequency cepstral coefficients. The formula is as follows:

[0121] C[n] = ∑ L[m] * cos(πn(m + 0.5) / N)

[0122] In the formula, C[n] represents the Mel frequency cepstral coefficients; n is the index after discrete cosine transform; N is the number of filter banks.

[0123] In this embodiment, through FFT, power spectrum calculation, Mel frequency mapping, filter bank weighted averaging, logarithm taking, and DCT, high-quality feature parameters (MFCC) can be extracted from the speech signal. The obtained MFCC feature parameters can effectively improve the performance of the speech recognition system, especially in applications in multilingual environments, dialect recognition, and noisy environments.

[0124] Specifically, in a preferred embodiment of the present application, the characteristic parameters of the processed voice signal further include:

[0125] Taking the logarithm of the power spectrum to obtain the logarithmic power spectrum;

[0126] Arranging the logarithmic power spectra in chronological order to form a spectrogram.

[0127] In this embodiment, the preprocessed voice signal is subjected to a fast Fourier transform again to obtain the spectrum of each frame. Using time as the horizontal axis (the serial number of the frame) and frequency as the vertical axis (the frequency points after FFT), the spectral amplitude values of each frame are arranged to form a spectrogram. Each point on the spectrogram represents the amplitude of the voice signal at a specific time and frequency, intuitively showing the time-frequency characteristics of the voice signal.

[0128] The embodiment of the present application further provides an emergency call system triggered by voice, which is applied to the aforementioned emergency call method triggered by voice recognition, and includes:

[0129] A data acquisition unit for acquiring the voice signal between the user and the chatbot;

[0130] A data processing unit for preprocessing the voice signal;

[0131] A feature extraction unit for extracting the characteristic parameters of the processed voice signal;

[0132] A distress signal recognition unit for inputting the extracted characteristic parameters into an acoustic recognition model trained based on LSTM-HMM to recognize the user's emergency distress signal;

[0133] A rescue activation unit for activating an emergency rescue process according to the recognized user's emergency distress signal;

[0134] The distress signal recognition unit includes:

[0135] A corpus construction module for constructing a distress signal corpus related to the language type used by the user and the distress signal;

[0136] A data processing module for preprocessing the distress signal corpus and then extracting the characteristic parameters;

[0137] A model learning and training module for inputting the extracted characteristic parameters into an LSTM-HMM model for learning and training to obtain an acoustic recognition model capable of recognizing the user's language type for calling for help.

[0138] The function explanations of each unit and each module in this embodiment are the same as those of the emergency call method triggered by voice, and the technical effects are the same, so they will not be repeated here.

[0139] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the foregoing emergency call method triggered by voice recognition.

[0140] The technical effects of this embodiment are the same as those of the emergency call method triggered by voice in the embodiment, and will not be repeated here.

[0141] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.

[0142] An embodiment of the present application also provides a processor. The processor is used to run a program. When the program runs, it executes the foregoing emergency call method triggered by voice recognition.

[0143] Compared with the prior art, the beneficial effects of the present invention are:

[0144] The technical effects of this embodiment are the same as those of the emergency call method triggered by voice in the embodiment, and will not be repeated here.

[0145] The processor in this embodiment may be a central processing unit (CPU), a controller, a microcontroller, or other data processing chips.

[0146] Those of ordinary skill in the art can realize that the units of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components of each example have been generally described according to their functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0147] In the embodiments provided by the present invention, it should be understood that the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units can be combined into one unit, one unit can be split into multiple units, or some features can be ignored, etc.

[0148] In addition, in each embodiment of the present invention, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0149] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0150] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of each embodiment of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.

Claims

1. An emergency call method based on voice recognition triggering, characterized in that: include: Obtain the voice signals between the user and the chat robot; Preprocessing the speech signal; Extracting characteristic parameters of the processed speech signal; The extracted feature parameters are input into the acoustic recognition model trained based on LSTM-HMM to identify the user's emergency distress signal; Initiate emergency rescue process based on identifying the user's emergency distress signal; The training process of the acoustic recognition model based on the LSTM-HMM training includes: Construct a distress signal corpus related to the language type used by users and distress signals; After preprocessing the distress signal corpus, feature parameters are extracted; The extracted feature parameters are input into the LSTM-HMM model for learning and training to obtain an acoustic recognition model capable of identifying the language type used by the user to call for help.

2. The method for emergency rescue based on voice recognition triggering according to claim 1, characterized in that: The preprocessing of the speech signal comprises: Pre-emphasis processing is performed on the speech signal by using a preset pre-emphasis coefficient; The pre-emphasized voice signal is framed and divided into voice frames of preset time. Use windowing to smooth the speech frames; The double threshold method is used to detect the endpoints of the windowed speech frames.

3. The method for emergency rescue based on voice recognition triggering according to claim 2, characterized in that: The characteristic parameters include Mel-frequency cepstral coefficients and spectrogram features.

4. The method for emergency rescue based on voice recognition triggering according to claim 3, characterized in that: The characteristic parameters of the extracted speech signal include: Perform fast Fourier transform on the processed speech signal to obtain the spectrum of each frame, and calculate the power spectrum of each frame; The spectrum is mapped to the Mel frequency scale, and the power spectrum is weighted averaged through a preset filter bank to obtain the output value; the logarithm of each value output by the filter bank is taken and a discrete cosine transform is performed to obtain the Mel frequency cepstrum coefficients.

5. The method for emergency rescue based on voice recognition triggering according to claim 3, characterized in that: The feature parameters of the extracted speech signal also include: Take the logarithm of the power spectrum to obtain the logarithmic power spectrum; Arrange the logarithmic power spectra in chronological order to form a spectrogram.

6. The method for emergency rescue based on voice recognition triggering according to claim 2, characterized in that: The construction of the distress signal corpus of the language type used by the user comprises: Acquire voice distress call data associated with the distress call information in the language type used by the user; Using a preset common language to annotate the corpus text corresponding to the voice call for help data; The pronunciation, initial consonant, final vowel and tone corresponding to the corpus text are updated into the dictionary to obtain a distress signal corpus.

7. An emergency call system based on voice triggering, characterized in that: The emergency call method based on voice recognition triggering applied to any one of claims 1 to 6 comprises: A data acquisition unit, used to acquire voice signals between the user and the chat robot; A data processing unit, used for preprocessing the speech signal; A feature extraction unit, used to extract feature parameters of the processed speech signal; A distress signal recognition unit is used to input the extracted feature parameters into the acoustic recognition model trained based on LSTM-HMM to recognize the user's emergency distress signal; A rescue initiation unit, used to initiate an emergency rescue process based on identifying an emergency distress signal from a user; The distress signal recognition unit comprises: A corpus construction module is used to construct a distress signal corpus of the language type used by the user and related to the distress signal; A data processing module is used to pre-process the distress signal corpus and extract feature parameters; The model learning and training module is used to input the extracted feature parameters into the LSTM-HMM model for learning and training, so as to obtain an acoustic recognition model capable of recognizing the language type used by the user to call for help.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the emergency call method based on voice recognition triggering as described in any one of claims 1 to 6.

9. A processor, characterized in that: The processor is used to run a program, wherein the program, when running, executes the emergency call method based on voice recognition triggering as described in any one of claims 1 to 6.