A 4G / 2G / 5G CAT.1 communication system based on offline intelligent voice interaction

By designing a 4G/2G/5G CAT.1 communication system containing multiple modules, the problem of low offline intelligent voice interaction in the prior art is solved, efficient voice interaction and user feature recognition are achieved, and user experience is improved.

CN118571226BActive Publication Date: 2025-06-13SHENZHEN SHENYUAN SOFTWARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410960676.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-17
Publication Date
2025-06-13
Estimated Expiration
2044-07-17

AI Technical Summary

Technical Problem

In the prior art, offline intelligent voice interaction is inefficient and cannot effectively improve the efficiency and user experience of voice interaction.

Method used

A 4G/2G/5G CAT.1 communication system based on offline intelligent voice interaction was designed, including voice input module, voiceprint recognition module, manual compensation module, voice analysis module, voice algorithm recognition module, voice storage module, digital baseband module, emergency optimization module, user feedback module, function execution module and 4G/2G/5G CAT.1 communication module. Through the coordinated work of these modules, efficient voice interaction and user feature recognition are achieved.

Benefits of technology

By collecting voice vibration signals and user voiceprints in real time, accurate identification of user identities and features can be achieved, identity fraud is prevented, the efficiency and accuracy of voice interactions are improved, and the user experience of key users is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118571226B_ABST
    Figure CN118571226B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of speech recognition technology, and particularly to a 4G / 2G / 5G CAT.1 communication system based on offline intelligent voice interaction, including a voice input module; a voiceprint recognition module; an artificial compensation module for compensating the user feature recognition result; a voice analysis module for screening the audio signal features, a voice algorithm recognition module for performing algorithm recognition; a voice storage module for storing and updating a preset audio signal feature sequence; a digital baseband module for processing the algorithm recognition instruction; an emergency optimization module for optimizing the execution judgment process of an emergency instruction; a user feedback module for judging the emergency optimization effect; a function execution module for executing the device functions; a 4G / 2G / 5G CAT.1 communication module for realizing remote calls; and a contact setting module for inputting, modifying, and querying contacts. The present invention improves the efficiency of offline intelligent voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition technology, and particularly to a 4G / 2G / 5G CAT.1 communication system based on offline intelligent voice interaction. Background Art

[0002] With the progress of current algorithms in offline intelligent speech recognition technology, offline speech technology is closer to recognizing normal human-machine conversations and is more widely applied in various electronic products. As the functions of mobile phones become gradually powerful, the menus and functions are particularly rich, and the operations are becoming increasingly complex. To facilitate user operations, directly enter the function menu required by the user through voice operations and directly operate the function, such as entering the camera function, entering the text message function, and then continuing to use voice to implement shooting and text message writing.

[0003] Chinese Patent Publication No.: CN102339604A discloses a voice intelligent interaction system, an intelligent tactile voice dialogue and timed playback learning system capable of performing multi-level voice conversations, including a content storage and playback module (1), a voice microphone collector (2), a startup device module (3), a playback system (4), and a voice recognition processing chip (5). The voice recognition processing chip (5) and the voice microphone collector (2) are connected to the content storage and playback module (1), the content storage and playback module (1) is connected to the playback system (4), and the startup device module (3) is interconnected with the content storage and playback module (1). This solution does not implement the offline intelligent voice interaction function, has a single application scenario, and does not make special voice interaction functions for key users, and cannot improve the voice interaction efficiency. Summary of the Invention

[0004] To this end, the present invention provides a 4G / 2G / 5G CAT.1 communication system based on offline intelligent voice interaction to overcome the problem of low efficiency of offline intelligent voice interaction in the prior art.

[0005] To achieve the above object, the present invention provides a 4G / 2G / 5G CAT.1 communication system based on offline intelligent voice interaction, including:

[0006] A voice input module for real-time collecting voice vibration signals and user voiceprints, and converting the voice vibration signals into audio signals;

[0007] A voiceprint recognition module for verifying the user identity according to the user voiceprint and recognizing the user characteristics;

[0008] An artificial compensation module for compensating the user characteristic recognition result according to the artificial check result;

[0009] A voice analysis module, used to screen and output audio signal features, and also used to identify emergency instructions according to preset audio keywords;

[0010] A voice algorithm recognition module, used to perform algorithm recognition on audio signal features according to a preset audio signal feature sequence, and output an algorithm recognition instruction;

[0011] A voice storage module, used to store and update a preset audio signal feature sequence;

[0012] A digital baseband module, used to process the algorithm recognition instruction, and transfer the processed instruction to a function execution module, and also used to judge the execution of the emergency instruction according to a waiting time interval, and process the emergency instruction according to the judgment result;

[0013] An emergency optimization module, used to optimize the execution judgment process of the emergency instruction according to the user feature recognition result;

[0014] A user feedback module, used to judge the emergency optimization effect according to the user feedback satisfaction and the number of user manual operations within a feedback period, adjust the emergency optimization process according to the judgment result, and give a voice storage update prompt;

[0015] A function execution module, used to execute the device function according to the analog signal input by the digital baseband module;

[0016] A 4G / 2G / 5G CAT.1 communication module, used to realize remote calls;

[0017] A contact setting module, used to input, modify, and query contacts.

[0018] Further, the voiceprint recognition module compares the user's voiceprint with a preset voiceprint, and verifies the user's identity according to the comparison result, where:

[0019] When the user's voiceprint is consistent with the preset voiceprint, the voiceprint recognition module determines that the user's identity passes the verification;

[0020] When the user's voiceprint is inconsistent with the preset voiceprint, the voiceprint recognition module determines that the user's identity fails to pass the verification;

[0021] The voiceprint recognition module also trains the recurrent neural network model based on the user feature voice samples. 70% of the user feature voice samples are divided into the training set, and 30% of the user feature voice samples are divided into the test set. The training set is input into the recurrent neural network model for training, and the test set is input into the trained recurrent neural network model to optimize and iterate the parameters in the recurrent neural network model until the accuracy rate of the output result of the recurrent neural network model test set reaches 98%. Then, the recurrent neural network model is output as the user feature recognition model, and the audio signal is input into the user feature recognition model, and the user features are output by the user feature recognition model.

[0022] Further, the artificial compensation module pushes a feature option window to the user terminals of non-key users. The non-key users check the feature options in the feature option window, and judge the compensation situation of the user feature recognition result according to the checked results, where:

[0023] When the feature option checked by the non-key user is a non-key user option, the artificial compensation module determines not to compensate the user feature recognition result;

[0024] When the feature option checked by the non-key user is a key user option, the artificial compensation module determines to compensate the user feature recognition result and compensate the non-key user as a key user.

[0025] Further, the speech analysis module trains the convolutional recurrent neural network model according to the audio signal feature samples. 70% of the audio signal feature samples are divided into the training set, and 30% of the audio signal feature samples are divided into the test set. The training set is input into the convolutional recurrent neural network model for training, and the test set is input into the trained convolutional recurrent neural network model to optimize and iterate the parameters in the convolutional recurrent neural network model until the accuracy rate of the output result of the convolutional recurrent neural network model test set reaches 99%. Then, the convolutional recurrent neural network model is output as the audio signal feature screening model, and the audio signal is input into the audio signal feature screening model, and the audio signal features are output by the audio signal feature screening model;

[0026] The speech analysis module also compares the text information in the audio signal with the preset audio keywords, and identifies the emergency instructions according to the comparison results, where:

[0027] When there is a keyword in the text information of the audio signal that is the same as the preset audio keyword, it is recognized that the audio signal includes an emergency instruction, and the current waiting time ΔT is obtained;

[0028] When there is no keyword in the text information of the audio signal that is the same as the preset audio keyword, it is recognized that the audio signal does not include an emergency instruction.

[0029] Further, the speech algorithm recognition module trains the hidden Markov model according to a preset audio signal feature sequence, divides 70% of the preset audio signal feature sequence into a training set, divides 30% of the preset audio signal feature sequence into a test set, inputs the training set into the hidden Markov model for training, and inputs the test set into the trained hidden Markov model to optimize and iterate the parameters in the hidden Markov model until the accuracy rate of the output result of the hidden Markov model test set reaches 90%. Then, the hidden Markov model is output as a speech algorithm recognition model, and the audio signal features are input into the speech algorithm recognition model, and the algorithm recognition instruction is output by the speech algorithm recognition model.

[0030] Further, when the audio signal includes an emergency instruction, the digital baseband module compares the current waiting time ΔT with the preset waiting time ΔT0, and judges the execution of the emergency instruction according to the comparison result, where:

[0031] When ΔT < ΔT0, it is determined that the emergency instruction is not executed;

[0032] When ΔT ≥ ΔT0, it is determined that the emergency instruction is executed.

[0033] Further, the emergency optimization module urgently optimizes the execution judgment process of the emergency instruction according to the user feature recognition result, where:

[0034] When the user feature recognition result is a non-key user, the preset waiting time ΔT0 in the execution judgment process of the emergency instruction is not urgently optimized;

[0035] When the user feature recognition result is a key user, the preset waiting time ΔT0 in the execution judgment process of the emergency instruction is urgently optimized.

[0036] Further, the user feedback module judges the emergency optimization effect according to the user feedback satisfaction M, the preset user satisfaction M0 and the number of user manual operations within the feedback period, and adjusts the emergency optimization process according to the judgment result, where:

[0037] When M ≥ M0, it is determined that the user is satisfied, and the number of user manual operations p is compared with the preset number of manual operations p0, and the emergency optimization effect is judged according to the comparison result, where:

[0038] If p < p0, it is determined that the emergency optimization effect reaches the standard;

[0039] If p ≥ p0, it is determined that the emergency optimization effect does not reach the standard;

[0040] When M < M0, it is determined that the user is not satisfied and the emergency optimization effect does not reach the standard;

[0041] When the emergency optimization effect does not meet the standard, the way to adjust the emergency optimization process is to set an adjustment coefficient R, where R = 0.3 + 0.7×e -0.3×(p-p0) , e is the base of the natural logarithm. According to the adjustment coefficient R, the preset waiting time ΔTy0 after optimization is adjusted, and the adjusted preset waiting time is set as ΔTyr0, where ΔTyr0 = R×ΔTy0.

[0042] Furthermore, the function execution module compares the analog signal input by the digital baseband module with the preset analog signal in the function database, and identifies the execution of the device function according to the comparison result, where:

[0043] When there is a preset analog signal in the function database that is consistent with the analog signal, the preset instruction corresponding to the preset analog signal is used as the execution instruction;

[0044] When there is no preset analog signal in the function database that is consistent with the analog signal, the analog signal is fed back to the administrator user terminal, and the administrator optimizes the function database;

[0045] The 4G / 2G / 5G CAT.1 communication module uses the user-selected contact as the communication object for 4G / 2G / 5G CAT.1 communication, and adjusts the communication object according to the connection result of the user-selected contact, where:

[0046] When the user-selected contact has been connected, the communication object is not adjusted;

[0047] When the user-selected contact has not been connected, the communication object is adjusted, and the communication object is adjusted to the emergency communication object;

[0048] The 4G / 2G / 5G CAT.1 communication module adjusts the emergency communication object according to the connection result of the emergency communication object, where:

[0049] When the emergency communication object has been connected, the emergency communication object is not adjusted;

[0050] When the emergency communication object has not been connected, the emergency communication object is adjusted, the communication frequencies of each contact within the preset communication cycle are obtained, and the contacts are sorted according to the communication frequencies from large to small to obtain the communication order. According to the communication order, the contacts are used as the emergency communication object in turn until the emergency communication object is connected;

[0051] Furthermore, the contact setting module includes:

[0052] A contact entry unit for entering contacts through voice interaction;

[0053] A contact temporary storage unit for temporarily storing the contact entry form during the contact entry process;

[0054] A contact permanent storage unit for permanently storing the temporarily stored contact entry form according to the contact entry form and voice interaction;

[0055] A contact modification and query unit for modifying and querying the permanently stored contact entry form according to voice interaction;

[0056] The contact entry unit compares the real-time collected voice vibration signal with the preset contact entry voice vibration signal, and judges the contact entry situation according to the comparison result, where:

[0057] When the real-time collected voice vibration signal is inconsistent with the preset contact entry voice vibration signal, the contact entry unit determines not to enter the contact;

[0058] When the real-time collected voice vibration signal is consistent with the preset contact entry voice vibration signal, the contact entry unit determines to enter the contact, and conducts voice interaction with the user according to the contact entry form, where:

[0059] When the contact entry unit enters the numbers corresponding to the preset digits in the contact entry form, it emits an entry voice signal to prompt the user to enter the contact name. After the user enters the contact name, it prompts the user to enter the numbers corresponding to the preset digits, and obtains the voice of the preset number of times input by the user. When the voice contents of each preset number of times are the same, it recognizes the numbers in the voice content of the preset number of times, and uses this number as the number corresponding to the preset digit. When the voice contents of each preset number of times are different, it prompts the user to re-enter the numbers corresponding to the preset digits. After all the numbers corresponding to the preset digits in the contact entry form are entered, it arranges all the numbers corresponding to the preset digits as the contact phone number;

[0060] During the contact entry process, before all the numbers corresponding to the preset digits in the contact entry form are entered, the contact temporary storage unit temporarily stores the contact entry form;

[0061] After all the numbers corresponding to the preset digits in the contact entry form are entered, the contact permanent storage unit permanently stores the temporarily stored contact entry form;

[0062] The contact modification and query unit compares the real-time collected voice vibration signal with the preset modification and query voice vibration signal, and judges the modification and query situations of the contact entry form according to the comparison result, where:

[0063] When the voice vibration signal collected in real time is inconsistent with the voice vibration signal for modifying the preset contact entry form, the contact modification query unit determines not to modify the contacts permanently stored;

[0064] When the voice vibration signal collected in real time is consistent with the voice vibration signal for modifying the preset contact entry form, the contact modification query unit determines to modify the contacts permanently stored, deletes the contact phone numbers of the contacts permanently stored, and enters the contacts through the contact entry unit, the contact temporary storage unit, and the contact permanent storage unit, and uses the entered contacts as the modified contacts;

[0065] When the voice vibration signal collected in real time is inconsistent with the voice vibration signal for querying the preset contact entry form, the contact modification query unit determines not to query the contacts permanently stored;

[0066] When the voice vibration signal collected in real time is consistent with the voice vibration signal for querying the preset contact entry form, the contact modification query unit determines to query the contacts permanently stored and reads aloud the contact phone numbers corresponding to the contacts permanently stored;

[0067] When the voice vibration signal collected in real time is consistent with the voice vibration signal for overall querying, the contact modification query unit determines to conduct an overall query of the contacts permanently stored and reads aloud all the contacts permanently stored and their corresponding contact phone numbers;

[0068] When the voice vibration signal collected in real time is inconsistent with the voice vibration signal for overall querying, the contact modification query unit determines not to conduct an overall query of the contacts permanently stored.

[0069] Compared with the prior art, the beneficial effects of the present invention are as follows. The system collects voice vibration signals and user voiceprints in real time through the voice input module, and converts the voice vibration signals into audio signals, so as to facilitate the system to transmit and process the voice vibration signals for voice interaction. The system verifies the user identity through the voiceprint recognition module and identifies user characteristics, effectively preventing identity fraud and transaction risks. The system also compensates the user characteristic recognition results through the manual compensation module to provide special settings for key users, thereby enhancing the usage experience of key users. The system screens and outputs audio signal features through the voice analysis module and identifies emergency instructions, so as to facilitate the system to analyze the audio signal features, thereby improving the efficiency of voice interaction. The system performs algorithm recognition on the audio signal features through the voice algorithm recognition module and outputs algorithm recognition instructions, so as to facilitate the system to process and transmit the algorithm recognition instructions, further improving the efficiency of voice interaction. The system stores and updates the preset audio signal feature sequence through the voice storage module, thereby improving the applicability of voice interaction. The system processes the algorithm recognition instructions through the digital baseband module, transfers the processed instructions to the function execution module, and also judges the execution of the emergency instructions and processes the emergency instructions according to the judgment results to improve the accuracy of voice recognition, thereby further improving the accuracy of voice interaction. The system optimizes the execution judgment process of the emergency instructions through the emergency optimization module to improve the usage experience of key users. The system also judges the emergency optimization effect through the user feedback module, adjusts the emergency optimization process according to the judgment results, and gives a voice storage update prompt to further improve the usage experience of key users. The system executes the device functions through the function execution module to realize the voice interaction function. The system realizes remote calls through the 4G / 2G / 5G CAT.1 communication module to provide a more comprehensive user service experience. The system inputs, modifies, and queries contacts through the contact setting module to facilitate users to perform addition, deletion, modification, and query operations on contacts. Description of the Drawings

[0070] Figure 1 It is a schematic structural diagram of the 4G / 2G / 5G CAT.1 communication system based on offline intelligent voice interaction in this embodiment;

[0071] Figure 2 It is a schematic structural diagram of the contact setting module in this example. Detailed Embodiments

[0072] In order to make the objectives and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0073] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.

[0074] It should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0075] Please refer to Figure 1 As shown, it is a schematic structural diagram of a 4G / 2G / 5G CAT.1 communication system based on offline intelligent voice interaction. The system includes:

[0076] A voice input module for real-time collecting voice vibration signals and user voiceprints, and converting the voice vibration signals into audio signals;

[0077] A voiceprint recognition module for verifying the user's identity according to the user's voiceprint and recognizing the user's characteristics. The voiceprint recognition module is connected to the voice input module;

[0078] An artificial compensation module for compensating the user characteristic recognition result according to the artificial check result. The artificial compensation module is connected to the voiceprint recognition module;

[0079] A voice analysis module for screening and outputting audio signal characteristics, and also for recognizing emergency instructions according to preset audio keywords. The voice analysis module is connected to the voice input module;

[0080] A voice algorithm recognition module for performing algorithm recognition on the audio signal characteristics according to a preset audio signal characteristic sequence and outputting an algorithm recognition instruction. The voice algorithm recognition module is connected to the voice analysis module;

[0081] A voice storage module for storing and updating the preset audio signal characteristic sequence. The voice storage module is connected to the voice algorithm recognition module;

[0082] A digital baseband module for processing the algorithm recognition instruction and transmitting the processed instruction to the function execution module, and also for judging the execution of the emergency instruction according to the waiting time interval and processing the emergency instruction according to the judgment result. The digital baseband module is connected to the voice algorithm recognition module;

[0083] An emergency optimization module is used to optimize the execution judgment process of emergency instructions according to the user feature recognition results. The emergency optimization module is connected to the voiceprint recognition module and the digital baseband module;

[0084] A user feedback module is used to judge the emergency optimization effect according to the user feedback satisfaction and the number of user manual operations within the feedback period, adjust the emergency optimization process according to the judgment result, and give a voice storage update prompt. The user feedback module is connected to the emergency optimization module;

[0085] A function execution module is used to execute the device functions according to the analog signals input by the digital baseband module. The function execution module is connected to the digital baseband module;

[0086] A 4G / 2G / 5G CAT.1 communication module is used to realize remote calls. The 4G / 2G / 5G CAT.1 communication module is connected to the function execution module;

[0087] A contact setting module is used to enter, modify and query contacts. The contact setting module is connected to the 4G / 2G / 5G CAT.1 communication module.

[0088] Specifically, the system is applied to platforms such as Spreadtrum SC8910, T107, ASR3601, ASR3603, 3603S, etc., to perform voice interaction with users. The voice input module collects voice vibration signals and user voiceprints in real time through a built-in microphone. When the user speaks, the microphone receives the voice and converts the mechanical wave signal generated by the voice into an audio signal.

[0089] Specifically, the 4G / 2G / 5G CAT.1 refers to the fourth-generation, second-generation, and fifth-generation communication technologies applied to communication terminals. In this embodiment, the implementation method of remote calls is not specifically limited, and those skilled in the art can freely set it according to the actual situation as long as the implementation requirements of remote calls are met. For example, the 4G / 2G / 5G CAT.1 communication module can also be 2G communication and 5G communication, that is, the implementation method of remote calls can also be the second-generation and fifth-generation communication technologies.

[0090] Specifically, the system collects voice vibration signals and user voiceprints in real time through a voice input module, and converts the voice vibration signals into audio signals to facilitate the transmission and processing of the voice vibration signals by the system for voice interaction. The system verifies the user identity through a voiceprint recognition module and identifies user characteristics to effectively prevent identity fraud and transaction risks. The system also compensates the user characteristic recognition results through an artificial compensation module to facilitate special settings for key users, thereby enhancing the usage experience of key users. The system screens and outputs audio signal features through a voice analysis module and identifies emergency instructions to facilitate the analysis of the audio signal features by the system, thereby improving the efficiency of voice interaction. The system performs algorithm recognition on the audio signal features through a voice algorithm recognition module and outputs algorithm recognition instructions to facilitate the processing and transmission of the algorithm recognition instructions by the system, further improving the efficiency of voice interaction. The system stores and updates a preset audio signal feature sequence through a voice storage module to improve the applicability of voice interaction. The system processes the algorithm recognition instructions through a digital baseband module, transmits the processed instructions to a function execution module, also judges the execution of emergency instructions, and processes the emergency instructions according to the judgment results to improve the accuracy of voice recognition, thereby further improving the accuracy of voice interaction. The system optimizes the execution judgment process of emergency instructions through an emergency optimization module to improve the usage experience of key users. The system also judges the emergency optimization effect through a user feedback module, adjusts the emergency optimization process according to the judgment results, and gives a voice storage update prompt to further improve the usage experience of key users. The system executes device functions through a function execution module to achieve the voice interaction function. The system realizes remote calls through a 4G / 2G / 5G CAT.1 communication module to facilitate a more comprehensive user service experience. The system enters, modifies, and queries contacts through a contact setting module to facilitate users' operations of adding, deleting, modifying, and querying contacts.

[0091] Specifically, the voice vibration signal of the voice input module refers to the mechanical wave signal generated during the generation and transmission of the user's voice. The user's voiceprint refers to a signal containing the user's voice characteristics. The audio signal refers to the electrical signal and physical quantity that continuously changes in time and amplitude. In this embodiment, the real-time acquisition method is not limited, and those skilled in the relevant art can freely set it as long as the acquisition requirements for the voice vibration signal and the user's voiceprint are met. For example, it can be set to use a microphone for monitoring. In this embodiment, the conversion method of converting the voice vibration signal into an audio signal is not limited, and those skilled in the relevant art can freely set it as long as the requirement of converting the voice vibration signal into an audio signal is met. For example, it can be set to convert the voice vibration signal into an audio signal by quantization. The quantization refers to the process of approximating the continuous values of the signal to a finite number of discrete values.

[0092] Specifically, the voiceprint recognition module compares the user's voiceprint with the preset voiceprint and verifies the user's identity according to the comparison result, where:

[0093] When the user's voiceprint is consistent with the preset voiceprint, the voiceprint recognition module determines that the user's identity passes the verification;

[0094] When the user's voiceprint is inconsistent with the preset voiceprint, the voiceprint recognition module determines that the user's identity fails to pass the verification.

[0095] Specifically, the preset voiceprint refers to the voiceprint data preset according to the user's first input voice sample. The voice sample refers to the audio record used for display, analysis, and reference. The acquisition method of the voice sample can be set such that the user needs to read a text and perform a voice interaction according to the system's prompt, and the read voice is used as the voice sample.

[0096] Specifically, the voiceprint recognition module trains the recurrent neural network model according to the user's characteristic voice sample. 70% of the user's characteristic voice sample is divided into the training set, and 30% of the user's characteristic voice sample is divided into the test set. The training set is input into the recurrent neural network model for training, and the test set is input into the trained recurrent neural network model to optimize and iterate the parameters in the recurrent neural network model until the accuracy rate of the output result of the recurrent neural network model's test set reaches 98%. Then, the recurrent neural network model is output as the user characteristic recognition model, and the audio signal is input into the user characteristic recognition model, and the user characteristics are output by the user characteristic recognition model.

[0097] Specifically, the user feature voice sample refers to a data group of audio signals - user features collected through big data. In this embodiment, the parameter settings of the recurrent neural network model are not specifically limited, and those skilled in the art can, according to the actual situation, set the loss function of the recurrent neural network model as the cross-entropy function, for example. The user feature refers to the feature attribute of whether the user needs to provide special settings identified according to the user audio signal, including key users and non-key users.

[0098] Specifically, the artificial compensation module pushes a feature option window to the user terminal of non-key users. The non-key users check the feature options in the feature option window, and the artificial compensation module judges the compensation situation of the user feature recognition result according to the check result, where:

[0099] When the feature option checked by the non-key user is the non-key user option, the artificial compensation module determines not to compensate the user feature recognition result;

[0100] When the feature option checked by the non-key user is the key user option, the artificial compensation module determines to compensate the user feature recognition result and compensates the non-key user as a key user.

[0101] Specifically, the feature option refers to the window option set according to the user feature, including the non-key user option and the key user option. In this embodiment, the push method is not specifically limited, and those skilled in the art can freely set it according to the actual situation as long as it meets the acquisition requirements of the user feature intention. For example, it can be set to execute window push through a preset instruction to push the feature option window to the terminal screen.

[0102] Specifically, the voice analysis module trains the convolutional recurrent neural network model according to the audio signal feature sample, divides 70% of the audio signal feature sample into the training set and 30% of the audio signal feature sample into the test set, inputs the training set into the convolutional recurrent neural network model for training, and inputs the test set into the trained convolutional recurrent neural network model to optimize and iterate the parameters in the convolutional recurrent neural network model until the accuracy rate of the output result of the test set of the convolutional recurrent neural network model reaches 99%. Then, the convolutional recurrent neural network model is output as the audio signal feature screening model, and the audio signal is input into the audio signal feature screening model, and the audio signal feature is output by the audio signal feature screening model.

[0103] Specifically, the audio signal feature sample refers to an audio signal feature data group obtained by collecting big data. This embodiment does not specifically limit the parameter setting of the convolutional recurrent neural network model. Technical personnel in this field can, according to actual conditions, set the loss function of the convolutional recurrent neural network model to a cross-entropy loss function. The convolutional recurrent neural network model refers to a deep learning model that combines a convolutional neural network and a recurrent neural network. The audio signal feature refers to a series of attributes and parameters used to describe and distinguish different audio signals.

[0104] Specifically, the speech analysis module also compares the text information in the audio signal with the preset audio keywords, and identifies the emergency command according to the comparison result, wherein:

[0105] When there is a keyword consistent with a preset audio keyword in the text information of the audio signal, identifying that the audio signal includes an emergency instruction, and obtaining a current waiting time ΔT;

[0106] When there is no keyword consistent with the preset audio keyword in the text information of the audio signal, it is recognized that the audio signal does not include an emergency command.

[0107] Specifically, the current waiting time ΔT refers to the time interval from recognizing that the audio signal includes an emergency command to the current absence of other operations. The preset audio keywords refer to keywords and phrases pre-set to reflect whether the text information of the audio signal includes emergency commands. This embodiment does not limit the content of the preset audio keywords, and relevant personnel in this field can set them freely. It only needs to meet the requirement of reflecting whether the text information of the audio signal includes emergency commands. For example, the preset audio keywords can be set to "alarm", "first aid" and "fire extinguishing". The emergency command refers to authoritative and mandatory commands and instructions issued to quickly respond to and solve problems.

[0108] Specifically, the speech algorithm recognition module trains the hidden Markov model according to a preset audio signal feature sequence, divides 70% of the preset audio signal feature sequence into a training set, and divides 30% of the preset audio signal feature sequence into a test set, inputs the training set into the hidden Markov model for training, and inputs the test set into the trained hidden Markov model, optimizes and iterates the parameters in the hidden Markov model until the accuracy of the output result of the hidden Markov model test set reaches 90%, outputs the hidden Markov model as a speech algorithm recognition model, and inputs the audio signal features into the speech algorithm recognition model, and the speech algorithm recognition model outputs algorithm recognition instructions.

[0109] Specifically, the preset audio signal feature sequence refers to a data group of audio signal features - instructions collected through big data. In this embodiment, the parameter settings of the hidden Markov model are not specifically limited. Those skilled in the art can, according to the actual situation, set the initial probability distribution of the hidden Markov model to be represented by π. The algorithm recognition instruction refers to the commands and codes that can enable the device to execute operations output by the hidden Markov model.

[0110] Specifically, the voice storage module stores and updates the preset audio signal feature sequence through a storage chip.

[0111] Specifically, the storage chip refers to a semiconductor component used for data storage. In this embodiment, the type of the storage chip is not limited. Those skilled in the relevant art can freely set it as long as it meets the requirements for storing and updating the preset audio signal feature sequence. For example, the type of the storage chip can be set as an electrically erasable programmable read-only memory. In this embodiment, the update method of the voice storage module is not limited. Those skilled in the relevant art can freely set it as long as it meets the requirements for updating the stored data. For example, it can be set that when the update period is reached, a notification prompt is sent to the user, and the user sends the terminal device to the administrator to update the preset audio signal feature sequence. The update period refers to the preset period for updating the preset audio signal feature sequence. For example, the update period can be set to one year.

[0112] Specifically, the digital baseband module converts the algorithm recognition instruction output by the voice algorithm recognition model into an analog signal according to the modulator, and transmits the converted analog signal to the function execution module, the function menu module, and the 4G / 2G / 5G CAT.1 communication module.

[0113] Specifically, the modulator refers to a device that modulates a low-frequency digital signal into a high-frequency digital signal through digital signal processing technology for signal transmission. In this embodiment, the type of the modulator is not limited. Those skilled in the relevant art can freely set it as long as it meets the requirement of converting the algorithm recognition instruction into an analog signal. For example, the type of the modulator can be set as a digital modulator and an analog modulator. The analog signal refers to the information represented by a continuously changing physical quantity.

[0114] Specifically, when the audio signal includes an emergency instruction, the digital baseband module compares the current waiting time ΔT with the preset waiting time ΔT0, and judges the execution of the emergency instruction according to the comparison result, where:

[0115] When ΔT < ΔT0, it is determined not to execute the emergency instruction;

[0116] When ΔT ≥ ΔT0, it is determined to execute the emergency instruction.

[0117] Specifically, the preset waiting time ΔT0 refers to the time interval preset to reflect whether to execute an emergency instruction. In this embodiment, the value of the preset waiting time ΔT0 is not limited, and those skilled in the relevant art can freely set it as long as it meets the requirement of reflecting whether to execute an emergency instruction. For example, the preset waiting time ΔT0 can be set to 5s. In this embodiment, the content of the executed emergency instruction is not limited, and those skilled in the relevant art can freely set it as long as it meets the requirement of reflecting the execution of the emergency instruction. For example, the content of the emergency instruction can be set to call the police and send the current location to the police.

[0118] Specifically, the emergency optimization module performs emergency optimization on the execution judgment process of the emergency instruction according to the user feature recognition result, where:

[0119] When the user feature recognition result is a non-key user, no emergency optimization is performed on the preset waiting time ΔT0 in the execution judgment process of the emergency instruction;

[0120] When the user feature recognition result is a key user, emergency optimization is performed on the preset waiting time ΔT0 in the execution judgment process of the emergency instruction.

[0121] Specifically, the process of performing emergency optimization on the preset waiting time ΔT0 in the execution judgment process of the emergency instruction is to set an emergency optimization coefficient Y, where Y = 0.4 + 0.8×e -0.3×(ΔT-ΔT0) , where e is the base of the natural logarithm. The preset waiting time is emergently optimized according to the emergency optimization coefficient Y, and the preset waiting time after emergency optimization is set as ΔTy0, where ΔTy0 = Y×ΔT0.

[0122] Specifically, the user feedback module judges the emergency optimization effect according to the user feedback satisfaction M, the preset user satisfaction M0, and the number of user manual operations within the feedback period, and adjusts the emergency optimization process according to the judgment result, where:

[0123] When M≥M0, it is determined that the user is satisfied, and the number of user manual operations p is compared with the preset number of manual operations p0, and the emergency optimization effect is judged according to the comparison result, where:

[0124] If p < p0, it is determined that the emergency optimization effect meets the standard;

[0125] If p≥p0, it is determined that the emergency optimization effect does not meet the standard;

[0126] When M < M0, it is determined that the user is not satisfied, and the emergency optimization effect does not meet the standard;

[0127] When the emergency optimization effect fails to meet the standard, the way to adjust the emergency optimization process is to set an adjustment coefficient R, where R = 0.3 + 0.7×e -0.3×(p-p0) , where e is the base of the natural logarithm. The optimized preset waiting time ΔTy0 is adjusted according to the adjustment coefficient R, and the adjusted preset waiting time is set as ΔTyr0, where ΔTyr0 = R×ΔTy0.

[0128] Specifically, the feedback period refers to the expected time interval for collecting user feedback satisfaction and the number of user manual operations. In this embodiment, the time interval of the feedback period is not limited, and those skilled in the relevant art can freely set it as long as it meets the requirement of judging the emergency optimization effect. For example, the feedback period can be set to one month. The user feedback satisfaction M refers to the score given by the user to the emergency optimization effect within the feedback period. The preset user satisfaction M0 refers to the preset value to reflect whether the emergency optimization effect meets the standard. The number of user manual operations p refers to the number of times of manually dialing the emergency call within the user feedback period. The preset number of manual operations p0 refers to the preset number of user manual operations to reflect the judgment result of the emergency optimization effect. In this embodiment, the value of the preset number of manual operations is not limited, and those skilled in the relevant art can freely set it as long as it meets the requirement of reflecting the judgment result of the emergency optimization effect. For example, the preset number of manual operations p0 = 3 times.

[0129] Specifically, the function execution module compares the analog signal input by the digital baseband module with the preset analog signal in the function database, and identifies the execution of the device function according to the comparison result, where:

[0130] When there is a preset analog signal in the function database that is the same as the analog signal, the preset instruction corresponding to the preset analog signal is used as the execution instruction;

[0131] When there is no preset analog signal in the function database that is the same as the analog signal, the analog signal is fed back to the administrator user terminal, and the administrator optimizes the function database.

[0132] Specifically, the function database refers to a set of analog signal - execution instruction data that determines how the device operates and performs specific tasks. The preset analog signal refers to the analog signal preset according to the function database. The preset instruction refers to the execution instruction preset according to the function database. The execution instruction refers to the process of determining that the device executes a program according to the requirements of the instruction. The administrator refers to a person who has the management responsibility for the communication device. The administrator client refers to the terminal used by the administrator who has the ability to set the function database. In this embodiment, the manner in which the administrator optimizes the function database is not limited. Those skilled in the relevant art can freely set it as long as it meets the requirement of the administrator to optimize the function database. For example, it can be set to push a collection window to the administrator client, and the administrator inputs in the collection window, and the content input by the administrator in the collection window is obtained as the optimization solution given by the administrator client.

[0133] Specifically, the 4G / 2G / 5G CAT.1 communication module obtains the execution instruction of the function execution module and performs 4G / 2G / 5G CAT.1 communication according to the execution instruction. The 4G / 2G / 5G CAT.1 communication module uses the user - selected contact as the communication object to perform 4G / 2G / 5G CAT.1 communication, and adjusts the communication object according to the connection result of the user - selected contact. Among them:

[0134] When the user - selected contact is connected, the communication object is not adjusted;

[0135] When the user - selected contact is not connected, the communication object is adjusted, and the communication object is adjusted to an emergency communication object;

[0136] The 4G / 2G / 5G CAT.1 communication module adjusts the emergency communication object according to the connection result of the emergency communication object. Among them:

[0137] When the emergency communication object is connected, the emergency communication object is not adjusted;

[0138] When the emergency communication object is not connected, the emergency communication object is adjusted. The communication frequencies of each contact within the preset communication cycle are obtained, and the contacts are sorted in descending order according to the communication frequencies to obtain the communication order. The contacts are sequentially used as the emergency communication object according to the communication order until the emergency communication object is connected.

[0139] Specifically, the communication object refers to the contact selected by the user actively for communication. The emergency communication object refers to the individuals, teams, and organizations that need to be contacted and communicated with in case of emergency. The emergency communication object is preset by the user on the terminal. The preset communication cycle refers to the preset cycle for adjusting the emergency communication object. In this embodiment, the time interval of the preset communication cycle is not limited, and those skilled in the relevant art can freely set it according to actual needs, as long as the requirement for adjusting the emergency communication object is met. For example, the preset communication cycle can be set to one month. The communication frequency of each contact refers to the cumulative number of communications between the user and each contact in the call record within the preset communication cycle.

[0140] It can be understood that the content of 4G / 2G / 5G CAT.1 communication is not specifically limited in this embodiment, and those skilled in the art can freely set it according to the actual situation. For example, when the execution instruction is to make a call, the contact can be selected for dialing to achieve a remote call.

[0141] Please refer to Figure 2 as shown, which is the structural schematic diagram of the contact setting module of this instance. The contact setting module includes:

[0142] A contact input unit for inputting contacts through voice interaction;

[0143] A contact temporary storage unit for temporarily storing the contact input form during the contact input process;

[0144] A contact permanent storage unit for permanently storing the temporarily stored contact input form according to the contact input form and voice interaction;

[0145] A contact modification and query unit for modifying and querying the permanently stored contact input form according to voice interaction.

[0146] Specifically, the contact input unit compares the real-time collected voice vibration signal with the preset contact input voice vibration signal, and judges the input situation of the contact according to the comparison result, where:

[0147] When the real-time collected voice vibration signal is inconsistent with the preset contact input voice vibration signal, the contact input unit determines not to input the contact;

[0148] When the real-time collected voice vibration signal is consistent with the preset contact input voice vibration signal, the contact input unit determines to input the contact and conducts voice interaction with the user according to the contact input form, where:

[0149] When the contact entry unit enters the numbers corresponding to the preset number of digits in the contact entry form, it emits an entry voice signal to prompt the user to enter the contact name. After the user enters the contact name, it prompts the user to enter the numbers corresponding to the preset number of digits and obtains the voices of the preset number of times entered by the user. When the voice contents of each preset number of times are the same, it recognizes the numbers in the voice content of the preset number of times and uses this number as the number corresponding to the preset number of digits. When the voice contents of each preset number of times are different, it prompts the user to re-enter the numbers corresponding to the preset number of digits. After the numbers of each preset number of digits in the contact entry form are entered, the numbers of each preset number of digits are arranged as the contact phone number.

[0150] Specifically, the contact entry unit compares the real-time collected voice vibration signal with the preset contact entry voice vibration signal and enters the contact according to the comparison result, so as to be compatible with the user's use requirement of using dialect for voice interaction through waveform comparison, and avoid a large amount of data storage through simple waveform comparison, realize the storage of offline data, and realize offline voice interaction. The contact entry unit recognizes the numbers in the voice content of the preset number of times and uses this number as the number corresponding to the preset number of digits when the voice contents of each preset number of times are the same, so as to ensure the accuracy of the input when the user uses dialect for voice interaction.

[0151] Specifically, the voice vibration signal refers to waveform data represented as a time series generated according to the user's voice, which contains information such as the amplitude and frequency of the voice signal. The preset contact input voice vibration signal refers to the preset voice vibration waveform data that is pre-stored and is consistent with the voice of the user reading the content for inputting contacts. The contact input table refers to a table of individual digits of the phone number corresponding to the contact name, which is usually an 11-digit number table. The preset number of digits refers to the number of digits of the phone number preset according to the quantity of individual digits of the phone number, such as 11 digits. In this embodiment, the manner of prompting the user to input the digits corresponding to the preset number of digits is not limited, and those skilled in the art can freely set it according to the actual situation, as long as the prompting input requirements for the user are met. For example, it can be set to prompt the user by reading the text content of "Please enter the contact number" through the microphone, or it can also be set to prompt the user by flashing a red indicator light. The voice of the preset number of repetitions refers to the preset value of the number of repetitions when interacting with the user by voice to input the digits corresponding to the preset number of digits. For example, if the preset number of repetitions is set to 3, then when the user inputs the digit corresponding to the first digit, the user reads 1 three times. The voice content of each preset number of repetitions being the same means that the digits corresponding to the preset number of digits repeated by the user are the same. For example, if the user reads 1 three times when inputting the digit corresponding to the first digit, then the digit corresponding to the first digit is 1. If the user reads 3 three times when inputting the digit corresponding to the second digit, then the digit corresponding to the second digit is 3. The voice content of each preset number of repetitions being different means that the digits corresponding to the preset number of digits repeated by the user are different. For example, if the digits read three times by the user when inputting the digit corresponding to the second digit are 3, 2, 3, then the voice content of each preset number of repetitions is different, and the user is prompted to re-enter the digit corresponding to the second digit. The input of the digits corresponding to each preset number of digits being completed means that the number of digits of the phone number preset according to the quantity of individual digits of the phone number is inputted, such as the input of 11 digits being completed.

[0152] Specifically, during the process of inputting contacts, the contact temporary storage unit temporarily stores the contact input table before the input of the digits corresponding to each preset number of digits in the contact input table is completed.

[0153] Specifically, the contact temporary storage unit temporarily stores the contact input table before the input of the digits corresponding to each preset number of digits in the contact input table is completed, so as to prevent the data of the digits corresponding to each preset number of digits that have not been inputted from being lost when the system powers off, thereby ensuring the security of the input process of the digits corresponding to each preset number of digits.

[0154] It can be understood that in this embodiment, the manner of temporarily storing the contact input table is not limited, and those skilled in the art can freely set it according to the actual situation. For example, the contact input table can be temporarily stored through a random access memory (RAM).

[0155] Specifically, after the numbers in each preset digit of the contact entry form are entered, the contact permanent storage unit permanently stores the temporarily stored contact entry form.

[0156] Specifically, the contact modification and query unit compares the real-time collected voice vibration signal with the preset modification and query voice vibration signal, and judges the modification and query situations of the contact entry form according to the comparison result, where:

[0157] When the real-time collected voice vibration signal is inconsistent with the preset voice vibration signal for modifying the contact entry form, the contact modification and query unit determines not to modify the permanently stored contacts;

[0158] When the real-time collected voice vibration signal is consistent with the preset voice vibration signal for modifying the contact entry form, the contact modification and query unit determines to modify the permanently stored contacts, deletes the contact phone numbers of the permanently stored contacts, and enters the contacts through the contact entry unit, contact temporary storage unit and contact permanent storage unit, and uses the entered contacts as the modified contacts;

[0159] When the real-time collected voice vibration signal is inconsistent with the preset voice vibration signal for querying the contact entry form, the contact modification and query unit determines not to query the permanently stored contacts;

[0160] When the real-time collected voice vibration signal is consistent with the preset voice vibration signal for querying the contact entry form, the contact modification and query unit determines to query the permanently stored contacts and reads out the corresponding contact phone numbers of the permanently stored contacts;

[0161] When the real-time collected voice vibration signal is consistent with the preset overall query voice vibration signal, the contact modification and query unit determines to perform an overall query on the permanently stored contacts and reads out all the permanently stored contacts and their corresponding contact phone numbers;

[0162] When the real-time collected voice vibration signal is inconsistent with the preset overall query voice vibration signal, the contact modification and query unit determines not to perform an overall query on the permanently stored contacts.

[0163] Specifically, the preset contact entry form modified voice vibration signal refers to the preset voice vibration waveform data stored in advance that is consistent with the voice of the user reading the content of modifying the contact. The preset contact entry form query voice vibration signal refers to the preset voice vibration waveform data stored in advance that is consistent with the voice of the user reading the content of querying the contact. In this embodiment, the reading method of the contact number corresponding to the permanently stored contact is not limited, and those skilled in the art can freely set it according to the actual situation, as long as the voice interaction requirement for the contact number corresponding to the permanently stored contact is met. For example, it can be set that when the voice vibration signal collected in real time is consistent with the preset contact entry form query voice vibration signal, the system performs voice broadcast through the built-in microphone, and the broadcast content is the name of the contact corresponding to the permanently stored contact and their phone number. The preset overall query voice vibration signal refers to the preset voice vibration waveform data stored in advance that is consistent with the voice of the user reading the content of querying all contacts.

[0164] It can be understood that in this embodiment, the contact setting module can also operate on the address book through the LCD interface menu. The LCD interface menu refers to the user interface displayed on the liquid crystal display screen for navigating, selecting, and configuring various functions, settings, and options. The content of operating on the address book includes entering contacts, modifying contacts, and querying contacts. In this embodiment, the operation method of the LCD interface menu on the address book is not limited, and relevant technicians in the art can freely set it according to actual needs, as long as the operation requirement of the LCD interface menu on the address book is met. For example, it can be set that the operation method of the LCD interface menu on the address book is through manual operation by the user.

[0165] So far, the technical solution of the present invention has been described in combination with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or replacements to the relevant technical features, and the technical solutions after these changes or replacements will all fall within the protection scope of the present invention.

Claims

1. A 4G / 2G / 5G CAT.1 communication system based on offline intelligent voice interaction, characterized in that: include: The voice input module is used to collect voice vibration signals and user voiceprints in real time and convert the voice vibration signals into audio signals; Voiceprint recognition module, used to verify the user's identity based on the user's voiceprint and identify the user's characteristics; The manual compensation module is used to compensate the user feature recognition result according to the manual selection result, including: the manual compensation module pushes a feature option window to the user terminal of the non-key user, and the non-key user selects the feature option in the feature option window, and the compensation of the user feature recognition result is judged according to the selection result, wherein: When the feature option selected by the non-key user is a non-key user option, the manual compensation module determines not to compensate the user feature recognition result; When the feature option selected by the non-key user is a key user option, the manual compensation module determines to compensate the user feature recognition result and compensates the non-key user to a key user; A speech analysis module, which is used to filter and output audio signal features, and also to identify emergency commands based on preset audio keywords; A speech algorithm recognition module is used to perform algorithm recognition on audio signal features according to a preset audio signal feature sequence and output an algorithm recognition instruction; A speech storage module, used to store and update a preset audio signal feature sequence; The digital baseband module is used to process the algorithm recognition instruction and pass the processed instruction to the function execution module, and is also used to judge the execution of the emergency instruction according to the waiting time interval and process the emergency instruction according to the judgment result; When the audio signal includes an emergency command, the digital baseband module compares the current waiting time ΔT with the preset waiting time ΔT0, and determines the execution of the emergency command according to the comparison result, wherein: When ΔT<ΔT0, it is determined that the emergency instruction is not executed; When ΔT≥ΔT0, it is determined that the emergency instruction is executed; The emergency optimization module is used to optimize the execution judgment process of the emergency command according to the user feature recognition results, where: When the user feature identification result is a non-key user, the preset waiting time ΔT0 in the execution judgment process of the emergency command is not urgently optimized; When the user feature identification result is a key user, the preset waiting time ΔT0 in the execution judgment process of the emergency command is urgently optimized; The user feedback module is used to judge the emergency optimization effect according to the user feedback satisfaction and the number of user manual operations within the feedback cycle, adjust the emergency optimization process according to the judgment result, and provide voice storage update prompts; A function execution module is used to execute the device function according to the analog signal input by the digital baseband module; 4G / 2G / 5G CAT.1 communication module for remote calls; The contact setting module is used to enter, modify and query contacts.

2. The 4G / 2G / 5GCAT.1 communication system based on offline intelligent voice interaction according to claim 1 is characterized in that: The voiceprint recognition module compares the user's voiceprint with the preset voiceprint and verifies the user's identity based on the comparison result, wherein: When the user's voiceprint is consistent with the preset voiceprint, the voiceprint recognition module determines that the user's identity has been verified; When the user's voiceprint is inconsistent with the preset voiceprint, the voiceprint recognition module determines that the user's identity has not been verified; The voiceprint recognition module also trains the recurrent neural network model according to the user characteristic voice samples, divides 70% of the user characteristic voice samples into a training set, and divides 30% of the user characteristic voice samples into a test set, inputs the training set into the recurrent neural network model for training, and inputs the test set into the trained recurrent neural network model, optimizes the parameters in the recurrent neural network model iteratively until the accuracy of the output result of the recurrent neural network model test set reaches 98%, outputs the recurrent neural network model as a user characteristic recognition model, and inputs the audio signal into the user characteristic recognition model, and the user characteristic recognition model outputs the user characteristic.

3. The 4G / 2G / 5GCAT.1 communication system based on offline intelligent voice interaction according to claim 1, characterized in that: The speech analysis module trains the convolutional recurrent neural network model according to the audio signal feature samples, divides 70% of the audio signal feature samples into a training set, divides 30% of the audio signal feature samples into a test set, inputs the training set into the convolutional recurrent neural network model for training, and inputs the test set into the trained convolutional recurrent neural network model, optimizes and iterates the parameters in the convolutional recurrent neural network model until the accuracy of the output result of the convolutional recurrent neural network model test set reaches 99%, outputs the convolutional recurrent neural network model as an audio signal feature screening model, inputs the audio signal into the audio signal feature screening model, and the audio signal feature screening model outputs the audio signal feature; The speech analysis module also compares the text information in the audio signal with the preset audio keywords and identifies the emergency instructions according to the comparison results, wherein: When there is a keyword consistent with a preset audio keyword in the text information of the audio signal, identifying that the audio signal includes an emergency instruction, and obtaining a current waiting time ΔT; When there is no keyword consistent with the preset audio keyword in the text information of the audio signal, it is recognized that the audio signal does not include an emergency command.

4. The 4G / 2G / 5G CAT.1 communication system based on offline intelligent voice interaction according to claim 1, characterized in that: The speech algorithm recognition module trains the hidden Markov model according to a preset audio signal feature sequence, divides 70% of the preset audio signal feature sequence into a training set, divides 30% of the preset audio signal feature sequence into a test set, inputs the training set into the hidden Markov model for training, and inputs the test set into the trained hidden Markov model, optimizes and iterates the parameters in the hidden Markov model until the accuracy of the output result of the hidden Markov model test set reaches 90%, outputs the hidden Markov model as a speech algorithm recognition model, inputs the audio signal feature into the speech algorithm recognition model, and the speech algorithm recognition model outputs an algorithm recognition instruction.

5. The 4G / 2G / 5G CAT.1 communication system based on offline intelligent voice interaction according to claim 1, characterized in that: The user feedback module judges the emergency optimization effect according to the user feedback satisfaction M, the preset user satisfaction M0 and the number of user manual operations within the feedback cycle, and adjusts the emergency optimization process according to the judgment result, wherein: When M≥M0, the user is determined to be satisfied, and the number of manual operations p of the user is compared with the preset number of manual operations p0, and the emergency optimization effect is judged according to the comparison result, where: If p<p0, it is determined that the emergency optimization effect meets the standard; If p≥p0, it is determined that the emergency optimization effect does not meet the standard; When M<M0, it is determined that the user is dissatisfied and the emergency optimization effect does not meet the standard; When the emergency optimization effect does not meet the standard, the emergency optimization process is adjusted by setting the adjustment coefficient R, where R = 0.3 + 0.7 × e -0.3×(p-p0) , e is the base of the natural logarithm, and the optimized preset waiting time ΔTy0 is adjusted according to the adjustment coefficient R, and the adjusted preset waiting time is set to ΔTyr0, wherein ΔTyr0=R×ΔTy0.

6. The 4G / 2G / 5G CAT.1 communication system based on offline intelligent voice interaction according to claim 1, characterized in that: The function execution module compares the analog signal input by the digital baseband module with the preset analog signal in the function database, and identifies the execution of the device function according to the comparison result, wherein: When there is a preset analog signal consistent with the analog signal in the function database, the preset instruction corresponding to the preset analog signal is used as the execution instruction; When there is no preset analog signal consistent with the analog signal in the function database, the analog signal is fed back to the administrator user end, and the administrator optimizes the function database; The 4G / 2G / 5G CAT.1 communication module uses the contact selected by the user as the communication object for 4G / 2G / 5G CAT.1 communication, and adjusts the communication object according to the connection result of the contact selected by the user, wherein: When the user selects a contact that has been connected, the communication object will not be adjusted; When the user selects a contact that is not connected, the communication object is adjusted to an emergency communication object; The 4G / 2G / 5G CAT.1 communication module adjusts the emergency communication object according to the connection result of the emergency communication object, wherein: When the emergency communication object is connected, no adjustment is made to the emergency communication object; When the emergency communication object is not connected, the emergency communication object is adjusted, the communication frequency of each contact within the preset communication cycle is obtained, and the contacts are sorted in descending order according to the communication frequency to obtain the communication order, and the contacts are used as emergency communication objects in turn according to the communication order until the emergency communication object is connected.

7. The 4G / 2G / 5G CAT.1 communication system based on offline intelligent voice interaction according to claim 1, characterized in that: The contact setting module includes: A contact input unit, used to input contacts through voice interaction; A contact temporary storage unit, used to temporarily store the contact entry form during the contact entry process; A contact permanent storage unit, used to permanently store the temporarily stored contact entry form according to the contact entry form and voice interaction; A contact modification and query unit, used to modify and query the permanently stored contact entry form according to voice interaction; The contact input unit compares the real-time collected voice vibration signal with the preset contact input voice vibration signal, and judges the input status of the contact according to the comparison result, wherein: When the real-time collected voice vibration signal is inconsistent with the preset contact recording voice vibration signal, the contact recording unit determines not to record the contact; When the real-time collected voice vibration signal is consistent with the preset contact entry voice vibration signal, the contact entry unit determines to enter the contact and performs voice interaction with the user according to the contact entry form, wherein: When the contact entry unit enters the number corresponding to the preset digit in the contact entry form, an entry voice signal is sent to prompt the user to enter the contact name, and after the user enters the contact name, the user is prompted to enter the number corresponding to the preset digit, and the voice of the preset number of times input by the user is obtained. When the voice content of each preset number of times is the same, the number of the voice content of the preset number of times is identified and the number is used as the number corresponding to the preset digit. When the voice content of each preset number of times is different, the user is prompted to re-enter the number corresponding to the preset digit. After the numbers of each preset digit in the contact entry form are entered, the numbers of each preset digit are arranged as the contact phone number; During the contact entry process, the contact temporary storage unit temporarily stores the contact entry form before the entry of each preset number of digits in the contact entry form is completed; The contact permanent storage unit permanently stores the temporarily stored contact entry form after the input of the numbers of each preset digit in the contact entry form is completed; The contact modification query unit compares the real-time collected voice vibration signal with the preset modification query voice vibration signal, and judges the modification and query status of the contact entry form according to the comparison result, wherein: When the voice vibration signal collected in real time is inconsistent with the preset contact entry form modification voice vibration signal, the contact modification query unit determines not to modify the permanently stored contact; When the voice vibration signal collected in real time is consistent with the voice vibration signal of the preset contact entry form modification, the contact modification query unit determines to modify the permanently stored contact, deletes the contact phone number of the permanently stored contact, and enters the contact through the contact entry unit, the contact temporary storage unit and the contact permanent storage unit, and uses the entered contact as the modified contact; When the voice vibration signal collected in real time is inconsistent with the voice vibration signal of the preset contact entry form query, the contact modification query unit determines not to query the permanently stored contacts; When the voice vibration signal collected in real time is consistent with the voice vibration signal of the preset contact entry form query, the contact modification query unit determines to query the permanently stored contacts and reads aloud the contact phone number corresponding to the permanently stored contacts; When the voice vibration signal collected in real time is consistent with the preset overall query voice vibration signal, the contact modification query unit determines to perform an overall query on the permanently stored contacts, and reads aloud all permanently stored contacts and their corresponding contact telephone numbers; When the voice vibration signal collected in real time is inconsistent with the preset overall query voice vibration signal, the contact modification query unit determines not to perform an overall query on the permanently stored contacts.

Citation Information

Patent Citations

  • Speech intelligent interaction system

    CN102339604A

  • Intelligent voice interaction system and method

    CN116597839A