AI intelligence-based human voice simulation response voice terminal and method

Through the AI-based intelligent simulated vocal response voice terminal, the inquiry information is automatically collected and processed, and the standard response is matched and played, which solves the vocal cord health risk caused by frequent and high-intensity voice communication of inquiry station personnel, and the effect of reducing the risk of vocal cord disease and improving work efficiency is achieved.

CN120199224APending Publication Date: 2025-06-24SHENZHEN LONGGANG DISTRICT MATERUITY & CHILD HEALTHCARE HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510324156.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The inquiry desk staff faces serious vocal cord health risks due to frequent and high-intensity voice communication, resulting in increased risk of vocal cord fatigue and disease.

Method used

The response and response voice terminal based on AI intelligence is adopted, including a dialogue information collection module, a response matching module and a reply play module. By automatically collecting and processing inquiry information, standard responses are matched and played, reducing the number of voices of staff.

Benefits of technology

It significantly reduces the voice work burden of staff, reduces the risk of vocal cord diseases caused by long-term repeated speaking, improves work efficiency and service quality, and provides timely and accurate responses to inquirers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199224A_ABST
    Figure CN120199224A_ABST
Patent Text Reader

Abstract

The invention discloses an AI intelligence-based voice-simulating response voice terminal and method. The AI intelligence-based voice-simulating response voice terminal comprises a dialogue information acquisition module, a response matching module and a reply playing module, wherein the dialogue information acquisition module acquires inquiry information and converts the inquiry information into original target data; the response matching module comprises a target database and a target database, the response matching module matches the original target data with the target database to obtain reference target data, and the reference target data is mapped to the target database to obtain reference target data; and automatically converting the data of the benchmark into an audio file for playing, or manually selecting the data of the benchmark for playing. According to the invention, through cooperative work of the modules, in inquiry scenes such as hospitals, stations, scenic spots and the like, the voice workload of workers on the inquiry table is reduced, the risk that the workers suffer from vocal nodules and other diseases due to long-time excessive throat use is reduced, the working efficiency and the service quality are improved, and meanwhile, timely and accurate reply can be provided for inquirers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice response terminals, and more specifically, to an answering response voice terminal and method for simulating human voices based on AI intelligence. Background Art

[0002] In various service scenarios in modern society, inquiry desk staff face high-intensity voice communication pressure. In hospitals, with the continuous growth of medical needs, large hospitals receive a large number of patients and their families every day. Hospital guides not only have to face a large number of patient consultations but also deal with repeated inquiries due to patients' anxious emotions. For example, in the outpatient hall of a comprehensive Class III Grade A hospital, the daily average number of inquiries received at the inquiry desk can reach thousands of person-times, covering various aspects such as registration, medical treatment procedures, department locations, and locations of examination items. Prolonged high-intensity speaking causes the vocal cords of the guides to be in a state of fatigue for a long time, and the risk of diseases such as vocal cord nodules increases significantly.

[0003] In addition, as a densely populated place for people flow, whether it is a railway passenger station or a long-distance bus station, service staff need to answer common questions such as train schedules, waiting areas, ticket purchase and refund procedures for passengers. During peak passenger flow periods such as holidays, a station service staff may receive hundreds or even thousands of inquiries per day. Continuous high-intensity voice work places a heavy burden on their vocal cords and has a negative impact on their daily work and life.

[0004] Scenic area service staff are no exception. In popular scenic areas, especially during the peak tourist season, the number of tourists surges. Scenic area service staff need to provide tourists with information such as scenic spot introductions, tour routes, and locations of scenic area facilities. During the peak tourist season, scenic area service staff have to repeatedly answer a large number of similar questions every day, resulting in vocal cord fatigue and affecting work efficiency and service quality.

[0005] In summary, inquiry desk staff in these industries face serious vocal cord health risks in their daily work due to frequent and high-intensity voice communication needs.

[0006] The above deficiencies need to be improved. Summary of the Invention

[0007] In order to solve or alleviate the problem that inquiry desk staff in the above-mentioned existing technologies face serious vocal cord health risks in their daily work due to frequent and high-intensity voice communication needs, the present invention provides an answering response voice terminal and method for simulating human voices based on AI intelligence.

[0008] The technical solution of the present invention is as follows:

[0009] An answering response voice terminal for simulating human voices based on AI intelligence includes:

[0010] A dialogue information collection module, which collects inquiry information and converts it into original targeted data;

[0011] A response matching module, which includes a targeted database and a target database. The response matching module matches the original targeted data with the targeted database to obtain benchmark targeted data, and maps the benchmark targeted data to the target database to obtain benchmark target data;

[0012] A reply playback module, which automatically converts the benchmark target data into an audio file for playback, or manually selects the benchmark target data for playback.

[0013] Furthermore, the dialogue information collection module includes:

[0014] A language recognition unit, which integrates a multi-language recognition engine and is used to recognize the original language of the inquiry information;

[0015] A real-time translation unit, which converts the inquiry information into a preset language, and then converts the translated inquiry information into original targeted data.

[0016] Furthermore, the dialogue information collection module includes a sign language action recognition unit, which collects the sign language action pictures of the inquirer, recognizes the inquiry information in a preset language, and then converts the inquiry information into original targeted data;

[0017] The reply playback module includes a picture display unit. The reply playback module converts the benchmark target data into a video file, and the picture display unit plays the video file.

[0018] Furthermore, the reply playback module includes:

[0019] A voice recording unit, which collects the voice material of the staff reading the text corresponding to the benchmark target data;

[0020] A voice synthesis unit, which inputs the voice material as training data based on deep learning and generates a corresponding voice synthesis model of the staff,

[0021] A voice application unit, which calls the voice synthesis model when the reply playback module plays the audio file and plays the audio file with the voice of the staff.

[0022] An AI intelligent-based simulated human voice response method, which is applicable to the above-mentioned AI intelligent-based simulated human voice response voice terminal, and includes the following steps:

[0023] S1. The dialogue information collection module collects the inquiry information in the surrounding environment and converts the collected inquiry information into original target data;

[0024] S2. The response matching module matches the original target data with the target database, and screens out the data most similar to the original target data from the target database as the benchmark target data;

[0025] Based on the benchmark target data, according to the pre-set mapping rules, perform a mapping search in the target database to obtain the benchmark target data;

[0026] S3. The reply playback module converts the obtained benchmark target data into an audio file and plays it through an audio playback device;

[0027] Or manually select the benchmark target data, convert the selected benchmark target data into an audio file and play it.

[0028] Furthermore, in S1, it includes the following steps:

[0029] S101. Data collection: Collect voice data and training samples in different regions of the target application scenario;

[0030] S102. Data preprocessing: Input the voice samples containing noise into the deep denoising autoencoder algorithm for training. After the training is completed, perform noise reduction processing on the actually collected voice data to separate the pure voice part; adopt the normalization method, calculate the maximum amplitude and minimum amplitude of each voice sample, and uniformly adjust the amplitude of the voice signal to the range of [-1, 1] through linear transformation;

[0031] S103. Feature extraction: Use the pitch period detection algorithm to segment the preprocessed voice data at a certain time interval, calculate the autocorrelation function of each segment of the voice signal at different delay times, determine the pitch period by finding the peak position to obtain the pitch information; adopt the Mel-frequency cepstral coefficient algorithm to convert the preprocessed voice signal to the Mel-frequency domain through the Mel filter bank, and then perform discrete cosine transform to obtain the MFCC coefficients to extract the timbre characteristics; calculate the root mean square energy of the voice signal, and perform a sliding calculation on the preprocessed voice signal with a fixed time window to obtain the intensity characteristics;

[0032] S104. Model training: Select the Gaussian mixture model-hidden Markov model as the speech recognition model; determine the number of mixture components of the Gaussian mixture model and the number of states of the hidden Markov model parameters, and use the extracted speech features and the corresponding text label data for initial training; during the training process, adopt the stochastic gradient descent algorithm combined with the backpropagation algorithm, and divide the training data into multiple small batches and input them into the model in turn;

[0033] S105, Model Testing: Divide a part of the collected voice data as the test data set. The test data set is independent of the training data set and has similar distribution characteristics. Use recognition accuracy and robustness as evaluation metrics. The recognition accuracy is obtained by calculating the ratio of the number of voice samples correctly recognized by the model to the total number of test samples. The robustness evaluation is carried out by adding interference data to the test data set and observing the change in the recognition accuracy of the model under interference conditions.

[0034] Furthermore, in S102, use the Python language and the TensorFlow framework to build a deep denoising autoencoder model, and divide the collected voice samples containing noise into a training set, a validation set, and a test set according to the ratio of 8:1:1 for training.

[0035] Furthermore, in S103, for pitch feature extraction, write an autocorrelation-based fundamental frequency detection program using Matlab software, and segment the preprocessed voice data at a time interval of 10 milliseconds; for timbre feature extraction, use the Mel Frequency Cepstral Coefficient algorithm implemented by the Librosa library in Python, set the number of Mel filter banks to 40, and the number of discrete cosine transform coefficients to 13; for intensity feature extraction, write a program in Python and perform sliding calculations with a time window of 50 milliseconds.

[0036] Furthermore, in S104, use the HMMlearn library in Python to build a model, and divide the extracted voice features and the corresponding text label data into a training set, a validation set, and a test set according to the ratio of 7:2:1; adopt the stochastic gradient descent algorithm, set the initial learning rate to 0.01, the learning rate decay factor to 0.99, update the learning rate every 100 iterations, divide the training data into small batches of size 128, and perform 500 iterations of training.

[0037] Furthermore, in S105, divide 20% of the data from the collected voice data as the test data set according to the similar characteristics of the training data set; use a Python-written evaluation program to calculate the recognition accuracy of the model and its robustness performance under different interference conditions.

[0038] According to the present invention of the above scheme, its beneficial effect is that, by setting up a dialogue information collection module, the present invention can accurately collect the information of inquirers such as patients, passengers, and tourists, and convert it into original target data, realizing the effective collection and standardized processing of various inquiry information in different scenarios, and laying the foundation for subsequent accurate responses. The original target data is matched with the target database through the response matching module, and then the data of the benchmark target is obtained from the target database, realizing the rapid and accurate positioning of high-frequency questions and the accurate extraction of the reply content, avoiding the staff from repeatedly manually searching and organizing the language. The data of the benchmark target is automatically converted into an audio file for playback through the reply playback module, or supports manual selection of playback, realizing that when facing a large number of repeated inquiries, the voice terminal replaces the staff to reply, significantly reducing the number of staff utterances, and allowing the staff's vocal cords to get full rest. Through the collaborative work of each module, it is realized in the inquiry scenes such as hospitals, stations, and scenic spots to reduce the voice workload of the staff at the inquiry desk, reduce their risk of hoarseness or even vocal cord nodules caused by long-term repeated speaking and excessive use of the voice, improve work efficiency and service quality, and also provide timely and accurate replies to the inquirers. In addition, it can improve the physical fitness of staff, show care and concern for employees, and ensure the stability of the team and their work enthusiasm. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0040] Figure 1 It is a schematic diagram of the system architecture of the present invention;

[0041] Figure 2 It is a diagram of the overall method steps of the present invention;

[0042] Figure 3 This is a step diagram of the conversation information collection method of the present invention.

[0043] Among them, the reference numerals in the figure are: 1. dialogue information collection module; 2. response matching module; 3. reply playback module. DETAILED DESCRIPTION

[0044] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0045] It should be noted that when a component is referred to as "fixed" or "set" or "connected" to another component, it can be directly or indirectly located on that other component. The orientations or positions indicated by terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientations or positions shown in the drawings, and are only for convenience of description and should not be construed as limiting the technical solution of the present invention. Terms such as "first", "second", etc. are only for convenience of description purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of technical features. The meaning of "plural" is two or more, unless otherwise clearly and specifically defined. The meaning of "several" is one or more, unless otherwise clearly and specifically defined.

[0046] As Figure 1 shown, in one embodiment of the present invention, the AI intelligent-based simulated human voice response voice terminal includes:

[0047] A dialogue information acquisition module 1, which acquires inquiry information and converts it into original targeted data;

[0048] A response matching module 2, which includes a targeted database and a target database. The response matching module 2 matches the original targeted data with the targeted database to obtain reference targeted data, and maps the reference targeted data to the target database to obtain reference target data;

[0049] A reply playback module 3, which automatically converts the reference target data into an audio file for playback, or manually selects the reference target data for playback.

[0050] Through a high-sensitivity microphone array, the pickup angle range is large, and it can effectively capture voice information in a complex environment. For example, a MEMS microphone array, through the collaborative work of multiple microphones, enhances the ability to collect sounds from different directions and reduces the interference of environmental noise.

[0051] Integrated speech recognition software, with an end-to-end speech recognition model based on deep learning as the core. Through the learning of a large amount of speech data, it can accurately convert speech into text. For example, using the Kaldi speech recognition toolkit, using its pre-trained model and fine-tuning according to the actual application scenario. Input the collected speech signal into the model, and after steps such as feature extraction, acoustic model matching, and language model decoding, output the corresponding text information.

[0052] For the collected or converted inquiry information, it is processed according to the UTF-8 encoding rules to ensure data compatibility between different systems and platforms. It is organized into a specific original targeted data format, such as JSON format, and a data structure containing fields such as information type (e.g., voice, text), keywords (key semantic words extracted from the inquiry information), and timestamp (recording the information collection time) is constructed to facilitate subsequent module processing.

[0053] Database construction:

[0054] The targeted database is constructed using the relational database MySQL. In the database, the data is carefully classified and labeled manually. For example, in the hospital scenario, different types of inquiry information such as registration process, department location, and visit instructions are classified separately, and clear feature tags such as "registration - process" and "department - location - internal medicine" are added to each data item.

[0055] The target database is also constructed using MySQL and is associated with the targeted database. Through mapping rules such as predefined foreign key relationships, each reference targeted data can accurately correspond to the corresponding reference target data in the target database. For example, the data item "registration - process" in the targeted database is linked to the text, audio, or video data in the target database that details the registration process through a foreign key.

[0056] The cosine similarity algorithm is implemented using the cosine_similarity function in the scikit - learn library of Python. When the original targeted data is input, the algorithm calculates the cosine similarity between it and each data item in the targeted database. The algorithm measures the similarity by calculating the cosine value of the angle between two vectors (converting both the original targeted data and the data items in the database into vector form). The data with the highest similarity is selected as the reference targeted data. For example, assuming there are 100 records in the targeted database, the similarity between the original targeted data and these 100 records is calculated, and the one with the highest similarity is selected as the reference targeted data.

[0057] According to the predefined mapping rules, a search is performed in the target database based on the reference targeted data. For example, through an SQL query statement, according to the primary key of the data item in the targeted database (which is used as a foreign key to associate with the target database), the corresponding reference target data is retrieved from the target database.

[0058] If the data of the reference target is in text form, a mainstream text-to-speech engine, such as the TTS engine of iFlytek, is adopted. By calling the API provided by it, the text data is sent to the engine, and the voice style (such as Mandarin female voice, male voice, speech rate, intonation and other parameters) is set. The engine returns the corresponding audio file. In terms of hardware, a high-fidelity speaker is equipped to ensure clear sound quality and moderate volume during audio playback.

[0059] Adopt a graphical user interface (GUI) or a web-based interface. Provide a manual selection button on the interface, and clear prompt information, such as "Manual Selection Reply", is marked on the button. When the staff clicks the button, a list of reference target data for selection pops up. After the staff selects specific data, the system calls the text-to-speech engine to convert it into an audio file and play it. At the same time, an automatic playback switch is set on the interface. When the automatic playback function is turned on, the system automatically converts the obtained reference target data into an audio file and plays it.

[0060] The dialogue information collection module 1 includes:

[0061] A language recognition unit, the language recognition unit integrates a multi-language recognition engine, and the language recognition unit is used to recognize the original language of the inquiry information;

[0062] A real-time translation unit, the real-time translation unit converts the inquiry information into a preset language, and then converts the translated inquiry information into the original target data.

[0063] On the hardware device, a stable network connection function is required to support the call of the online multi-language recognition engine. The multi-language API of Google Cloud Speech Recognition and the multi-language service of Baidu Translate's speech recognition are selected as the main recognition engines. By registering a developer account, the corresponding API keys are obtained and securely configured into the system so that the engine services can be legally called when needed.

[0064] At the software level, a calling interface is written in the Python language. The multi-threading library (such as the threading library) is used to achieve parallel processing. When the voice of the inquiry information is collected, multiple threads are created, and each thread is responsible for sending the voice data to different recognition engines respectively. For example, create a thread to send the voice data to the Google Cloud Speech Recognition multi-language API, and another thread to send the same voice data to the Baidu Translate's speech recognition multi-language service, making full use of the performance of the multi-core processor to shorten the recognition time.

[0065] Recognition Result Judgment: After receiving the speech data, each recognition engine will perform rapid processing and return the recognition result, along with a confidence score. The confidence score is used to reflect the reliability of the recognition result. After receiving the results returned by each engine, compare their confidence scores. For example, the confidence of the result returned by Google Cloud Speech Recognition is 0.85, and the confidence of the result returned by Baidu Translate Speech Recognition is 0.88. In this case, select the result of Baidu Translate Speech Recognition with a higher confidence as the determination of the original language. By writing Python code, parse the results returned by each engine, extract the confidence score and the recognized language information, and then make comparisons and selections.

[0066] Translation API Access: Taking Tencent Translate API as an example, configure the parameters required for API access in the system, such as API key, application ID, etc. Ensure a stable network connection so that requests can be sent to the API and responses can be received in a timely manner. Use the request library in Python (such as the requests library) to build a communication interface with Tencent Translate API. After the language recognition unit determines the original language of the inquiry information, use the recognized text information and the preset language (such as Chinese) as parameters, and send a POST request to the translation interface of Tencent Translate API through the requests library.

[0067] In the request, set the request header information to ensure the security and correctness of data transmission. For example, set Content-Type to application / json, encapsulate the request data in JSON format, including information such as the text to be translated, source language, and target language.

[0068] Processing of Translated Data:

[0069] After receiving the request, Tencent Translate API performs translation processing and returns the translation result. Write code in the system to parse the received translation result and extract the translated text content. Process the translated inquiry information in the same way as the original targeted data conversion method of the above-mentioned dialogue information collection module 1. For example, construct the translated text into JSON format data containing fields such as information type (in this case, the translated text), keywords (extract key semantic words from the translated text), and timestamp (record the translation completion time). Use the JSON processing library in Python (such as the json library) to organize the translated text into a data structure that meets the requirements of the original targeted data format, so that the subsequent response matching module 2 can process it correctly.

[0070] The said dialogue information collection module 1 includes a sign language action recognition unit. The sign language action recognition unit collects the sign language action pictures of the inquirer, recognizes the inquiry information in the preset language, and then converts the inquiry information into the original targeted data;

[0071] The reply playback module 3 includes a screen display unit. The reply playback module converts the data of the reference target into a video file, and the screen display unit plays the video file.

[0072] Hardware configuration: Select a high-definition camera as the image acquisition device. To ensure that the details of sign language movements can be clearly captured, the resolution of the camera should be no less than 1080P, and the frame rate is set at 60fps or above to ensure that the captured images are smooth and free of stuttering. For example, select a common USB high-definition camera and stably connect it to the device through the USB interface.

[0073] Software algorithms and recognition process: Adopt a sign language movement recognition model based on deep learning, such as a model architecture that combines a convolutional neural network (CNN) and a recurrent neural network (RNN). First, collect a large amount of rich sign language movement video data, covering various common inquiry and reply-related sign language movements, and accurately label each movement. The labeled content includes the inquiry information text in the corresponding preset language (such as Chinese).

[0074] Use Python and a deep learning framework (such as TensorFlow or PyTorch) to process the collected data and train the model. During the training process, divide the video data into a training set, a validation set, and a test set according to a certain ratio. Use the training set data to iteratively train the model, and continuously adjust the model parameters to enable the model to learn the characteristics and patterns of sign language movements.

[0075] In practical applications, when the camera captures the sign language movement image of the inquirer, the video frames are sequentially input into the trained model. The model extracts features from each frame of the image through the CNN, and then uses the RNN to process the time series information, thereby identifying the meaning expressed by the sign language movement and converting it into the inquiry information text in the preset language.

[0076] Finally, convert the recognized inquiry information text into the original target data according to the established rules. For example, construct JSON format data containing fields such as information type (sign language recognition text), keywords (extracted from the inquiry information text), and timestamp (recording the recognition time) to facilitate subsequent module processing.

[0077] Video decoding and conversion: When the reply playback module 3 receives the data of the reference target, first judge the data type. If the data is related to video content, use a professional video decoding library for processing. For example, use the video decoding function in the OpenCV library to decode the data according to the encoding format of the video data (such as H.264, MPEG-4, etc.) and convert it into a playable video file format.

[0078] During the decoding process, according to the display parameters and performance of the device, parameters such as the resolution and frame rate of the video are appropriately adjusted to ensure that the video can be played smoothly on the device and has a good display effect. For example, if the device screen resolution is 1920×1080, the video resolution is adjusted to match it to avoid problems such as stretching or blurring.

[0079] Use the video playback API provided by the operating system to implement the video playback function. In the Windows system, the MediaPlayer API can be called; in the Linux system, the VLC API can be used. Write a control program to implement basic operations for video playback, such as play, pause, fast forward, rewind, etc. Set corresponding control buttons on the operation interface of the device, and the staff can control the video playback by clicking the buttons.

[0080] To improve the inquirer experience, a subtitle display function can be added to the video playback interface. If the data of the reference target contains text information corresponding to the video, it is converted into a subtitle format (such as the SRT format), and the subtitles are synchronously displayed during video playback to facilitate the inquirer to understand the video content.

[0081] The reply playback module 3 includes:

[0082] A tone recording unit that collects voice materials of the staff reading the text corresponding to the data of the reference target;

[0083] A voice synthesis unit that inputs the voice materials as training data based on deep learning and generates a corresponding voice synthesis model for the staff,

[0084] A tone application unit that calls the voice synthesis model when the reply playback module 3 plays an audio file and plays the audio file using the staff's tone.

[0085] Recording environment: Select a quiet room with good sound insulation as the recording venue. Sound absorption treatment is carried out in the room, such as laying sound-absorbing cotton on the walls to reduce echo interference. Ensure that the background noise in the recording environment is below 30 decibels to ensure the purity of the recorded voice materials. Control the indoor temperature at 22℃ - 25℃ and the relative humidity at 40% - 60%. Such environmental conditions help the staff maintain a good vocal state and are also beneficial to the stable operation of the recording equipment.

[0086] Recording device selection and configuration: Use a microphone with high sensitivity and wide frequency response, such as a large diaphragm condenser microphone, which can accurately capture the voice details of the staff and restore the true timbre. For example, the Rode NT-USB Mini microphone can be directly connected to the recording device through the USB interface, which is convenient and fast. It is paired with a professional audio interface, such as the Focusrite Scarlett Solo, to improve the conversion quality of the microphone signal and ensure that the recorded audio has high fidelity. Use professional audio recording software, such as Adobe Audition, and set appropriate sampling rates (usually 44.1 kHz or 48 kHz) and bit depths (16 bits or 24 bits) to ensure the quality of the recorded audio.

[0087] Recording process: The staff should do appropriate vocal warm-ups before recording, such as humming simple scales and practicing breathing control, etc., to achieve the best vocal state. Provide the staff with a clear text script, and the content of the script is the text corresponding to the data of the reference target. When the staff reads, maintain a stable speaking speed (about 150 - 180 words per minute), natural intonation, and avoid stuttering, repetition, or abnormal intonation. Record each text paragraph multiple times, generally recording 3 - 5 times for each paragraph. During the recording process, the staff can adjust according to their own feelings and recording effects. After the recording is completed, select the audio segments with clear voices, natural emotional expressions, and no obvious defects from the multiple recorded audios as the final materials.

[0088] Deep learning framework selection: Select a deep learning-based speech synthesis framework, such as TensorFlow or PyTorch. These frameworks have powerful computing capabilities and rich toolkits, which are convenient for building and training speech synthesis models. Taking TensorFlow as an example, use its high-level API, such as Keras, to build the basic architecture of the speech synthesis model.

[0089] Model construction and training: Adopt advanced speech synthesis model architectures such as WaveNet or Tacotron. WaveNet achieves speech synthesis by generating audio waveforms, while Tacotron first generates Mel spectrograms and then converts them into waveforms through a vocoder.

[0090] Divide the voice materials collected by the timbre recording unit into a training set, a validation set, and a test set according to a ratio of 8:1:1. The training set is used for parameter training of the model, the validation set is used to evaluate the performance of the model during training, and the test set is used to finally evaluate the generalization ability of the model.

[0091] During the training process, set appropriate hyperparameters. For example, set the initial learning rate to 0.001, adopt an exponential decay strategy, and decay by 0.96 every 1000 steps; set the batch size to 32 and the number of iterations to 50000. By continuously adjusting the model parameters, the model learns the timbre characteristics of the staff, including pitch, timbre, intonation, etc.

[0092] Use the training set data to iteratively train the model. During the training process, adjust the model parameters according to the evaluation results of the validation set to prevent overfitting. When the loss function value on the validation set no longer decreases, stop the training and save the trained model.

[0093] Model integration and invocation: Integrate the trained speech synthesis model into the system of the reply playback module 3. When designing the system architecture, reserve a dedicated interface for invoking the speech synthesis model. When the reply playback module 3 needs to play an audio file, first determine whether the timbre of the staff needs to be used. If so, input the text corresponding to the benchmark target data into the speech synthesis model. Write an invocation function using a programming language (such as Python) to interact with the speech synthesis model through the interface, and the model generates an audio simulating the timbre of the staff according to the input text.

[0094] Audio playback and quality control: After the audio is generated, play it through an audio playback device, such as a speaker or headphones. Before playing, adjust the volume of the audio appropriately to ensure that the volume is moderate, neither too large to cause noise interference nor too small to be inaudible.

[0095] Regularly evaluate the effect of the timbre application unit, which can be done by comparing the similarity between the synthesized audio and the original recorded audio, and the feedback from the inquirer, etc. If it is found that the timbre quality of the synthesized audio has decreased or there is a large difference from the real timbre of the staff, retrain the speech synthesis model or adjust the relevant parameters to ensure that the audio file can always be played with the timbre of the staff with high quality.

[0096] Such as Figure 2 and Figure 3 As shown, the AI intelligence-based method for simulating human voice response in an embodiment of the present invention is applicable to the above-mentioned AI intelligence-based speech terminal for simulating human voice response, and includes the following steps:

[0097] S1. The dialogue information collection module 1 collects the inquiry information in the surrounding environment and converts the collected inquiry information into raw target data;

[0098] S2. The response matching module 2 matches the raw target data with the target database and screens out the data most similar to the raw target data from the target database as the benchmark target data;

[0099] Based on the benchmark target data, perform a mapping search in the target database according to the pre-set mapping rules to obtain the benchmark target data;

[0100] S3. The playback module 3 converts the obtained benchmark target data into an audio file and plays it through an audio playback device;

[0101] Or manually select the benchmark target data, convert the selected benchmark target data into an audio file and play it.

[0102] In S1, the following steps are included:

[0103] S101. Data collection: Collect voice data and training samples in different regions of the target application scenario;

[0104] S102. Data preprocessing: Input the voice samples containing noise into the deep denoising autoencoder algorithm for training. After the training is completed, perform noise reduction processing on the actually collected voice data to separate the pure voice part; Use the normalization method to calculate the maximum amplitude and minimum amplitude of each voice sample, and uniformly adjust the amplitude of the voice signal to the range of [-1, 1] through linear transformation;

[0105] Use the Python language and the TensorFlow framework to build a deep denoising autoencoder model, and divide the collected voice samples containing noise into a training set, a validation set, and a test set according to a ratio of 8:1:1 for training.

[0106] S103. Feature extraction: Use the pitch period detection algorithm to segment the preprocessed voice data at a certain time interval, calculate the autocorrelation function of each segment of the voice signal at different delay times, determine the pitch period by finding the peak position, and obtain the pitch information; Use the Mel-frequency cepstral coefficients algorithm to convert the preprocessed voice signal to the Mel-frequency domain through the Mel filter bank, and then perform a discrete cosine transform to obtain the MFCC coefficients to extract the timbre features; By calculating the root mean square energy of the voice signal, perform a sliding calculation on the preprocessed voice signal with a fixed time window to obtain the intensity features;

[0107] For pitch feature extraction, write an autocorrelation method pitch period detection program using Matlab software, and segment the preprocessed voice data at a time interval of 10 milliseconds; For timbre feature extraction, use the Librosa library of Python to implement the Mel-frequency cepstral coefficients algorithm, set the number of Mel filter banks to 40, and the number of discrete cosine transform coefficients to 13; For intensity feature extraction, write a program using Python to perform a sliding calculation with a time window of 50 milliseconds.

[0108] S104. Model Training: Select the Gaussian Mixture Model - Hidden Markov Model as the speech recognition model; determine the number of mixture components of the Gaussian mixture model and the number of states of the hidden Markov model, and perform initialization training using the extracted speech features and corresponding text label data; during the training process, adopt the Stochastic Gradient Descent algorithm combined with the Backpropagation algorithm, and divide the training data into multiple small batches and input them into the model sequentially;

[0109] Use the HMMlearn library in Python to build the model, and divide the extracted speech features and corresponding text label data into a training set, a validation set, and a test set according to the ratio of 7:2:1; adopt the Stochastic Gradient Descent algorithm, set the initial learning rate to 0.01, the learning rate decay factor to 0.99, update the learning rate every 100 iterations, divide the training data into small batches of size 128, and perform 500 iterations of training.

[0110] S105. Model Testing: Divide a part of the collected speech data as the test data set. The test data set is independent of the training data set and has similar distribution characteristics; use the recognition accuracy and robustness as evaluation metrics. The recognition accuracy is obtained by calculating the ratio of the number of speech samples correctly recognized by the model to the total number of test samples, and the robustness evaluation is obtained by adding interference data to the test data set and observing the change in the recognition accuracy of the model under interference.

[0111] Divide 20% of the data from the collected speech data as the test data set according to the similar characteristics of the training data set; use Python to write an evaluation program to calculate the recognition accuracy of the model and its robustness performance under different interference conditions.

[0112] Specifically,

[0113] I. Data Collection

[0114] In the hospital scenario, medical staff familiar with the medical environment carry high-fidelity recording equipment, such as professional portable recorders like Zoom H6, and collect speech data in different time periods such as in the morning, afternoon, and on weekends on weekdays in areas such as the registration desk, waiting area, and various department wards. At the registration desk, focus on collecting conversations between patients and registration staff regarding the registration process, source query of appointment numbers, etc.; in the waiting area, collect conversations among patients and inquiries from patients to the guides about department locations, matters needing attention during treatment, etc.; in the ward, record conversations between patients and medical staff about the condition and inquiries about nursing needs, etc.

[0115] II. Data Preprocessing

[0116] Noise Reduction Processing

[0117] Select Python as the development language and use the TensorFlow deep learning framework to build a Deep Denoising Autoencoder (DDAE) model. First, preprocess the collected speech samples containing various noises, and uniformly convert the audio files into the WAV format with a sampling rate of 16 kHz and a quantization depth of 16 bits. Then, divide these samples into a training set, a validation set, and a test set according to the ratio of 8:1:1. During the training process, set the model parameters, such as setting the learning rate to 0.001, the number of iterations to 500 times, and optimizing and adjusting the number of neurons in the hidden layer according to the characteristics of the speech samples. Through continuous training, the DDAE model gradually learns the characteristic patterns of the noises. After training is completed, use this model to perform noise reduction processing on the actually collected speech data, and accurately separate the pure speech from the mixed speech signal.

[0118] Normalization operation

[0119] With the help of Python's audio processing libraries such as Librosa, read the amplitude data of each speech sample. Write a custom function to calculate the maximum amplitude and minimum amplitude of each sample, and then according to the formula:

[0120]

[0121] Uniformly map the amplitude of the speech samples to the range of [-1, 1]. Through the normalization operation, eliminate the recognition errors caused by the amplitude differences of the speech signals, and provide unified standard data input for subsequent feature extraction and model training.

[0122] III. Feature extraction

[0123] Pitch feature extraction

[0124] Use the signal processing function of Matlab software to write a program for detecting the fundamental period by the autocorrelation method. Segment the preprocessed speech data at a fixed time interval of 10 milliseconds. For each segment of the speech signal, calculate its autocorrelation function within the delay time range of 0 to 20 milliseconds. Through the peak search algorithm, accurately find the peak position of the autocorrelation function, thereby determining the fundamental period of the segment of the speech signal, and then obtaining the corresponding pitch information. These pitch information can reflect the prosodic changes of the speech and provide a feature basis for subsequent speech recognition.

[0125] Timbre feature extraction

[0126] Implement the Mel Frequency Cepstral Coefficient (MFCC) algorithm using the Librosa library in Python. Input the preprocessed speech signal into the MFCC function of the Librosa library. According to the characteristics of the speech signal and the experimental optimization results, set the number of Mel filter banks to 40 and the number of discrete cosine transform coefficients to 13. After a series of complex signal processing and transformation operations, obtain the MFCC coefficients. These coefficients contain rich timbre information and can effectively distinguish the speech features of different speakers, improving the adaptability of the speech recognition system to the speech of different speakers.

[0127] Pre-emphasis processing: Before calculating the MFCC of the preprocessed speech signal, perform pre-emphasis processing first. Since the energy of the speech signal is relatively weak in the high-frequency part, pre-emphasis can enhance the energy of the high-frequency part and improve the clarity and intelligibility of the speech. It is implemented through a first-order high-pass filter, and its transfer function is H(z) = 1 - αz -1 , where α generally takes values between 0.95 and 0.97. For example, when α = 0.97, filter the input speech signal x(n) to obtain the pre-emphasized signal y(n) = x(n) - 0.97x(n - 1), making the high-frequency features of the speech signal more prominent in subsequent processing.

[0128] Framing and windowing: The pre-emphasized speech signal is framed at a fixed length. The frame length is usually set to 20 - 40 milliseconds, and the frame shift is set to 10 - 20 milliseconds. Assuming the frame length is 32 milliseconds and the sampling rate is 16 kHz, then each frame contains sample points; when the frame shift is 10 milliseconds, there is partial overlap between adjacent frames. After framing, window each frame of the signal. The commonly used Hamming window n = 0, 1,..., N - 1 (N is the frame length). Windowing reduces spectral leakage, making each frame of the signal smoothly transition at the boundary and avoiding high-frequency components generated by truncation from interfering with subsequent spectral analysis. For example, for a frame of signal s(n), the windowed signal ensures that the signal can more accurately reflect its frequency characteristics during spectral analysis.

[0129] Fast Fourier Transform (FFT): Perform FFT transformation on each windowed frame of the signal to convert the time-domain signal into a frequency-domain signal and obtain its spectrum. Through FFT, a time-domain sequence of length N is converted into a frequency-domain sequence X(k) of the same length, k = 0, 1,..., N - 1, This step converts the time-domain characteristics of the speech signal into frequency-domain characteristics, facilitating subsequent analysis and processing in the frequency domain and clearly showing the energy distribution of the speech signal at different frequencies.

[0130] Mel filter bank filtering: Construct a set of Mel filter banks, and the center frequencies of this filter bank are distributed according to the Mel frequency scale. The Mel frequency scale is more in line with the frequency perception characteristics of the human auditory system, and the conversion relationship with the linear frequency f is The filter bank generally contains 20 - 40 triangular band-pass filters. If set to 40, these filters are equally spaced in the Mel frequency and then converted back to the linear frequency for design. The spectrum X(k) after FFT is successively passed through the Mel filter bank, and the output of each filter is m = 1, 2, …, M (M is the number of filter banks), where H m (k) is the response of the m-th Mel filter at frequency k, and k1 and k2 are the frequency ranges of this filter. This step enables the energy integration of the speech signal on the frequency scale that conforms to the human auditory characteristics, highlighting the frequency features that are more critical for speech recognition.

[0131] Logarithmic operation and discrete cosine transform (DCT): Perform a logarithmic operation on the energy S m output by the Mel filter bank to obtain L m = log(S m ), converting the multiplication operation into an addition operation, which conforms to the logarithmic characteristics of the human auditory and enhances the dynamic range of the signal at the same time. Then perform a discrete cosine transform (DCT) on L m , n = 1, 2, …, N d (N d is the number of DCT coefficients, generally taking 12 - 13). Through DCT, the signal is converted from the Mel frequency domain to the cepstrum domain to obtain the MFCC coefficients. These coefficients remove the redundant information in the spectrum, retain the most representative features of the speech signal, can effectively distinguish the speeches of different speakers, and improve the adaptability of the speech recognition system to the speeches of different interrogators.

[0132] Extraction of intensity features

[0133] The preprocessed speech signal is calculated by sliding with a 50-millisecond time window. In each time window, according to the formula: where x(n) is the amplitude of the speech signal at time n, and N is the number of sampling points in the time window, accurately calculate the root mean square energy (RMS) of the speech signal. In this way, the intensity of the speech signal at different time periods can be accurately obtained, providing strong support for the speech recognition system to capture the key information emphasized by the interrogator.

[0134] IV. Model training

[0135] Model selection and construction

[0136] Taking the Gaussian Mixture Model - Hidden Markov Model (GMM - HMM) as an example, the HMMlearn library in Python is used to construct the model. During the construction process, through a large number of experiments and parameter tuning, the number of mixture components of the Gaussian mixture model is determined to be 16, and the number of states of the hidden Markov model is 8. The extracted speech features and the corresponding text label data are sorted and labeled, and divided into a training set, a validation set, and a test set according to the ratio of 7:2:1. The training set is used for the initial training of the model, enabling the model to initially learn the mapping relationship between speech features and text content; the validation set is used to evaluate the performance of the model during training, adjust the model parameters in a timely manner to prevent overfitting; the test set is used to finally evaluate the generalization ability and recognition accuracy of the model.

[0137] Training Process Optimization

[0138] During the training process, the Stochastic Gradient Descent (SGD) algorithm is used to optimize the model parameters. The initial learning rate is set to 0.01, the learning rate decay factor is 0.99, and the learning rate is updated every 100 iterations. At the same time, combined with the backpropagation algorithm, the error between the predicted result and the true label is propagated backward from the output layer to the input layer, and the gradient of each parameter is accurately calculated through the chain rule, and then the parameters are updated. To improve the training efficiency and model performance, the training data is divided into small batches of size 128 and input into the model for training in sequence. After 500 iterations of training, the model parameters are continuously adjusted, so that the recognition accuracy of the model on the training data set is continuously improved and gradually reaches the optimal state.

[0139] V. Model Testing

[0140] Test Data Set Construction

[0141] From the collected speech data, 20% of the data is divided as the test data set in strict accordance with the characteristics such as the similar scenarios and speaker distributions of the training data set. In the test data set of the hospital scenario, it is ensured that various problems of patients in different departments are included, such as the condition consultation of internal medicine patients, the surgical - related problems of surgical patients, etc., as well as speech samples with different accents and different emotional states. Through this scientific and reasonable way of constructing the test data set, the performance of the model in practical applications can be comprehensively evaluated.

[0142] Calculation of Evaluation Metrics

[0143] By carefully comparing the text results predicted by the model with the true text labels, accurately counting the number of correctly recognized speech samples, and then dividing by the total number of test samples, the recognition accuracy is obtained. For the robustness evaluation, different intensities of noise are added to the test dataset, such as simulating the equipment noise in a hospital, the broadcast noise at a station, etc., and the speech rate (speeded up or slowed down by 20%) and volume (increased or decreased by 10 dB) of the speech samples are changed. The recognition accuracy of the model under these interference conditions is calculated respectively. By observing the changes in the recognition accuracy, the robustness performance of the model is comprehensively evaluated.

[0144] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An AI-based intelligent voice-simulating answering and response voice terminal, characterized in that: include: A dialogue information collection module, which collects query information and converts it into original targeting data; A response matching module, the response matching module includes a targeting database and a target database, the response matching module matches the original targeting data with the targeting database to obtain reference targeting data, and the reference targeting data is mapped to the target database to obtain reference target data; The reply playing module automatically converts the data of the reference mark into an audio file for playing, or manually selects the data of the reference mark for playing.

2. The AI ​​intelligent human voice simulating answering and response voice terminal according to claim 1, characterized in that: The dialogue information collection module includes: A language recognition unit, wherein the language recognition unit integrates a multi-language recognition engine and is used to recognize the original language of the query information; A real-time translation unit converts the query information into a preset language and then converts the translated query information into original targeting data.

3. The AI ​​intelligent human voice simulating answering and response voice terminal according to claim 1, characterized in that: The dialogue information collection module includes a sign language action recognition unit, which collects the sign language action picture of the interrogator, recognizes it as the inquiry information in a preset language, and then converts the inquiry information into original targeting data; The reply playback module includes a screen display unit, the reply playback module converts the data of the reference mark into a video file, and the screen display unit plays the video file.

4. The AI ​​intelligent human voice simulating answering and response voice terminal according to claim 1, characterized in that: The reply playback module includes: A voice recording unit, the voice recording unit collecting voice material of a staff member reading aloud the text corresponding to the data of the reference mark; A speech synthesis unit, wherein the speech synthesis unit inputs the speech material as training data based on deep learning to generate a speech synthesis model for the corresponding staff member, The timbre application unit calls the speech synthesis model when the reply playback module plays the audio file, and uses the timbre of the staff to play the audio file.

5. The method for simulating human voice based on AI intelligence is characterized in that: The AI ​​intelligent human voice simulating answering and response voice terminal applicable to any one of claims 1 to 4 comprises the following steps: S1, the dialogue information collection module collects the inquiry information in the surrounding environment and converts the collected inquiry information into original targeting data; S2, the response matching module matches the original targeting data with the targeting database, and selects the data most similar to the original targeting data from the targeting database as the benchmark targeting data; Based on the benchmark targeting data, according to the pre-set mapping rules, a mapping search is performed in the target database to obtain the benchmark target data; S3, the reply playback module converts the acquired reference mark data into an audio file and plays it through the audio playback device; Or manually select the data of the benchmark, convert the selected benchmark data into an audio file and play it.

6. The method for simulating human voice based on AI intelligence according to claim 5, characterized in that: In S1, the following steps are included: S101, data collection: collect voice data and training samples in different areas of the target application scenario; S102, data preprocessing: input the speech sample containing noise into the deep noise reduction autoencoder algorithm for training, and after the training is completed, perform noise reduction processing on the actually collected speech data to separate the pure speech part; use the normalization method to calculate the maximum amplitude and minimum amplitude of each speech sample, and adjust the amplitude of the speech signal to the range of [-1,1] uniformly through linear transformation; S103, feature extraction: using the pitch period detection algorithm, the pre-processed speech data is segmented at a certain time interval, the autocorrelation function of each speech signal at different delay times is calculated, the pitch period is determined by finding the peak position, and the pitch information is obtained; using the Mel frequency cepstral coefficient algorithm, the pre-processed speech signal is converted to the Mel frequency domain through the Mel filter bank, and then the discrete cosine transform is performed to obtain the MFCC coefficient to extract the timbre feature; By calculating the root mean square energy of the speech signal, the preprocessed speech signal is subjected to sliding calculation in a fixed time window to obtain the sound intensity feature; S104, model training: selecting a Gaussian mixture model-hidden Markov model as a speech recognition model; determining the number of mixed components of the Gaussian mixture model and the number of states of the hidden Markov model, and performing initialization training using the extracted speech features and corresponding text label data; During the training process, the stochastic gradient descent algorithm is combined with the back propagation algorithm to divide the training data into multiple small batches and input them into the model in sequence; S105, model testing: a part of the collected speech data is divided as a test data set. The test data set is independent of the training data set and has similar distribution characteristics. Recognition accuracy and robustness are used as evaluation indicators. The recognition accuracy is obtained by calculating the ratio of the number of speech samples correctly recognized by the model to the total number of test samples. The robustness is evaluated by adding interference data to the test data set and observing the changes in the recognition accuracy of the model under interference.

7. The method for simulating human voice based on AI intelligence according to claim 5, characterized in that: In S102, a deep denoising autoencoder model is built using Python language and TensorFlow framework, and the collected speech samples containing noise are divided into a training set, a validation set, and a test set in a ratio of 8:1:1 for training.

8. The method for simulating human voice based on AI intelligence according to claim 5, characterized in that: In S103, the pitch feature extraction uses Matlab software to write an autocorrelation method fundamental frequency period detection program, and the preprocessed speech data is segmented at 10 milliseconds. The timbre feature extraction uses Python's Librosa library to implement the Mel frequency cepstral coefficient algorithm, and sets the number of Mel filter groups to 40 and the number of discrete cosine transform coefficients to 13. The sound intensity feature extraction uses Python to write a program and performs sliding calculations with a 50 millisecond time window.

9. The method for simulating human voice based on AI intelligence according to claim 5, characterized in that: In S104, the HMMlearn library of Python is used to build the model, and the extracted speech features and the corresponding text label data are divided into training set, validation set and test set in the ratio of 7:2:1; the stochastic gradient descent algorithm is used, the initial learning rate is set to 0.01, the learning rate decay factor is set to 0.99, the learning rate is updated every 100 iterations, the training data is divided into small batches of size 128, and 500 iterations of training are performed.

10. The method for simulating human voice based on AI intelligence according to claim 5, characterized in that: In S105, 20% of the collected speech data is divided as a test data set according to features similar to the training data set; an evaluation program is written in Python to calculate the recognition accuracy of the model and its robustness performance under different interference conditions.