Voice question answering method and system based on multi-mode voice recognition and semantic understanding

By adopting multimodal speech recognition and semantic understanding technology in the intelligent voice question-and-answer system, the problem of identifying questions in complex environments is solved, and higher recognition accuracy and user experience are achieved.

CN120046107APending Publication Date: 2025-05-27SHENZHEN LINGHAI INTELLIGENT CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510137402.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

It is difficult for existing intelligent voice Q&A systems to accurately determine which discourses are directly asked questions in complex environments, resulting in high misrecognition rates and affecting user experience and effective application of the system.

Method used

Using a method based on multimodal speech recognition and semantic understanding, visual and audio data are obtained through integrated cameras and multi-array microphones, and the audio separation and conversion of target speakers is performed. Combined with face detection, lip information extraction and deep learning fusion models, the target speech waveform is accurately positioned, and the validity of text content is evaluated through a large language model, generating and output answers.

Benefits of technology

It significantly improves the recognition accuracy and semantic understanding ability of voice interactions, reduces the misrecognition rate, and improves the user experience and the system's scenario adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046107A_ABST
    Figure CN120046107A_ABST
Patent Text Reader

Abstract

The invention discloses a voice question answering method and system based on multi-mode voice recognition and semantic understanding. The method comprises the following steps: acquiring an on-site visual image and an audio signal acquired by an integrated camera and a multi-array microphone to obtain multi-modal data; performing audio separation and conversion of a target speaker on the multi-modal data to obtain text content; performing semantic understanding and question effectiveness evaluation on the text content to obtain an evaluation result; when the evaluation result is that the text content belongs to the effective questions, generating a corresponding answer text according to the text content; converting the answer text into voice to obtain answer audio; and outputting the answer audio. By implementing the method provided by the invention, accurate voice interaction in a complex environment can be realized, the recognition accuracy, the semantic understanding ability, the user experience and the scene adaptability are remarkably improved, and the error recognition rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a voice question - answering method, and more specifically to a voice question - answering method and system based on multi - modal speech recognition and semantic understanding. Background Art

[0002] Currently, when dealing with the problems of speech recognition in complex acoustic environments such as multiple people asking questions simultaneously and cocktail parties, the industry has developed a set of multi - modal speech separation and speech recognition algorithms based on face lip - shapes and audio. This set of algorithms improves the accuracy of speech detection and recognition by combining visual information such as facial features, lip movements, and auditory information, especially in cases where background noise is high or there are multiple sound sources. It can effectively distinguish different speakers and improve the quality of speech signals, making the subsequent automatic speech recognition process more accurate. In addition, in the field of intelligent voice question - answering, existing systems have integrated the above - mentioned multi - modal speech separation and speech recognition algorithms and combined them with natural language processing models to achieve an automated process from speech input to text parsing and then to providing answers. These systems usually have a certain ability to understand conversations and maintain context, and can provide interactive information query services for users.

[0003] However, although existing intelligent voice question - answering systems have made significant progress in speech separation and recognition, they still face an important challenge in practical applications: it is difficult for them to accurately determine whether a specific utterance is a question directly addressed to the question - answering system. When there are multiple people talking in the environment, existing systems may not be able to precisely distinguish which words are clear questions for the question - answering system and which are just conversations among other people on - site or meaningless sounds. Due to the lack of an effective mechanism to evaluate and filter out irrelevant speech segments, the system may respond to inquiries that are not specifically addressed to it, resulting in misrecognition. Although some advanced question - answering systems can handle a certain degree of context information, they are not sensitive enough to capture the intentions of utterances in complex scenarios, especially in multi - person interaction environments. In summary, although current technologies perform well in speech separation and basic question - answering functions, there is still room for improvement in ensuring that only real questions are responded to in complex environments. This not only affects the user experience but also limits the effective application of such systems in a wider range of scenarios.

[0004] Therefore, it is necessary to design a new method to achieve precise voice interaction in complex environments, significantly improving recognition accuracy, semantic understanding ability, user experience, scene adaptability, and reducing the misrecognition rate. Summary of the Invention

[0005] The purpose of the present invention is to overcome the defects of the prior art and provide a voice question - answering method and system based on multi - modal speech recognition and semantic understanding.

[0006] To achieve the above object, the present invention adopts the following technical solutions: A voice question-and-answer method based on multi-modal speech recognition and semantic understanding, including:

[0007] Obtain the on-site visual images and audio signals collected by the integrated camera and multi-array microphones to obtain multi-modal data;

[0008] Perform audio separation and conversion of the target speaker on the multi-modal data to obtain text content;

[0009] Perform semantic understanding and evaluation of the validity of the question on the text content to obtain an evaluation result;

[0010] When the evaluation result indicates that the text content is a valid question, generate a corresponding answer text according to the text content;

[0011] Convert the answer text into speech to obtain an answer audio;

[0012] Output the answer audio.

[0013] Its further technical solution is: The performing audio separation and conversion of the target speaker on the multi-modal data to obtain text content includes:

[0014] Perform face detection, lip information extraction, and determination of the orientation of the target speaker on the multi-modal data to obtain a determination result;

[0015] Separate the target speech waveform from the multi-modal data according to the determination result;

[0016] Convert the target speech waveform into text content.

[0017] Its further technical solution is: The performing face detection, lip information extraction, and determination of the orientation of the target speaker on the multi-modal data to obtain a determination result includes:

[0018] Process each frame of the on-site visual image in the multi-modal data, identify the face therein, and calculate the position of the face relative to the center of the picture;

[0019] Obtain the specific position of the lips from the face by applying a face semantic segmentation model;

[0020] Based on the position information of the face relative to the center of the picture, calculate the angle deviating from the center line of the picture to determine the orientation of the target speaker;

[0021] Wherein, the determination result includes the information of the face, the specific position of the lips, and the orientation of the target speaker.

[0022] A further technical solution thereof is: separating the target speech waveform from the multimodal data according to the determination result includes:

[0023] Inputting the determination result and the audio signal into a deep learning fusion model to generate an audio mask, and separating the audio of the target speaker according to the audio mask to obtain the target speech waveform.

[0024] A further technical solution thereof is: converting the target speech waveform into text content includes:

[0025] Using the ASR algorithm service to convert the target speech waveform into text content.

[0026] A further technical solution thereof is: evaluating the semantic understanding and question validity of the text content to obtain an evaluation result, including:

[0027] Inputting the text content into a trained large language model to evaluate whether the text content constitutes a valid query to obtain an evaluation result.

[0028] A further technical solution thereof is: when the evaluation result is that the text content belongs to a valid question, generating a corresponding answer text according to the text content, including:

[0029] When the evaluation result is that the text content belongs to a valid question, calling a large language model according to the text content to generate an answer to obtain the answer text.

[0030] A further technical solution thereof is: converting the answer text into speech to obtain an answer audio, including:

[0031] Using a third-party TTS algorithm service that supports anthropomorphic voices to convert the answer text into an answer audio for output.

[0032] The present invention also provides a voice question-answering system based on multimodal speech recognition and semantic understanding, including:

[0033] A data acquisition unit for acquiring on-site visual images and audio signals collected by an integrated camera and a multi-array microphone to obtain multimodal data;

[0034] A separation and conversion unit for separating and converting the audio of the target speaker from the multimodal data to obtain text content;

[0035] An evaluation unit for evaluating the semantic understanding and question validity of the text content to obtain an evaluation result;

[0036] An answer generation unit, configured to generate a corresponding answer text according to the text content when the evaluation result indicates that the text content belongs to a valid question;

[0037] A conversion unit, configured to convert the answer text into speech to obtain an answer audio;

[0038] An output unit, configured to output the answer audio.

[0039] A further technical solution thereof is that: the separation and conversion unit includes:

[0040] A determination unit, configured to perform face detection, lip information extraction, and determination of the orientation of the target speaker on the multimodal data to obtain a determination result;

[0041] A separation unit, configured to separate a target speech waveform from the multimodal data according to the determination result;

[0042] A conversion unit, configured to convert the target speech waveform into text content.

[0043] The beneficial effects of the present invention compared with the prior art are: the present invention

[0044] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings

[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1 It is a schematic diagram of an application scenario of a voice question-answering method based on multimodal speech recognition and semantic understanding provided by an embodiment of the present invention;

[0047] Figure 2 It is a schematic flowchart of a voice question-answering method based on multimodal speech recognition and semantic understanding provided by an embodiment of the present invention;

[0048] Figure 3 It is a schematic sub-flowchart of a voice question-answering method based on multimodal speech recognition and semantic understanding provided by an embodiment of the present invention;

[0049] Figure 4 It is a schematic sub-flowchart of a voice question-answering method based on multimodal speech recognition and semantic understanding provided by an embodiment of the present invention;

[0050] Figure 5Schematic block diagram of a voice question - answering system based on multi - modal speech recognition and semantic understanding provided by an embodiment of the present invention;

[0051] Figure 6 Schematic block diagram of a separation and conversion unit of a voice question - answering system based on multi - modal speech recognition and semantic understanding provided by an embodiment of the present invention;

[0052] Figure 7 Schematic block diagram of a confirmation sub - unit of a voice question - answering system based on multi - modal speech recognition and semantic understanding provided by an embodiment of the present invention;

[0053] Figure 8 Schematic block diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0056] It should also be understood that the terms used in this specification of the present invention are merely for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0057] It should be further understood that the term "and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0058] Please refer to Figure 1 and Figure 2 , Figure 1 Schematic diagram of an application scenario of a voice question - answering method based on multi - modal speech recognition and semantic understanding provided by an embodiment of the present invention. Figure 2Schematic flowchart of the voice question-answering method based on multimodal speech recognition and semantic understanding provided by an embodiment of the present invention. The voice question-answering method based on multimodal speech recognition and semantic understanding is applied to a server. The server interacts with a device integrated with a camera and a multi-array microphone. Through the acquisition of the camera and microphone of multimodal data, combined with the audio separation and conversion of the target speaker, the system can clearly recognize speech in a complex environment and reduce noise interference; by adopting face detection, lip shape extraction and orientation determination technologies, the extraction accuracy of the target speaker's speech is effectively improved, especially in a multi-person environment; based on the deep learning fusion model and audio mask generation, the audio waveform of the target speaker can be accurately separated, thereby improving the accuracy of speech recognition; with the support of the ASR algorithm, the system can efficiently convert the target speech waveform into text content, ensuring the rapidity and accuracy of the conversion; combined with the semantic understanding and effectiveness evaluation of the text content by the large language model, efficient and accurate question recognition is realized, greatly enhancing the semantic understanding ability; finally, the TTS technology with anthropomorphic timbre is used to convert the answer into speech, which not only improves the user experience, but also optimizes the scene adaptability and overall interaction quality.

[0059] Figure 2 It is a schematic flowchart of the voice question-answering method based on multimodal speech recognition and semantic understanding provided by an embodiment of the present invention. As Figure 2 shown, the method includes the following steps S110 to S160.

[0060] S110. Obtain the on-site visual image and audio signal collected by the integrated camera and multi-array microphone to obtain multimodal data.

[0061] In this embodiment, the multimodal data refers to the on-site visual image and the audio signal.

[0062] The system first needs to configure and activate the integrated camera and multi-array microphone. These devices are designed to capture the on-site visual image and audio signal, providing the basic data for subsequent processing. To ensure the best data acquisition effect, the devices are usually optimized according to the specific usage scenario, such as adjusting the angle, focal length of the camera, and the sensitivity of the microphone and other parameters to adapt to spaces of different sizes and shapes.

[0063] The camera is strategically placed to cover the entire interaction area, ensuring that all potential primary speakers can be captured. It not only records the live video stream but may also include high-resolution static pictures. The video frames are analyzed by computer vision algorithms to identify and track individuals in the frame. Combining face recognition technology and lip movement analysis, it is possible to further determine who the current primary speaker is and whether they are speaking. In addition to identifying faces, gestures, postures, and other non-verbal cues can be used to enhance the understanding of user intentions.

[0064] A set of microphones arranged in a specific geometry is used to collect sound signals from different directions. This layout helps to achieve spatial audio capture, that is, not only recording the sound itself but also perceiving the direction of its source. By calculating the time difference and phase difference of the sound signals received by each microphone, the position of the sound source can be accurately determined, and the speech of the target speaker can be effectively distinguished even in a noisy environment. Built-in algorithms can reduce the impact of background noise and eliminate echo interference caused by speaker playback, thereby improving the quality of the speech signal.

[0065] To ensure the consistency between visual and audio data, each piece of data is marked with an accurate timestamp. This enables perfect matching of the two during post-processing. For example, when performing speaker separation, the audio-visual materials at the same moment can be used to more accurately lock the speaker.

[0066] Finally, the visual images and audio signals are integrated as complementary information sources to form a complete multi-modal dataset. This dataset not only contains rich raw information but also, after preprocessing, can directly support subsequent advanced processing tasks such as speaker separation, speech recognition, and semantic understanding.

[0067] In summary, step S110 is the starting point of the entire intelligent interaction system. Through carefully designed hardware configurations and advanced software algorithms, it ensures the successful acquisition of high-quality multi-modal data, laying a solid foundation for subsequent intelligent analysis and service response.

[0068] S120. Perform audio separation and conversion of the target speaker on the multi-modal data to obtain text content.

[0069] In this embodiment, step S120 focuses on in-depth analysis of the collected multi-modal data (including visual images and audio signals), aiming to effectively separate the voice of the target speaker and convert this voice into text content in real time. This process not only improves the speech recognition accuracy in a multi-person conversation scenario but also ensures that the speech of the target speaker can be accurately captured and understood even in a noisy environment.

[0070] In this embodiment, the text content refers to the text formed by converting the audio of the target speaker.

[0071] In one embodiment, please refer to Figure 3 , the above step S120 may include steps S121 to S123.

[0072] S121. Perform face detection, lip information extraction, and determination of the orientation of the target speaker on the multimodal data to obtain a determination result.

[0073] In one embodiment, please refer to Figure 4 , the above step S121 may include steps S1211 to S1213.

[0074] S1211. Process each frame of the on-site visual image in the multimodal data, identify the face therein, and calculate the position of the face relative to the center of the screen.

[0075] In this embodiment, for each frame of the on-site visual image obtained from the camera, the system will use computer vision algorithms to identify the face therein. Then, by comparing the position relationship between the identified face and the center of the screen, the angle by which the face deviates from the center line is calculated. This step helps to preliminarily determine who may be the current main speaker.

[0076] S1212. Obtain the specific position of the lip shape from the face by applying a face semantic segmentation model.

[0077] In this embodiment, a pre-trained face semantic segmentation model is used to accurately locate and extract the information of the lip area. This technology can provide additional clues about mouth movements to help more accurately distinguish the person who is speaking.

[0078] S1213. Calculate the angle deviating from the center line of the screen based on the position information of the face relative to the center of the screen to determine the orientation of the target speaker;

[0079] Among them, the determination result includes the information of the face, the specific position of the lip shape, and the orientation of the target speaker. The face information refers to the position information where the face is located.

[0080] In this embodiment, based on the data obtained in the previous two steps, especially the position information of the face relative to the center of the screen, the direction of the target speaker can be further accurately located. This information is crucial for subsequent audio processing because it provides an important reference basis for sound source localization.

[0081] S122. Separate the target speech waveform from the multimodal data according to the determination result.

[0082] In this embodiment, the target speech waveform refers to the audio content of the target speaker.

[0083] Specifically, the determination result and the audio signal are input into a deep learning fusion model to generate an audio mask, and the audio of the target speaker is separated according to the audio mask to obtain the target speech waveform.

[0084] In this embodiment, after obtaining the information related to the target speaker from the visual data, the next step is to combine the audio stream to perform more refined voice separation. Specifically:

[0085] The above determination results (i.e., face information, lip position, and speaker orientation) and the original audio signal are fed into a trained deep learning fusion model. This model generates an audio mask based on multi-modal features, which is used to highlight the voice features of the target speaker while suppressing other irrelevant voice components. Finally, the generated audio mask is used to separate the clear speech waveform of the target speaker from the mixed audio signal, that is, the target speech waveform.

[0086] S123. Convert the target speech waveform into text content.

[0087] In this embodiment, the ASR algorithm service is used to convert the target speech waveform into text content.

[0088] In this embodiment, the last step is to convert the separated target speech waveform into a text form that is easy to understand and process:

[0089] Apply the ASR (Automatic Speech Recognition) algorithm service: Using advanced automatic speech recognition technology, quickly and accurately convert the target speech waveform into the corresponding text content.

[0090] Text output: After the conversion is completed, the system can directly output or store this text content for subsequent application programs or users to use directly.

[0091] Through the above detailed description of step S120, it can be seen that the multi-modal data analysis method in this embodiment makes full use of the advantages of visual and auditory information, realizes high-precision separation of the target speaker's audio, and successfully converts it into text content, thus greatly improving the user experience and efficiency of the intelligent interaction system.

[0092] S130. Evaluate the semantic understanding and question validity of the text content to obtain an evaluation result.

[0093] In this embodiment, the evaluation result refers to the determination result of whether the text content is a complete and valid question.

[0094] Specifically, the text content is input into the trained large language model to evaluate whether the text content constitutes a valid query, so as to obtain an evaluation result.

[0095] In the intelligent voice Q&A system, step S130 focuses on deeply analyzing the text content converted from audio, with the aim of determining whether this text constitutes a valid query for the system. This process is crucial for ensuring that the system's response is both relevant and accurate. Specifically, to achieve the goal of this step, first, specific fine-tuning is required based on an open-source large language model (such as LLaMA-2 LLM). This fine-tuning process uses an annotated dataset prepared according to the actual application scenario, which contains dialogue texts in different scenarios and their corresponding meaningfulness level labels (for example: [TEXT] Is this the information you are looking for? - [TAG] High meaningfulness). Through such training, the model can better understand the language patterns in the specific application environment.

[0096] Once the model is trained, the next step is to input the text content obtained in step S123 into this trained large language model. The task of the model is to evaluate the input text, determine whether it constitutes a complete and valid question, and also consider its meaningfulness. Here, meaningfulness refers to whether the text content clearly expresses the user's question or request to the system, rather than irrelevant conversations or other background noises.

[0097] The large language model will output an evaluation result indicating whether the input text content is regarded as a complete and valid question. This evaluation is not only based on the grammatical structure (i.e., whether the sentence is a question), but also delves into the semantic level, considering the actual meaning and intention of the text.

[0098] Only when the evaluation result shows that the text is indeed a valid query for the system will the system's answer mechanism be triggered. This means that the system will not respond to non-query or meaningless conversation content, thereby improving the relevance and efficiency of the interaction.

[0099] By implementing step S130, this system can more accurately distinguish which utterances are queries for the system and which are not. This method not only improves the relevance and accuracy of the Q&A interaction, but also enhances the user experience, because it ensures that users get the information they really want, rather than a misresponse to irrelevant conversation content. In addition, through the effective utilization and fine-tuning of the large language model, the system's performance can be continuously improved to adapt to changing application requirements and language environments.

[0100] S140. When the evaluation result indicates that the text content belongs to a valid question, generate a corresponding answer text according to the text content.

[0101] In this embodiment, the answer text refers to the answer corresponding to the text content.

[0102] Specifically, when the evaluation result indicates that the text content is a valid question, the large language model is called according to the text content to generate an answer, so as to obtain the answer text.

[0103] For application scenarios that require rapid deployment and have few customization requirements, it is possible to choose to use large language model as a service (LLM-as-a-Service) provided by a third party. These services are usually widely trained, can handle various types of inquiries, and provide API interfaces that are easy to integrate.

[0104] For applications with specific domain requirements or those that hope to improve accuracy and relevance, it is possible to further fine-tune based on open-source large language models (such as LLaMA-2 LLM). This approach allows developers to prepare specialized training datasets according to the actual application scenario. These datasets contain [TEXT] question - [TAG] answer pairs for specific domains. In this way, the performance of the model is optimized to make it more in line with the specific business logic and service objects.

[0105] After the system determines that a question is valid, it passes this question as input to the selected large language model service or the fine-tuned model. During this process, the model does not generate a complete answer at once, but gradually constructs the answer in a streaming manner. This means that as the model's understanding of the question deepens, the answer will continue to expand and improve, thus ensuring the timeliness and coherence of the answer.

[0106] In order to provide high-quality answers, when generating answers, the large language model not only considers the content of the question itself, but also combines context information and other factors that may affect the answer. This helps to ensure that the generated answer is not only detailed but also highly relevant, meeting the user's query needs.

[0107] Finally, through the above process, the system can generate a corresponding answer text for each text content confirmed as a valid question. This answer text is the best response to the original question, reflecting both a deep understanding of the question and providing specific and useful information.

[0108] By implementing step S140, the intelligent voice answering system can, after confirming the validity of the question, utilize the powerful large language model capabilities to provide accurate, timely, and coherent answers to users, greatly enhancing the user experience and service quality.

[0109] S150. Convert the answer text into speech to obtain the answer audio.

[0110] In this embodiment, the answer audio refers to the speech content converted from the answer text.

[0111] Specifically, a third-party TTS algorithm service that supports anthropomorphic voices is used to convert the answer text into answer audio for output.

[0112] To ensure that the final output speech is both natural and emotional, the system recommends using TTS (Text-to-Speech) algorithm services provided by third parties, especially those that can provide "anthropomorphic" voices. Such services are usually specially designed to simulate the voice characteristics close to real human voices, including but not limited to tone, intonation, rhythm, etc., so that the speech generated by the machine sounds more cordial and real.

[0113] Considering the requirements of instantaneity and coherence in user interaction, the TTS service also adopts a streaming processing method. This means that the answer text is not converted into speech all at once, but is gradually converted as the text content is understood. The advantage of this is that it can start playing while understanding, reducing the user's waiting time and maintaining the coherence of the answer.

[0114] By using advanced speech synthesis technology, the TTS service can accurately capture the emotional color and tone changes in the text, and then produce a highly realistic speech effect. This helps to enhance the realism of the conversation and makes the user feel like communicating with a human customer service representative.

[0115] In addition to basic speech synthesis, some high-end TTS services also allow customization of the gender, age, accent, and even emotion of the speech to better meet the needs in different application scenarios. For example, a more lively children's voice can be selected in children's educational software; while in enterprise-level customer service, a mature and steady adult voice may be preferred.

[0116] S160. Output the answer audio.

[0117] In this embodiment, at this stage, the system will send the already converted speech data to the user's terminal device, such as a smartphone, tablet computer, or smart speaker, etc., and play it through the speaker to respond to the user's inquiry.

[0118] Due to the adoption of an efficient TTS algorithm and a streaming processing mechanism, the entire process of converting text to speech is almost instantaneously completed, and the user hardly feels any delay, greatly enhancing the fluency of the interaction experience. Considering the diverse devices used by modern users, high-quality TTS services usually also have good cross-platform compatibility to ensure consistent high-quality speech output whether on iOS, Android, or other operating systems.

[0119] In summary, by implementing steps S150 and S160, the intelligent Q&A system not only achieves efficient conversion from text to speech, but more importantly, by introducing anthropomorphic voices and other advanced functions, it greatly improves the communication effect with users, making the entire interaction process more user-friendly and attractive.

[0120] The method of this embodiment can address the challenges of voice interaction in scenarios with multiple people asking questions simultaneously and complex background noise environments (such as cocktail party scenarios). It integrates multi-modal data processing, advanced voice separation technology, automatic speech recognition (ASR), semantic evaluation and generation capabilities of large language models (LLMs), and text-to-speech (TTS) synthesis technology, aiming to provide accurate and efficient voice Q&A services.

[0121] The integrated camera and multi-array microphones are used to collect on-site visual images and audio signals, laying the foundation for speaker separation, speech recognition, and semantic understanding. The images captured by the camera are used to assist in determining the identity of the main speaker, while the multi-array microphones ensure the accurate determination of the direction of the sound source. Deep learning algorithms are employed to analyze the multi-modal data to effectively separate the voice of the target speaker. The specific process includes extracting face detection information, lip shape information, and the orientation of the main speaker from the video stream; extracting noisy multi-channel speech mixtures and the direction of the target speaker from the audio stream as inputs, and finally outputting an audio mask of the target speaker to separate the clear target speech waveform. The separated audio stream is real-time converted into text content through the ASR algorithm service for further processing. The large language model (LLM) is used to semantically evaluate each transcribed text to determine whether it constitutes a complete and meaningful question. Only questions confirmed to be addressed to the system will trigger the answering mechanism, thereby improving the relevance and accuracy of the Q&A interaction. For valid questions, the large language model service is called to streamingly generate detailed and relevant answers. This process is dynamic, ensuring the timeliness and coherence of the response. Finally, the generated answer text is synthesized into natural and fluent speech using the TTS algorithm service to respond to the user's query, making the interaction more user-friendly.

[0122] Combining face recognition and lip shape detection, precisely lock in the main speaker, reducing interference from background noise and other speakers. Using the large language model to analyze the text content to ensure the relevance and accuracy of the answer. Providing clear and natural voice responses, reducing waiting time, and enhancing user satisfaction. Effectively filtering non-target voices, especially in multi-person conversation scenarios, improving the reliability and stability of the system.

[0123] The method of this embodiment is applicable to a variety of complex voice interaction scenarios, such as:

[0124] Smart Home and Internet of Things (IoT): Achieve precise control and collaborative work among devices, enhancing the level of home intelligence. Specifically, users can interact with the smart home system through natural language to control various devices. The system identifies the speaker through a camera and uses lip detection and audio separation technologies to ensure that when multiple people speak simultaneously, the system only responds to the correct commands. The intelligent system can integrate multiple devices in the home to achieve automated control. For example, when the user says "I'm home", the system automatically unlocks the door, adjusts the temperature, plays music, etc., providing comprehensive services.

[0125] Intelligent Customer Service and Virtual Assistants: Ensure that each user's instructions are correctly recognized and processed in customer service and virtual assistant applications. In intelligent customer service or virtual assistant applications, the system can effectively handle the situation where multiple users ask questions simultaneously, ensuring that each user's instructions are correctly recognized and responded to. For example, in a large shopping mall or exhibition, the system can determine the main speaker through face recognition and lip detection to provide personalized services. Combined with large language models, the system can more accurately understand user questions and generate natural and fluent answers, improving the user experience.

[0126] Conference and Speech Occasions: Serve as a real-time Q&A assistant to help the host or speaker interact with the audience and support multilingual communication. In conferences or speeches, the system can serve as a real-time Q&A assistant to help the host or speaker interact with the audience. Through lip detection and audio separation technologies, the system can accurately capture audience questions, avoid background noise interference, and ensure that each question is responded to in a timely manner. Combined with large model ASR (Automatic Speech Recognition) and TTS (Text-to-Speech) technologies, the system can translate and answer questions in different languages in real time, suitable for international conferences or internal communication within multinational enterprises, enhancing the efficiency of cross-cultural communication.

[0127] Education and Training: Provide personalized teaching guidance, simulate real work scenarios, and help trainees practice operations. In the field of education, the system can act as an intelligent teacher and provide personalized guidance according to students' questions. Through lip detection and speech recognition, the system can evaluate students' learning progress in real time and adjust teaching content to improve learning effects. In vocational skills training, the system can simulate the actual working environment to help trainees conduct practical training. For example, in medical training, the system can simulate patients to have natural conversations with trainees to help them practice diagnosis and treatment.

[0128] Public Facilities and City Services: As a smart tour guide or citizen service platform, it answers inquiries and provides guidance. In public places such as museums and exhibition halls, the system can act as a smart tour guide, offering personalized explanations and guidance. Through face recognition and lip detection, the system can accurately respond to tourists' questions, enhancing the visiting experience. In urban management, the system can serve as a citizen service platform, answering policy inquiries, handling affairs, etc. Through accurate speech recognition and semantic understanding, the system can quickly and effectively respond to citizens' questions, improving the efficiency and transparency of city services.

[0129] In-vehicle Voice Assistant: Assists the driver in safe driving while handling the needs of multiple passengers. In the in-vehicle system, the voice assistant can help the driver with operations such as navigation, answering calls, playing music, etc. Using lip detection and audio separation technologies, the system can accurately recognize the driver's commands in a noisy in-vehicle environment, avoiding misoperations and ensuring driving safety. The system can also handle the commands of multiple passengers simultaneously, ensuring that the needs of each passenger can be quickly responded to, improving the riding experience.

[0130] Remote Healthcare and Health Monitoring: Assists doctors in consultations and reminds users of health management matters. In remote healthcare, the system can act as a smart assistant to help doctors communicate with patients. Through lip detection and speech recognition, the system accurately records the patient's symptoms and transmits the information to the doctor, assisting in the generation of diagnosis and treatment suggestions. As a personal health management assistant, the system can remind users to take medicine on time, exercise, and answer health questions. Through emotional interaction technology, the system can also provide psychological support and companionship for the elderly and patients with chronic diseases.

[0131] Entertainment and Media: Interacts with audiences as a virtual idol or streamer, or acts as an NPC to talk with players in games and virtual reality environments. In the entertainment industry, the system can interact with audiences as a virtual idol or streamer. Through lip synchronization and speech recognition, virtual characters can respond to audiences' questions in real time, enhancing interactivity and sense of participation. In virtual reality and games, the system acts as an NPC or tour guide, capable of having natural conversations with players, providing a more immersive experience, and enhancing the sense of interaction. Through lip synchronization and speech recognition, the system creates a more realistic virtual environment.

[0132] In summary, the method of this embodiment significantly improves the accuracy of voice interaction and the user experience in complex environments by integrating multiple advanced technologies, and provides new possibilities for the development of future intelligent voice assistants.

[0133] The above voice answering method based on multi-modal speech recognition and semantic understanding, through

[0134] Figure 5It is a schematic block diagram of a voice question-answering system 300 provided by an embodiment of the present invention based on multimodal speech recognition and semantic understanding. As Figure 5 shown, corresponding to the above voice question-answering method based on multimodal speech recognition and semantic understanding, the present invention also provides a voice question-answering system 300 based on multimodal speech recognition and semantic understanding. The voice question-answering system 300 based on multimodal speech recognition and semantic understanding includes units for executing the above voice question-answering method based on multimodal speech recognition and semantic understanding, and this system can be configured in a server. Specifically, please refer to Figure 5 , the voice question-answering system 300 based on multimodal speech recognition and semantic understanding includes a data acquisition unit 301, a separation and conversion unit 302, an evaluation unit 303, an answer generation unit 304, a conversion unit 305, and an output unit 306.

[0135] The data acquisition unit 301 is used to acquire on-site visual images and audio signals collected by an integrated camera and a multi-array microphone to obtain multimodal data; the separation and conversion unit 305302 is used to perform audio separation and conversion of the target speaker on the multimodal data to obtain text content; the evaluation unit 303 is used to perform semantic understanding and evaluation of the validity of the question on the text content to obtain an evaluation result; the answer generation unit 304 is used to generate a corresponding answer text according to the text content when the evaluation result indicates that the text content belongs to a valid question; the conversion unit 305 is used to convert the answer text into speech to obtain answer audio; the output unit 306 is used to output the answer audio.

[0136] In one embodiment, as Figure 6 shown, the separation and conversion unit 302 includes: a determination subunit 3021, a separation subunit 3022, and a conversion subunit 3023.

[0137] The determination subunit 3021 is used to perform face detection, lip information extraction, and determination of the orientation of the target speaker on the multimodal data to obtain a determination result; the separation subunit 3022 is used to separate the target speech waveform from the multimodal data according to the determination result; the conversion subunit 3023 is used to convert the target speech waveform into text content.

[0138] In one embodiment, as Figure 7 shown, the determination subunit 3021 includes a face recognition module 30211, a lip recognition module 3012, and an orientation determination module 3013.

[0139] The face recognition module 30211 is used to process each frame of on-site visual image in the multi-modal data, recognize the face therein, and calculate the position of the face relative to the center of the picture; the lip shape recognition module 3012 is used to obtain the specific position of the lip shape from the face by applying a face semantic segmentation model; the orientation determination module 3013 is used to calculate the angle deviating from the center line of the picture based on the position information of the face relative to the center of the picture to determine the orientation of the target speaker; wherein, the determination result includes the information of the face, the specific position of the lip shape, and the orientation of the target speaker.

[0140] In one embodiment, the separation sub-unit 3022 is used to input the determination result and the audio signal into a deep learning fusion model to generate an audio mask, and separate the audio of the target speaker according to the audio mask to obtain a target speech waveform.

[0141] In one embodiment, the conversion sub-unit 3023 is used to convert the target speech waveform into text content by using an ASR algorithm service.

[0142] In one embodiment, the evaluation unit 303 is used to input the text content into a trained large language model to evaluate whether the text content constitutes a valid question to obtain an evaluation result.

[0143] In one embodiment, the answer generation unit 304 is used to, when the evaluation result is that the text content belongs to a valid question, call a large language model according to the text content to generate an answer to obtain an answer text.

[0144] In one embodiment, the conversion unit 305 is used to convert the answer text into an answer audio output by using a third-party TTS algorithm service that supports anthropomorphic voices.

[0145] It should be noted that those skilled in the art can clearly understand that the specific implementation processes of the above-mentioned speech question-answering system 300 based on multi-modal speech recognition and semantic understanding and each unit can refer to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and conciseness of description, they will not be elaborated here.

[0146] The above-mentioned speech question-answering system 300 based on multi-modal speech recognition and semantic understanding can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 8 shown.

[0147] Please refer to Figure 8 , Figure 8It is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 may be a server. Among them, the server may be an independent server or a server cluster composed of multiple servers.

[0148] Refer to Figure 8 , the computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory may include a non-volatile storage medium 503 and an internal memory 504.

[0149] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions. When the program instructions are executed, the processor 502 can be made to execute a voice question-and-answer method based on multi-modal speech recognition and semantic understanding.

[0150] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0151] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be made to execute a voice question-and-answer method based on multi-modal speech recognition and semantic understanding.

[0152] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 8 the structure shown in

[0153] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.

[0154] Obtain the on-site visual images and audio signals collected by the integrated camera and multi-array microphone to obtain multi-modal data; perform audio separation and conversion on the multi-modal data for the target speaker to obtain text content; perform semantic understanding and evaluation of the validity of the question on the text content to obtain an evaluation result; when the evaluation result is that the text content belongs to a valid question, generate a corresponding answer text according to the text content; convert the answer text into speech to obtain an answer audio; output the answer audio.

[0155] In one embodiment, when the processor 502 implements the step of performing audio separation and conversion of the target speaker on the multi-modal data to obtain text content, the specific implementation is as follows:

[0156] Perform face detection, lip information extraction, and determination of the orientation of the target speaker on the multi-modal data to obtain a determination result; separate the target speech waveform from the multi-modal data according to the determination result; convert the target speech waveform into text content.

[0157] In one embodiment, when the processor 502 implements the step of performing face detection, lip information extraction, and determination of the orientation of the target speaker on the multi-modal data to obtain a determination result, the specific implementation is as follows:

[0158] Process each frame of the on-site visual image in the multi-modal data, identify the face therein, and calculate the position of the face relative to the center of the picture; obtain the specific position of the lips from the face by applying a face semantic segmentation model; calculate the angle deviating from the center line of the picture based on the position information of the face relative to the center of the picture to determine the orientation of the target speaker;

[0159] Wherein, the determination result includes information of the face, the specific position of the lips, and the orientation of the target speaker.

[0160] In one embodiment, when the processor 502 implements the step of separating the target speech waveform from the multi-modal data according to the determination result, the specific implementation is as follows:

[0161] Input the determination result and the audio signal into a deep learning fusion model to generate an audio mask, and separate the audio of the target speaker according to the audio mask to obtain a target speech waveform.

[0162] In one embodiment, when the processor 502 implements the step of converting the target speech waveform into text content, the specific implementation is as follows:

[0163] Use the ASR algorithm service to convert the target speech waveform into text content.

[0164] In one embodiment, when the processor 502 implements the step of performing semantic understanding and evaluation of the validity of the question on the text content to obtain an evaluation result, the specific implementation is as follows:

[0165] Input the text content into a trained large language model to evaluate whether the text content constitutes a valid query to obtain an evaluation result.

[0166] In one embodiment, when the processor 502 implements the step of generating a corresponding answer text according to the text content when the evaluation result is that the text content belongs to a valid question, the specific implementation is as follows:

[0167] When the evaluation result is that the text content belongs to a valid question, call a large language model according to the text content to generate an answer to obtain the answer text.

[0168] In one embodiment, when the processor 502 implements the step of converting the answer text into speech to obtain the answer audio, the specific implementation is as follows:

[0169] Use a third-party TTS algorithm service that supports anthropomorphic voices to convert the answer text into answer audio for output.

[0170] It should be understood that in the embodiments of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0171] Those of ordinary skill in the art can understand that all or part of the processes of the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0172] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the following steps:

[0173] Obtain the on-site visual images and audio signals collected by the integrated camera and multi-array microphone to obtain multi-modal data; perform audio separation and conversion of the target speaker on the multi-modal data to obtain text content; perform semantic understanding and evaluation of the effectiveness of the question on the text content to obtain an evaluation result; when the evaluation result indicates that the text content belongs to a valid question, generate a corresponding answer text according to the text content; convert the answer text into speech to obtain an answer audio; output the answer audio.

[0174] In one embodiment, when the processor executes the computer program to implement the step of performing audio separation and conversion of the target speaker on the multi-modal data to obtain text content, the specific implementation is as follows:

[0175] Perform face detection, lip information extraction, and determination of the orientation of the target speaker on the multi-modal data to obtain a determination result; separate the target speech waveform from the multi-modal data according to the determination result; convert the target speech waveform into text content.

[0176] In one embodiment, when the processor executes the computer program to implement the step of performing face detection, lip information extraction, and determination of the orientation of the target speaker on the multi-modal data to obtain a determination result, the specific implementation is as follows:

[0177] Process each frame of the on-site visual image in the multi-modal data, identify the face therein, and calculate the position of the face relative to the center of the picture; obtain the specific position of the lips from the face by applying a face semantic segmentation model; calculate the angle deviating from the center line of the picture based on the position information of the face relative to the center of the picture to determine the orientation of the target speaker;

[0178] Wherein, the determination result includes the information of the face, the specific position of the lips, and the orientation of the target speaker.

[0179] In one embodiment, when the processor executes the computer program to implement the step of separating the target speech waveform from the multi-modal data according to the determination result, the specific implementation is as follows:

[0180] Input the determination result and the audio signal into a deep learning fusion model for audio mask generation, and separate the audio of the target speaker according to the audio mask to obtain the target speech waveform.

[0181] In one embodiment, when the processor executes the computer program to implement the step of converting the target speech waveform into text content, the specific implementation is as follows:

[0182] Convert the target speech waveform into text content using the ASR algorithm service.

[0183] In one embodiment, when the processor executes the computer program to implement the step of evaluating the semantic understanding and question validity of the text content to obtain an evaluation result, the specific implementation is as follows:

[0184] Input the text content into the trained large language model to evaluate whether the text content constitutes a valid query to obtain an evaluation result.

[0185] In one embodiment, when the processor executes the computer program to implement the step of generating a corresponding answer text according to the text content when the evaluation result indicates that the text content belongs to a valid question, the specific implementation is as follows:

[0186] When the evaluation result indicates that the text content belongs to a valid question, call the large language model according to the text content to generate an answer text.

[0187] In one embodiment, when the processor executes the computer program to implement the step of converting the answer text into speech to obtain an answer audio, the specific implementation is as follows:

[0188] Use a third-party TTS algorithm service that supports anthropomorphic voices to convert the answer text into an answer audio output.

[0189] The storage medium can be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.

[0190] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0191] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0192] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the system embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0193] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.

[0194] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A voice question-answering method based on multimodal speech recognition and semantic understanding, characterized in that: include: Acquire live visual images and audio signals collected by integrated cameras and multi-array microphones to obtain multimodal data; Separating and converting the audio of a target speaker from the multimodal data to obtain text content; Evaluate the semantic understanding and question validity of the text content to obtain an evaluation result; When the evaluation result shows that the text content is a valid question, generating a corresponding answer text according to the text content; Convert the answer text into speech to obtain answer audio; The answer audio is output.

2. The voice question-answering method based on multimodal speech recognition and semantic understanding according to claim 1, characterized in that: The step of separating and converting the audio of a target speaker from the multimodal data to obtain text content includes: Performing face detection, lip shape information extraction, and position determination of a target speaker on the multimodal data to obtain a determination result; separating a target speech waveform from the multimodal data according to the determination result; The target speech waveform is converted into text content.

3. The voice question-answering method based on multimodal speech recognition and semantic understanding according to claim 2, characterized in that: The performing face detection, lip shape information extraction and target speaker position determination on the multimodal data to obtain a determination result includes: Processing each frame of the on-site visual image in the multimodal data, identifying a human face therein, and calculating a position of the human face relative to the center of the picture; Obtaining the specific position of the lip shape from the face by applying a face semantic segmentation model; Based on the position information of the face relative to the center of the picture, the angle of deviation from the center line of the picture is calculated to determine the direction of the target speaker; The determination result includes information of the face, specific position of the lip shape and the position of the target speaker.

4. The voice question-answering method based on multimodal speech recognition and semantic understanding according to claim 2, characterized in that: The step of separating a target speech waveform from the multimodal data according to the determination result includes: The determination result and the audio signal are input into a deep learning fusion model to generate an audio mask, and the audio of the target speaker is separated according to the audio mask to obtain a target speech waveform.

5. The voice question-answering method based on multimodal speech recognition and semantic understanding according to claim 2, characterized in that: The step of converting the target speech waveform into text content comprises: The target speech waveform is converted into text content using an ASR algorithm service.

6. The voice question-answering method based on multimodal speech recognition and semantic understanding according to claim 1, characterized in that: The evaluation of semantic understanding of the text content and question validity to obtain an evaluation result includes: The text content is input into the trained large language model to evaluate whether the text content constitutes a valid query, so as to obtain an evaluation result.

7. The voice question-answering method based on multimodal speech recognition and semantic understanding according to claim 6, characterized in that: When the evaluation result is that the text content belongs to a valid question, generating a corresponding answer text according to the text content includes: When the evaluation result shows that the text content is a valid question, a large language model is called to generate an answer based on the text content to obtain an answer text.

8. The voice question-answering method based on multimodal speech recognition and semantic understanding according to claim 1, characterized in that: The step of converting the answer text into speech to obtain the answer audio comprises: A third-party TTS algorithm service that supports anthropomorphic voice is used to convert the answer text into an answer audio output.

9. A voice question-answering system based on multimodal speech recognition and semantic understanding, characterized in that: include: A data acquisition unit, used to acquire on-site visual images and audio signals collected by an integrated camera and a multi-array microphone to obtain multimodal data; A separation and conversion unit, configured to separate and convert the audio of a target speaker from the multimodal data to obtain text content; An evaluation unit, used to evaluate the semantic understanding of the text content and the effectiveness of the question to obtain an evaluation result; An answer generating unit, configured to generate a corresponding answer text according to the text content when the evaluation result is that the text content belongs to a valid question; A conversion unit, used for converting the answer text into speech to obtain an answer audio; An output unit is used to output the answer audio.

10. The voice question-answering system based on multimodal speech recognition and semantic understanding according to claim 9, characterized in that: The separation conversion unit comprises: A determination subunit, configured to perform face detection, lip shape information extraction, and position determination of a target speaker on the multimodal data to obtain a determination result; a separation subunit, configured to separate a target speech waveform from the multimodal data according to the determination result; The conversion subunit is used to convert the target speech waveform into text content.

Citation Information

Cited By

  • Simulation digital human real-time intelligent voice interaction system and method based on vision and large model

    CN120998199A

  • A visual and large model-based simulated digital human real-time intelligent voice interaction system and method thereof

    CN120998199B

  • Intelligent response interaction method and device based on voice input

    CN121331138A