Voice control methods, devices, vehicles, electronic devices and storage media

CN116825100BActive Publication Date: 2026-08-14APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

因此,语音识别系统的识别能力会在一定程度上影响设备或应用的控制效果

Benefits of technology

[0021]根据本公开的技术,基于针对第一设备的第一语音,得到语音识别文本;基于预设的多个语义识别模型,对语音识别文本进行语义识别,得到第一语义识别结果集;基于第一设备的状态信息以及第一语义识别结果集中各语义识别结果的置信度,在第一语义识别结果集中确定目标语义结果;基于所述目标语义结果对应的语音指令,控制所述第一设备。从而,采用本公开的技术,能够提高语音的语义识别准确率,进而提供准确的指令来控制设备。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116825100B_ABST
    Figure CN116825100B_ABST
Patent Text Reader

Abstract

This disclosure provides a voice control method, device, electronic device, vehicle, and storage medium, relating to the field of artificial intelligence technology, particularly natural language processing and the Internet of Vehicles (IoV). The specific implementation scheme is as follows: based on a first voice address to a first device, speech-recognized text is obtained; based on multiple preset semantic recognition models, semantic recognition is performed on the speech-recognized text to obtain a first semantic recognition result set; based on the state information of the first device and the confidence level of each semantic recognition result in the first semantic recognition result set, a target semantic result is determined in the first semantic recognition result set; based on the voice command corresponding to the target semantic result, the first device is controlled. Using the technical solution of this disclosure can improve the accuracy of semantic recognition, thereby accurately controlling the device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to the fields of natural language processing and vehicle networking, specifically to a voice control method, device, vehicle, electronic device, and storage medium. Background Technology

[0002] With the continuous development of voice recognition technology, it is becoming increasingly common to integrate voice recognition systems into hardware devices. For example, in vehicles, voice recognition systems allow users to control navigation applications via voice commands while the vehicle is in motion. Similarly, voice recognition systems integrated into touchscreens allow users to use their voice to interact with controls displayed on the screen instead of their hands. Therefore, the recognition capability of a voice recognition system will, to some extent, affect the control effectiveness of the device or application. Summary of the Invention

[0003] This disclosure provides a voice control method, apparatus, vehicle, electronic device, and storage medium.

[0004] According to one aspect of this disclosure, a voice control method is provided, comprising:

[0005] Based on the first speech received from the first device, the speech-recognized text is obtained;

[0006] Based on multiple preset semantic recognition models, semantic recognition is performed on the speech recognition text to obtain a first semantic recognition result set;

[0007] Based on the status information of the first device and the confidence level of each semantic recognition result in the first semantic recognition result set, the target semantic result is determined in the first semantic recognition result set;

[0008] The first device is controlled based on the voice command corresponding to the target semantic result.

[0009] According to another aspect of this disclosure, a voice control device is provided, comprising:

[0010] The speech recognition module is used to obtain speech-recognized text based on the first speech received from the first device.

[0011] The semantic recognition module is used to perform semantic recognition on the speech recognition text based on multiple preset semantic recognition models to obtain a first semantic recognition result set;

[0012] The target semantic determination module is used to determine the target semantic result in the first semantic recognition result set based on the state information of the first device and the confidence level of each semantic recognition result in the first semantic recognition result set.

[0013] The voice control module is used to control the first device based on the voice commands corresponding to the target semantic result.

[0014] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0015] At least one processor; and

[0016] The memory is communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the voice control methods in the embodiments of this disclosure.

[0018] According to another aspect of this disclosure, a vehicle is provided that includes any of the electronic devices described in the embodiments of this disclosure.

[0019] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the voice control methods according to embodiments of this disclosure.

[0020] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the voice control methods according to embodiments of this disclosure.

[0021] According to the technology disclosed herein, speech-recognized text is obtained based on a first speech received from a first device; semantic recognition is performed on the speech-recognized text based on multiple preset semantic recognition models to obtain a first semantic recognition result set; a target semantic result is determined in the first semantic recognition result set based on the state information of the first device and the confidence level of each semantic recognition result in the first semantic recognition result set; and the first device is controlled based on the voice command corresponding to the target semantic result. Therefore, by employing the technology disclosed herein, the accuracy of speech semantic recognition can be improved, thereby providing accurate commands to control the device.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0023] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0024] Figure 1 This is a flowchart of a voice control method according to an embodiment of the present disclosure;

[0025] Figure 2 This is an architectural block diagram of a voice system according to an embodiment of the present disclosure;

[0026] Figure 3 This is a flowchart of a speech recognition method according to another embodiment of this disclosure;

[0027] Figure 4 This is a flowchart of the data flow for online speech recognition according to an embodiment of the present disclosure;

[0028] Figure 5 This is a flowchart of a dialogue interaction according to an embodiment of the present disclosure;

[0029] Figure 6 This is a timing diagram of a conversational chat scenario according to an embodiment of this disclosure;

[0030] Figure 7 This is a structural block diagram of a voice control device according to an embodiment of the present disclosure;

[0031] Figure 8 This is a structural block diagram of a voice control device according to another embodiment of the present disclosure;

[0032] Figure 9 This is a block diagram of an electronic device used to implement the voice control method of the embodiments of this disclosure. Detailed Implementation

[0033] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0034] Figure 1 This is a flowchart of a voice control method according to an embodiment of the present disclosure. The method can be applied to an electronic device or a server. The electronic device can be an in-vehicle device, a smartphone, a computer, a smart wearable device, etc. In some embodiments, the electronic device receives voice data and provides it to a server, which then recognizes the data and returns the recognition result to the electronic device.

[0035] like Figure 1 As shown, the voice control method may include:

[0036] S110, based on the first speech for the first device, obtain speech-recognized text;

[0037] S120, based on multiple preset semantic recognition models, performs semantic recognition on the speech recognition text to obtain the first semantic recognition result set;

[0038] S130, based on the status information of the first device and the confidence level of each semantic recognition result in the first semantic recognition result set, determine the target semantic result in the first semantic recognition result set;

[0039] S140 controls the first device based on the voice command corresponding to the target semantic result.

[0040] For example, the first device may include in-vehicle equipment, smartphones, computers, smart wearable devices, etc.

[0041] For example, the first speech is recognized using Automatic Speech Recognition (ASR) technology to obtain speech-recognized text.

[0042] For example, a semantic recognition model may include an online semantic recognition model and an offline semantic recognition model. An offline semantic recognition model may include a text matching model and a pre-trained neural network model. The text matching model works by performing text matching on the speech recognition text within a pre-registered set of instruction texts based on pre-set text matching rules, obtaining the target instruction text, i.e., the semantic recognition result.

[0043] For example, semantic recognition of speech recognition text based on multiple preset semantic recognition models to obtain a first semantic recognition result set may include: performing semantic recognition of speech recognition text based on each of the multiple preset semantic recognition models to obtain a first semantic recognition result set.

[0044] For example, semantic recognition of speech recognition text based on multiple preset semantic recognition models to obtain a first semantic recognition result set may include: performing semantic recognition on speech recognition text model by model based on the execution order of each of the multiple preset semantic recognition models to obtain the first semantic recognition result set.

[0045] For example, based on multiple preset semantic recognition models, semantic recognition is performed on the speech-recognized text to obtain a first semantic recognition result set. This may include: if the semantic recognition model is an online semantic recognition model in a cloud server, sending the speech-recognized text to the cloud server, receiving the semantic recognition results returned by the cloud server, and adding the semantic recognition results to the first semantic recognition result set. The cloud server uses the online semantic recognition model to recognize the speech-recognized text to obtain the semantic recognition results.

[0046] In practical applications, cloud servers can simultaneously utilize multiple online semantic recognition models to recognize speech-to-text, thereby obtaining multiple online semantic recognition results.

[0047] Locally, multiple offline semantic recognition models are used to perform semantic recognition on speech recognition text, resulting in multiple offline semantic recognition results.

[0048] For example, the first semantic recognition result set includes one or more semantic recognition results. Semantic recognition results can also be called semantic understanding results. Semantic recognition results include semantic information obtained by semantic understanding of the speech recognition text. The semantic understanding method can be text matching or machine learning. Semantic recognition results can include offline semantic recognition results and online semantic recognition results.

[0049] For example, the status information of the first device may include the operating status of the first device, the working status of the components in the first device, and the working status of the applications in the first device. Components may be hardware devices that are compatible with the first device, such as sensors, speakers, network cards, cameras, etc. If the first device is a vehicle, the status information of the first device may include whether the vehicle is moving, whether the windows are open, whether the doors are open, and whether the multimedia application has been started.

[0050] For example, using the state information of the first device, the semantic recognition results that match the state information are filtered in the first semantic recognition result set to obtain at least one semantic recognition result; based on the confidence level of each semantic recognition result in the at least one semantic recognition result, the target semantic recognition result is determined in the at least one semantic recognition result.

[0051] According to the above implementation method, semantic recognition is performed on the speech recognition text based on multiple preset semantic recognition models to obtain one or more semantic recognition results. Then, the semantic recognition results are filtered by using the device's status information and confidence level to obtain a more accurate target semantic recognition result for understanding the first speech, thereby accurately controlling the device.

[0052] In some embodiments, the status information of the first device may include the signal strength of the network signal.

[0053] In one exemplary implementation, determining a target semantic result in the first semantic recognition result set based on the state information of the first device and the confidence level of each semantic recognition result in the first semantic recognition result set includes: when the signal strength of the network signal of the first device is lower than a set threshold, filtering semantic recognition results in the first semantic recognition result set based on the type information of the semantic recognition model corresponding to each semantic recognition result in the first semantic recognition result set to obtain a second semantic recognition result set; and determining the target semantic recognition result in the second semantic recognition result set based on the first confidence level of each semantic recognition result in the second semantic recognition result set.

[0054] For example, the semantic recognition model corresponding to the semantic recognition result is the model that outputs the semantic recognition result.

[0055] For example, semantic recognition models can be categorized as online models and offline models. They can also be categorized as text matching models and neural network models. Furthermore, they can be categorized as models running in a computational digital signal processor (cDSP) and models running in a non-computational digital signal processor (cDSP).

[0056] For example, when the signal strength of the network signal of the first device is lower than a set threshold, the semantic recognition results obtained based on the offline semantic recognition model are filtered from the first semantic recognition result set to obtain the second semantic recognition result set.

[0057] For example, when the signal strength of the network information of the first device is lower than a set threshold, the semantic recognition results obtained based on the semantic recognition model running in the computational digital signal processor (cDSP) are filtered from the first semantic recognition result set to obtain the second semantic recognition result set.

[0058] For example, the second semantic recognition result set includes one or more semantic recognition results.

[0059] In practical applications, using online semantic recognition models to perform semantic recognition on speech-to-text can improve recognition accuracy. Online semantic recognition models can employ GPT (Generative Pre-Training) models.

[0060] By using offline semantic recognition models to perform semantic recognition on speech-recognition text, such as text matching, semantic recognition results can be obtained quickly.

[0061] By utilizing the semantic recognition model running in a computational digital signal processor (cDSP), semantic recognition of speech-recognized text can be performed, and semantic recognition results can be obtained quickly.

[0062] For example, when the second semantic recognition result set includes multiple semantic recognition results, the target semantic recognition result is determined among these multiple semantic recognition results based on the first confidence level of each result in the multiple semantic recognition results.

[0063] For example, the first confidence level can be the confidence level provided by the semantic recognition model when outputting semantic recognition results for speech recognition text. The first confidence level can also be the confidence level obtained by evaluating the semantic recognition results according to a set confidence level evaluation method.

[0064] For example, a semantic recognition result with a first confidence level greater than a set threshold can be selected as the target semantic recognition result.

[0065] According to the above implementation method, when the network signal strength is weak, filtering the semantic recognition results based on the type of the model outputting the semantic recognition results can improve the speed of semantic recognition and avoid speech recognition lag in weak network conditions. Furthermore, further filtering the semantic recognition results based on their confidence level can further improve the accuracy of the semantic recognition results.

[0066] In the above embodiments, the confidence level can be corrected, and the corrected confidence level can be used to filter the semantic recognition results.

[0067] In one exemplary implementation, determining a target semantic recognition result in the second semantic recognition result set based on the first confidence level of each semantic recognition result in the second semantic recognition result set includes: for each semantic recognition result in the second semantic recognition result set, determining a first value of the semantic recognition result based on the type information of the voice command corresponding to the semantic recognition result and the type information of the semantic recognition model corresponding to the semantic recognition result; adjusting the first confidence level of the semantic recognition result based on the first value of the semantic recognition result to obtain a second confidence level of the semantic recognition result; and determining the target semantic recognition result in the second semantic recognition result set based on the second confidence level of each semantic recognition result in the second semantic recognition result set.

[0068] For example, the types of voice commands include commands for applications such as weather, stocks, news, and navigation. They can also be commands for controlling hardware devices, or commands for touching a specific control.

[0069] For example, the types of semantic recognition models can be referred to in the foregoing embodiments.

[0070] In practical applications, for commands that require real-time information feedback to the user, such as weather, stock quotes, news, or navigation, a semantic recognition model with higher accuracy is generally required. For commands that control hardware devices or controls, which do not require real-time feedback, a faster semantic recognition model is generally required. Therefore, by adjusting the confidence level of the semantic recognition results output by the model based on the command type and model type, the reliability of the confidence level can be improved, thereby increasing the accuracy of the selected semantic recognition results.

[0071] For example, the first value can be used as the weight value of the first confidence level, and the second confidence level can be obtained by multiplying the first value by the first confidence level.

[0072] For example, when the second semantic recognition result set includes multiple semantic recognition results, the semantic recognition results with a second confidence level greater than a set threshold can be filtered to obtain the target semantic recognition result.

[0073] According to the above implementation method, by using instruction type and model type to correct the confidence of the semantic recognition result output by the model, the reliability of the confidence can be improved, thereby improving the accuracy of the selected semantic recognition result.

[0074] In one exemplary implementation, for each semantic recognition result in the second semantic recognition result set, based on the type information of the voice command corresponding to the semantic recognition result and the type information of the semantic recognition model corresponding to the semantic recognition result, a first value of the semantic recognition result is determined, including:

[0075] The second value is determined based on the type information of the voice command corresponding to the semantic recognition result;

[0076] The third value is determined based on the type information of the semantic recognition model corresponding to the semantic recognition result;

[0077] Based on the second and third values, the first value of the semantic recognition result is determined.

[0078] For example, a first function can be pre-defined to calculate the type information of the voice command, resulting in a second value. The first function characterizes the relationship between the degree to which the voice command requires real-time feedback and the type of the voice command. Thus, the second value characterizes the relationship between the degree to which the voice command corresponding to the semantic recognition result requires real-time feedback and the type of the voice command.

[0079] For example, a second function can be pre-defined to calculate the type information of the semantic recognition model, resulting in three numerical values. The second function characterizes the recognition accuracy and speed of the semantic recognition model, and their relationship with the type of the semantic recognition model. Thus, the third numerical value characterizes the recognition accuracy and speed of the semantic recognition model corresponding to the semantic recognition result, and their relationship with the type of the semantic recognition model.

[0080] For example, the first value is obtained based on the product between the second and third values.

[0081] For example, a parameter is pre-defined to obtain a first value based on the sum of the product between the second and third values ​​and the parameter.

[0082] According to the above embodiment, the first value is determined by using a second value representing the relationship between the degree to which the voice command corresponding to the semantic recognition result needs to provide real-time information and the type of the voice command, and a second value representing the relationship between the recognition accuracy and recognition speed of the semantic recognition model corresponding to the semantic recognition result and the type of the semantic recognition model. Thus, the first value can accurately represent the relationship between the real-time performance, accuracy, and speed of the semantic recognition result, thereby accurately correcting the confidence level of the semantic recognition result and improving the selection accuracy of the semantic recognition result.

[0083] In one exemplary embodiment, obtaining speech-recognized text based on a first speech for a first device includes: if the source location of the first speech for the first device matches the speech recognition mode of the first device, verifying the voiceprint information of the first speech based on the first voiceprint information set of the first device to obtain a first verification result; if the first verification result is successful, performing speech recognition on the first speech to obtain speech-recognized text.

[0084] For example, if the first device is a vehicle, the source location of the voice can include a location inside the vehicle and a location outside the vehicle. The location inside the vehicle can further include the driver's seat, the front passenger seat, the driver's rear seat, and the front passenger's rear seat, etc. The location outside the vehicle can further include the front of the vehicle, the rear of the vehicle, the left side of the vehicle, and the right side of the vehicle, etc.

[0085] For example, if the first device is a vehicle, the voice recognition mode may include an in-vehicle voice recognition mode and an out-of-vehicle voice recognition mode. Different voice recognition modes are used to recognize voices from different sources, thereby improving recognition accuracy.

[0086] For example, for in-vehicle voice recognition mode, the voice recognition text can be matched using the command text set corresponding to that mode to obtain the target command text. For out-of-vehicle voice mode, the voice recognition text can be matched using the command text set corresponding to that mode to obtain the target command text.

[0087] According to the above implementation method, when the source location of the voice matches the current voice recognition mode of the device, the voice is verified by voiceprint, and the voice is recognized only after the verification is successful. This can avoid misrecognition of voice commands.

[0088] In one exemplary embodiment, the method may further include: determining a first voiceprint information set based on voiceprint information of a second voice for a first device; wherein the second voice includes voices identified within a preset historical time period.

[0089] For example, the historical time period can be a period from the first time to the current time, where the first time is before the current time, and the duration between the first time and the current time can be preset as needed.

[0090] In practical applications, if the voiceprint of the previously recognized speech matches the voiceprint of the current speech within a short period of time, then the speech recognition process continues.

[0091] According to the above implementation method, using the voiceprint of historically recognized speech to verify the voiceprint of the current speech can improve the accuracy of the verification.

[0092] In one exemplary embodiment, the above method may further include: if the first verification result is a verification failure, verifying the voiceprint information of the first speech based on the second voiceprint information set pre-registered by the first device to obtain a second verification result; if the second verification result is a verification success, performing speech recognition on the first speech to obtain speech recognition text.

[0093] For example, a user of the first device inputs a registration voice into the first device, and the first device registers based on the voiceprint information in the registration voice to obtain a second voiceprint information set.

[0094] In practical applications, if the voiceprint of the previously recognized speech does not match the voiceprint of the current speech within a short period of time, the pre-registered voiceprint information in the device will continue to be used to verify the voiceprint, which can prevent the user's voice from being intercepted.

[0095] In one exemplary embodiment, the above method can be applied to a voice tool built based on a software development kit, wherein the first voice comes from a first voice application, and the first device is controlled based on the voice command corresponding to the target semantic recognition result, including: the voice tool provides the target semantic recognition result to the first voice application, wherein the target semantic recognition result is used by the voice application to determine the voice command, and the voice command is used by the first voice application to control the first device.

[0096] For example, a voice tool is set up in the first device, which is built based on a software development kit (SDK).

[0097] For example, the first voice application may be set in the first device. The first device may include various applications, such as voice applications, multimedia applications, and system-adaptive applications.

[0098] According to the above implementation method, by using the software development kit (SDK), different types of semantic recognition models can be called to provide speech and semantic recognition functions for different speech applications, thereby improving the compatibility of speech and semantic recognition.

[0099] Figure 2 This is an architectural block diagram of a voice system according to an embodiment of the present disclosure.

[0100] like Figure 2 As shown, the following will take the first device as an in-vehicle device as an example to describe the working principle of the voice recognition system in the in-vehicle device, as follows:

[0101] The speech recognition system includes a terminal and a cloud. The terminal includes an application module and a speech SDK.

[0102] The application modules include voice application apps, business application apps, and system adaptation application apps. Business applications and voice applications are upper-layer applications for in-vehicle devices, such as map applications, music applications, mini-program applications, and video applications. System applications are lower-layer applications for in-vehicle devices, such as map service applications, vehicle control service applications, multimedia service applications, and telephone service applications.

[0103] Upon receiving a voice command, the in-vehicle voice application provides the voice command to the voice SDK. The voice SDK then calls the voice recognition engine, offline semantic recognition models, and online semantic recognition models to process the voice command, obtain the semantic recognition result, and return it to the in-vehicle voice application. The in-vehicle voice application then invokes business applications and / or system applications based on the voice command corresponding to the semantic recognition result.

[0104] Therefore, in the aforementioned speech recognition system, due to the embedding of the speech SDK, each application can operate independently, allowing for more flexible control of the device using voice.

[0105] The Voice SDK includes an Automatic Speech Recognition (ASR) engine and a Natural Language Understanding (NLU) engine. The Voice SDK module supports three working modes: online, hybrid online / offline, and fully offline (all-device-side). In the all-device-side mode, ASR and NLU algorithms, based on the NPU (Neural-Network Processing Unit) computing power, run on the computational digital signal processor (cDSP). This allows the all-device-side to achieve the same voice interaction capabilities in offline or weak network environments as in network-connected environments, and even with faster voice interaction response speeds. Examples include faster ASR recognition, text display, and semantic understanding result return. It's important to note that the Voice SDK interface is consistent with the interfaces of voice applications or other applications. This interface maintains forward compatibility as the Voice SDK continues to be upgraded.

[0106] In this example, the voice SDK is decoupled from the voice app, making model upgrades within the voice SDK easier. Furthermore, it achieves the same voice interaction capabilities in offline or weak network environments as in environments with a network connection.

[0107] The cloud platform can include an online speech recognition module (ASR), an online speech synthesis module (TTS, Text-to-Speech), and an online semantic understanding module (NLU). The cloud primarily focuses on semantic understanding, providing streaming text and image services. In multi-turn voice interaction scenarios, the cloud's advantages in processing speech are more pronounced. It can provide corresponding text and image information based on the specific scenario. For example, it can generate reasonable travel guides based on travel requirements, including attractions, restaurants, and routes, thus solving users' travel problems. In chat scenarios, the speech recognition system can freely converse and chat with users, much like a human.

[0108] According to the above implementation method, a voice SDK is set up in the voice recognition system. The voice SDK integrates online and offline engines for wake-up recognition, TTS, and NLU, which can be built in the form of an SDK. The voice SDK calls wake-up and recognition instances through its interface layer. It processes the voice using the offline engine and also accesses cloud APIs for online voice processing. The voice application app receives the offline and online recognition results from the voice SDK and displays them streamed on the screen. Simultaneously, after arbitrating the semantic results online and offline, the voice application app calls various business apps and system-adapted apps to begin performing operations such as multimedia playback, vehicle control, and navigation.

[0109] The voice app first configures the wake-up words and voice recognition mode using a configuration file. After initialization, the voice app calls the recording to begin the conversation. If the voice SDK operates in full-duplex mode, it needs to send the vehicle's terminal status information (e.g., whether Bluetooth is enabled, whether navigation is active, etc.) to the voice SDK to facilitate better judgment in semantic understanding. Simultaneously, the voice assistant server manages the registration information of various business application apps. After obtaining the semantic results, the voice SDK provides them to the Dialogue Management (DM) layer. The dialogue management layer manages each round of conversation. For cases where the previous round was based on online semantic recognition results and the next round on offline semantic recognition results, a dialogue flow tracker is established to ensure the voice app makes the best decisions.

[0110] Figure 3 This is a flowchart of a speech recognition method according to another embodiment of the present disclosure.

[0111] like Figure 3 As shown, the method includes the following steps:

[0112] S301, the voice SDK collects the terminal status and conversation details before recognizing the spoken dialogue. After noise reduction, the voice SDK performs offline and online speech recognition processing on the recorded data.

[0113] The S302 voice SDK uses a VAD (Voice Active Detection) model to detect the start point of the conversational speech. Upon successful detection, the conversational speech is provided to both the cDSP and the cloud.

[0114] S303, the voice SDK calls the speech recognition model (ASR) and semantic understanding model (NLU) in the computational digital signal processor (cDSP) to recognize speech, obtaining an offline first speech recognition result (ASR1), a first text semantic understanding result (VTS1), a first model semantic understanding result (NLU1), and first discriminant information (R1). The first discriminant information (R1) matches the terminal state with either the first text semantic understanding result (VTS1) or the first model semantic understanding result (NLU1) to determine whether to return either the first text semantic understanding result (VTS1) or the first model semantic understanding result (NLU1). When the first discriminant information (R1) determines whether to return the first text semantic understanding result (VTS1) or the first model semantic understanding result (NLU1), the voice SDK can provide this data to execute step S308.

[0115] S304, the speech SDK calls the "See-It-Speak" model in the computational digital signal processor (cDSP) to recognize speech, obtaining an offline second text semantic understanding result VTS2, a second model semantic understanding result NLU2, and second discriminant information R2. The second discriminant information R2 is matched against the end state and the second text semantic understanding result VTS2 or the second model semantic understanding result NLU2 to determine whether to return either the second text semantic understanding result VTS2 or the second model semantic understanding result NLU2. When it is determined, based on the second discriminant information R2, to return either the second text semantic understanding result VTS2 or the second model semantic understanding result NLU2, the speech SDK can provide this data to execute step S307.

[0116] S305, the voice SDK calls the model in the cloud to recognize the speech, obtaining the online second speech recognition result ASR2, the third text semantic understanding result VTS3, and the third discriminant information R3. The third discriminant information R3 is matched with the terminal state and either the third text semantic understanding result VTS3 or the third model semantic understanding result NLU3 to determine whether to return either the third text semantic understanding result VTS3 or the third model semantic understanding result NLU3. If, based on the first discriminant information R3, it is determined that the third text semantic understanding result VTS3 or the third model semantic understanding result NLU3 should be returned, then the cloud returns these data to the voice SDK. Simultaneously, the cloud also returns the second speech recognition result ASR2 to the voice SDK.

[0117] S306, the Voice SDK receives the online second speech recognition result ASR2, the third text semantic understanding result VTS3, and the third model semantic understanding result NLU3. Then, the Voice SDK returns the second speech recognition result ASR2, the third text semantic understanding result VTS3, and the third model semantic understanding result NLU3 to the Voice App.

[0118] S307, the voice SDK determines whether the confidence level of the second text semantic understanding result VTS2 is greater than 0.8. If it is greater, the voice SDK provides the second text semantic understanding result VTS2 to the voice APP as the offline recognition result; otherwise, proceed to step S308.

[0119] S308, the voice SDK determines whether the semantic understanding result NLU1 of the first model belongs to the whitelist. If it does, the voice SDK provides the semantic understanding result NLU1 of the first model to the voice APP; otherwise, the voice SDK provides the semantic understanding result NLU2 of the second model to the voice APP.

[0120] Offline speech recognition and semantic understanding results are processed together by a computational digital signal processor (cDSP) and returned simultaneously, resulting in faster processing speeds. Online speech recognition and semantic understanding results are retrieved from an online GPT model, leading to more accurate results.

[0121] In the example above, the voice SDK can send both offline and online recognition results back to the voice app. The voice app then arbitrates the offline and online results according to their priorities. For example, for results with the same priority, the one that returns the result faster will be selected.

[0122] In the example above, embedding a voice SDK into the speech recognition system integrates offline and online engines, allowing the access party to choose as needed, resulting in a more flexible architecture. Furthermore, it enables the same voice interaction capabilities in weak network environments as in well-connected environments.

[0123] Figure 4 This is a flowchart of the data flow for online speech recognition according to an embodiment of the present disclosure.

[0124] like Figure 4 As shown, the cloud provides online AI (Artificial Intelligence) services, such as GPT services. GPT is used to perform semantic recognition on speech-recognition text and return the recognition results to the voice SDK. The voice SDK then returns the results to the client, i.e., the voice APP.

[0125] In this scenario, multiple terminals can request services from the cloud synchronously or asynchronously. In this case, a DCS (Distributed Control System) can be used to identify the intent of the request before requesting services from the online AI.

[0126] Figure 5 This is a flowchart of a dialogue interaction according to an embodiment of the present disclosure.

[0127] like Figure 5 As shown, after the voice app begins recognizing the corresponding speech, the voice SDK sends a request to the cloud-based bot (BOT). The cloud-based bot parses the intent of the request and sends it to an online AI service, such as the GPT model. The online AI service processes the speech to obtain semantic understanding results. After performing operations such as filtering for pornographic content, checking for sensitive words, authentication, and limiting traffic on the semantic understanding results, the processed semantic understanding results are sent back to the voice SDK.

[0128] The voice SDK can request voice services from the online AI service synchronously or asynchronously. When requesting a service, the request data can include the following information: dialogue identifier, whether it is a multi-turn dialogue, the sequence number of the current speech in the multi-turn dialogue, whether the current speech is the last sentence in the multi-turn dialogue, the dialogue result, whether the dialogue result contains risks, and whether the dialogue needs to be terminated.

[0129] In this process, after the voice app calls the voice SDK to obtain the online dialogue results, it then calls the TTS SDK through the voice SDK to synthesize natural synthesized speech TTS.

[0130] Figure 6 This is a timing diagram of a conversational chat scenario according to an embodiment of this disclosure.

[0131] like Figure 6As shown, when initiating a small talk session, the voice app sends the voice message to the ASR (Automatic Speech Recognition) for speech recognition. The ASR then sends the request to the DCS (Distributed Control System), which in turn sends the intent to initiate the small talk session to the BOT (Bot). The BOT returns a multi-turn synthesized text-to-speech (TTS) message indicating the start of the immersive small talk session. The BOT can subsequently request online AI model services to enter the immersive small talk scenario. When ending a small talk session, the voice app sends the voice message to the ASR for speech recognition. The ASR then sends the request to the DCS, which in turn sends the intent to end the small talk session to the BOT. The BOT returns a synthesized text-to-speech (TTS) message indicating the end of the small talk session.

[0132] Based on the above implementation method, a unified voice service architecture integrating the client and cloud can be provided, offering greater flexibility and controllability. This architecture, tailored to specific business needs, more accurately grasps the user's true intent, facilitating easier integration for business users. It accelerates API call speed through caching and streaming, supporting multi-turn conversations. Furthermore, by applying a voice SDK, redundant model development is avoided, facilitating the rapid deployment of new models.

[0133] Figure 7 This is a structural block diagram of a voice control device according to an embodiment of the present disclosure.

[0134] like Figure 7 As shown, the voice control device may include:

[0135] The speech recognition module 710 is used to obtain speech-recognized text based on a first speech directed to the first device;

[0136] The semantic recognition module 720 is used to perform semantic recognition on the speech recognition text based on multiple preset semantic recognition models to obtain a first semantic recognition result set;

[0137] The target semantic determination module 730 is used to determine the target semantic result in the first semantic recognition result set based on the state information of the first device and the confidence level of each semantic recognition result in the first semantic recognition result set.

[0138] The voice control module 740 is used to control the first device based on the voice command corresponding to the target semantic result.

[0139] Figure 8 This is a structural block diagram of a voice control device according to another embodiment of the present disclosure.

[0140] like Figure 7 and Figure 8 As shown, Figure 8 The speech recognition module 810, semantic recognition module 820, target semantic determination module 830, and voice control module 840 in the text correspond to the speech recognition module 810, semantic recognition module 820, target semantic determination module 830, and voice control module 840 respectively. Figure 7The speech recognition module 710, semantic recognition module 720, target semantic determination module 730 and voice control module 740 have the same function or role, and will not be described in detail here.

[0141] In one exemplary implementation, such as Figure 8 As shown, the target semantic determination module 830 includes:

[0142] The first semantic filtering unit 831 is used to filter semantic recognition results in the first semantic recognition result set based on the type information of the semantic recognition model corresponding to each semantic recognition result in the first semantic recognition result set when the signal strength of the network signal of the first device is lower than a set threshold, thereby obtaining a second semantic recognition result set.

[0143] The second semantic filtering unit 832 is used to determine the target semantic recognition result in the second semantic recognition result set based on the first confidence level of each semantic recognition result in the second semantic recognition result set.

[0144] In one exemplary embodiment, the second semantic filtering unit 832 is specifically used for:

[0145] For each semantic recognition result in the second semantic recognition result set, a first value of the semantic recognition result is determined based on the type information of the voice command corresponding to the semantic recognition result and the type information of the semantic recognition model corresponding to the semantic recognition result;

[0146] Based on the first value of the semantic recognition result, the first confidence level of the semantic recognition result is adjusted to obtain the second confidence level of the semantic recognition result;

[0147] Based on the second confidence level of each semantic recognition result in the second semantic recognition result set, the target semantic recognition result is determined in the second semantic recognition result set.

[0148] In one exemplary implementation, determining a first value of the semantic recognition result for each semantic recognition result in the second semantic recognition result set, based on the type information of the voice command corresponding to the semantic recognition result and the type information of the semantic recognition model corresponding to the semantic recognition result, includes:

[0149] Based on the type information of the voice command corresponding to the semantic recognition result, a second value is determined;

[0150] Based on the type information of the semantic recognition model corresponding to the semantic recognition result, a third numerical value is determined;

[0151] Based on the second value and the third value, a first value of the semantic recognition result is determined.

[0152] In one exemplary implementation, such as Figure 8 As shown, the speech recognition module 810 includes:

[0153] The first verification unit 811 is used to verify the voiceprint information of the first speech based on the first voiceprint information set of the first device when the source location of the first speech of the first device matches the speech recognition mode of the first device, and obtain a first verification result.

[0154] The first speech recognition unit 812 is used to perform speech recognition on the first speech when the first verification result is successful, and obtain the speech recognition text.

[0155] In one exemplary implementation, such as Figure 8 As shown, the above-mentioned speech recognition module 810 also includes:

[0156] The voiceprint set determination unit 813 is used to determine the first voiceprint information set based on the voiceprint information of the second voice for the first device; wherein the second voice includes voices identified within a preset historical time period.

[0157] In one exemplary implementation, such as Figure 8 As shown, the above-mentioned speech recognition module 810 also includes:

[0158] The second verification unit 814 is used to verify the voiceprint information of the first voice based on the second voiceprint information set pre-registered by the first device when the first verification result is a verification failure, and to obtain a second verification result.

[0159] The second speech recognition unit 815 is used to perform speech recognition on the first speech when the second verification result is successful, so as to obtain the speech recognition text.

[0160] In one exemplary embodiment, the device is applied to a voice tool built based on a software development kit, wherein the first voice comes from a first voice application, and the voice control module 840 is specifically used for: the voice tool providing the target semantic recognition result to the first voice application, wherein the target semantic recognition result is used by the first voice application to determine a voice command, and the voice command is used by the voice application to control the first device.

[0161] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0162] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0163] According to embodiments of this disclosure, this disclosure also provides an electronic device, a vehicle, a readable storage medium, and a computer program product.

[0164] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0165] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.

[0166] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0167] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as a voice control method. For example, in some embodiments, a voice control method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of a voice control method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform a voice control method by any other suitable means (e.g., by means of firmware).

[0168] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0169] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0170] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0171] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0172] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0173] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0174] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0175] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A voice control method, comprising: Based on the first speech received from the first device, the speech-recognized text is obtained; Based on multiple preset semantic recognition models, semantic recognition is performed on the speech recognition text to obtain a first semantic recognition result set; Based on the status information of the first device and the confidence level of each semantic recognition result in the first semantic recognition result set, the target semantic result is determined in the first semantic recognition result set; Based on the voice command corresponding to the target semantic result, control the first device; The step of determining the target semantic result in the first semantic recognition result set based on the state information of the first device and the confidence level of each semantic recognition result in the first semantic recognition result set includes: When the signal strength of the network signal of the first device is lower than a set threshold, based on the type information of the semantic recognition model corresponding to each semantic recognition result in the first semantic recognition result set, the semantic recognition results are filtered in the first semantic recognition result set to obtain the second semantic recognition result set. For each semantic recognition result in the second semantic recognition result set, a first value of the semantic recognition result is determined based on the type information of the voice command corresponding to the semantic recognition result and the type information of the semantic recognition model corresponding to the semantic recognition result; Based on the first value of the semantic recognition result, the first confidence level of the semantic recognition result is adjusted to obtain the second confidence level of the semantic recognition result; Based on the second confidence level of each semantic recognition result in the second semantic recognition result set, the target semantic recognition result is determined as the target semantic result in the second semantic recognition result set.

2. The method according to claim 1, wherein, The step of determining a first value for each semantic recognition result in the second semantic recognition result set, based on the type information of the voice command corresponding to the semantic recognition result and the type information of the semantic recognition model corresponding to the semantic recognition result, includes: Based on the type information of the voice command corresponding to the semantic recognition result, a second value is determined; Based on the type information of the semantic recognition model corresponding to the semantic recognition result, a third numerical value is determined; Based on the second value and the third value, a first value of the semantic recognition result is determined.

3. The method according to claim 1 or 2, wherein, The process of obtaining speech-recognized text based on the first speech received from the first device includes: If the source location of the first voice from the first device matches the voice recognition mode of the first device, the voiceprint information of the first voice is verified based on the first voiceprint information set of the first device to obtain a first verification result. If the first verification result is successful, the first speech is subjected to speech recognition to obtain the speech recognition text.

4. The method according to claim 3, further comprising: Based on the voiceprint information of the second voice of the first device, the first voiceprint information set is determined; wherein the second voice includes voices identified within a preset historical time period.

5. The method according to claim 3, further comprising: If the first verification result is a verification failure, the voiceprint information of the first voice is verified based on the second voiceprint information set pre-registered by the first device to obtain a second verification result. If the second verification result is successful, the first speech is subjected to speech recognition to obtain the speech recognition text.

6. The method according to claim 1, wherein, The method is applied to a voice tool built based on a software development kit, wherein the first voice comes from a first voice application, and the step of controlling the first device based on the voice command corresponding to the target semantic recognition result includes: The voice tool provides the target semantic recognition result to the first voice application, wherein the target semantic recognition result is used by the first voice application to determine a voice command, and the voice command is used by the first voice application to control the first device.

7. A voice control device, comprising: The speech recognition module is used to obtain speech-recognized text based on the first speech received from the first device. The semantic recognition module is used to perform semantic recognition on the speech recognition text based on multiple preset semantic recognition models to obtain a first semantic recognition result set; The target semantic determination module is used to determine the target semantic result in the first semantic recognition result set based on the state information of the first device and the confidence level of each semantic recognition result in the first semantic recognition result set. The voice control module is used to control the first device based on the voice commands corresponding to the target semantic result; The target semantic determination module includes: The first semantic filtering unit is used to filter semantic recognition results in the first semantic recognition result set based on the type information of the semantic recognition model corresponding to each semantic recognition result in the first semantic recognition result set when the signal strength of the network signal of the first device is lower than a set threshold, so as to obtain a second semantic recognition result set. The second semantic filtering unit is configured to, for each semantic recognition result in the second semantic recognition result set, determine a first value of the semantic recognition result based on the type information of the voice command corresponding to the semantic recognition result and the type information of the semantic recognition model corresponding to the semantic recognition result; adjust the first confidence level of the semantic recognition result based on the first value of the semantic recognition result to obtain a second confidence level of the semantic recognition result; and determine a target semantic recognition result as the target semantic result in the second semantic recognition result set based on the second confidence level of each semantic recognition result in the second semantic recognition result set.

8. The apparatus according to claim 7, wherein, The step of determining a first value for each semantic recognition result in the second semantic recognition result set, based on the type information of the voice command corresponding to the semantic recognition result and the type information of the semantic recognition model corresponding to the semantic recognition result, includes: Based on the type information of the voice command corresponding to the semantic recognition result, a second value is determined; Based on the type information of the semantic recognition model corresponding to the semantic recognition result, a third numerical value is determined; Based on the second value and the third value, a first value of the semantic recognition result is determined.

9. The apparatus according to claim 7 or 8, wherein, The speech recognition module includes: The first verification unit is used to verify the voiceprint information of the first speech based on the first voiceprint information set of the first device when the source location of the first speech of the first device matches the speech recognition mode of the first device, and obtain a first verification result. The first speech recognition unit is configured to perform speech recognition on the first speech when the first verification result is successful, and obtain the speech recognition text.

10. The apparatus according to claim 9, further comprising: The voiceprint set determination unit is used to determine the first voiceprint information set based on the voiceprint information of the second voice for the first device; wherein the second voice includes voices identified within a preset historical time period.

11. The apparatus according to claim 9, further comprising: The second verification unit is used to verify the voiceprint information of the first speech based on the second voiceprint information set pre-registered by the first device when the first verification result is a verification failure, and to obtain a second verification result. The second speech recognition unit is used to perform speech recognition on the first speech when the second verification result is successful, so as to obtain the speech recognition text.

12. The apparatus according to claim 7, wherein, The device is applied to a voice tool built based on a software development kit, wherein the first voice comes from a first voice application, and the voice control module is specifically used for: The voice tool provides the target semantic recognition result to the first voice application, wherein the target semantic recognition result is used by the first voice application to determine a voice command, and the voice command is used by the first voice application to control the first device.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

14. A vehicle comprising the electronic device of claim 13.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.

16. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Voice interaction method, device and system, storage medium and processor

    CN110491383A

  • Electronic equipment unlocking method and device

    CN114444042A

  • Voice interaction method, voice interaction device, vehicle and readable storage medium

    CN115410579A