Voice processing method, apparatus and system
By correcting the language confidence level in the speech recognition system and combining user and scene features, the problem of insufficient speech recognition capability in multilingual environments is solved, achieving higher speech recognition accuracy and language recognition precision.
Patent Information
- Application Number
- CN202180001914.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-22
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2041-06-22
AI Technical Summary
Existing speech recognition technologies have a low recognition capability when recognizing different languages, and are prone to fluctuations, especially in multilingual environments.
By acquiring the user's input voice information, the language confidence is adjusted based on user characteristics and scene characteristics. By utilizing information such as historical language records, user-specified languages, and environmental characteristics, the preset weights are adjusted to improve the accuracy of language recognition.
It improves the accuracy of speech recognition and the precision of language identification, and enhances the adaptability of the speech processing system in multilingual environments.
Smart Images

Figure CN113597641B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a speech processing method, device and system. BACKGROUND
[0002] With the development of computer technology, speech recognition technology has been more and more widely applied. In addition, with the deepening of globalization, there are often scenes of mixed office and life of people using different languages. For example, on international flights, passengers are often people from different countries or regions, and the languages they use are not the same. Or, for example, in a country like Singapore, English is the first foreign language for most locals, and at the same time, because there are many Chinese and Chinese descendants, they usually use Chinese as their daily communication language. Therefore, both English and Chinese may appear in communication. Therefore, in order to cope with this situation, a technology capable of recognizing different language speech (for example, both Chinese speech and English speech can be recognized) has appeared.
[0003] However, the recognition result of the technology sometimes fluctuates, that is, the speech recognition ability is low, and there is still room for improvement in this regard. SUMMARY
[0004] The present application provides a speech processing method, device and system capable of improving speech recognition ability, for improving the accuracy of speech recognition.
[0005] The first aspect of the present application relates to a speech processing method, comprising the following contents: obtaining input speech information of a user; determining a plurality of first confidence degrees corresponding to the input speech information according to the input speech information, the plurality of first confidence degrees respectively corresponding to a plurality of languages; correcting the plurality of first confidence degrees to a plurality of second confidence degrees according to a user feature of the user; and determining a language of the input speech information according to the plurality of second confidence degrees.
[0006] By using the speech processing method as described above, the plurality of first confidence degrees is corrected to the plurality of second confidence degrees according to the user feature of the user, and the language of the input speech information is determined according to the plurality of second confidence degrees, that is, the language of the input speech information of the user is determined on the basis of considering the user feature, so that the language recognition precision can be improved, and the speech recognition ability can be improved.
[0007] As one possible implementation manner of the first aspect of the present application, the plurality of first confidence degrees is corrected to the plurality of second confidence degrees according to the user feature of the user, and specifically can include: when the plurality of first confidence degrees is less than a first threshold, the plurality of first confidence degrees is corrected to the plurality of second confidence degrees according to the user feature.
[0008] When the plurality of first confidence degrees are less than the first threshold, it is difficult to determine the language of the input speech information according to the first confidence degrees. When the plurality of first confidence degrees are corrected into the plurality of second confidence degrees according to the user feature at this time, the language of the input speech information is determined according to the second confidence degrees, which can improve the recognition accuracy of the language and improve the speech recognition capability.
[0009] The user feature can include one or more of a historical language record and a user specified language.
[0010] In this way, the first recognition confidence degree is corrected according to the historical language record and / or the user specified language of the user, and the language of the input speech is determined on this basis, so that the language recognition capability can be improved.
[0011] The historical language record of the user here refers to a record of the language to which the speech input by the user before the input of the input speech above belongs. The user specified language here refers to the type of the system language set by the user, and there can be only one user specified language, or there can be multiple user specified languages (i.e., the user has set multiple system languages).
[0012] As one possible implementation manner of the first aspect of the present application, the historical language record and the user specified language are obtained according to the voiceprint feature of the input speech information.
[0013] In this way, the historical language record or the user specified language is queried according to the voiceprint of the input speech information, which can avoid the language misrecognition caused by the misrecognition of the user (speaker) (the non-speaker is recognized as the speaker) compared with the way of querying according to the face information, iris, etc. In addition, in this way, the voiceprint can be obtained according to the input speech information, and the user image needs to be obtained in the way of querying according to the face information, iris, etc., so that the device required in the way of querying according to the voiceprint is less and the processing is faster.
[0014] As one possible implementation manner of the first aspect of the present application, the plurality of first confidence degrees are determined by the plurality of initial confidence degrees and the plurality of preset weights. The speech processing method can further include the following content: updating the plurality of preset weights according to the plurality of second confidence degrees.
[0015] In this way, the language recognition accuracy in the subsequent processing period can be improved.
[0016] As one possible implementation manner of the first aspect of the present application, the plurality of preset weights are updated according to the plurality of second confidence degrees, specifically including: when there is a second confidence degree greater than the first threshold in the plurality of second confidence degrees, the plurality of preset weights are updated according to the plurality of second confidence degrees.
[0017] Therefore, when the second confidence is greater than the first threshold, the result of the current processing cycle is more reliable, and the preset weights are updated according to the second confidence, so that the language recognition accuracy in the subsequent processing cycle is improved more reliably.
[0018] As a possible implementation of the first aspect of the application, the semantic of the input speech information is determined according to the input speech information and the language of the input speech information.
[0019] In this way, the language recognition accuracy and the semantic understanding accuracy can be improved.
[0020] As a possible implementation of the first aspect of the application, the plurality of languages are pre-set.
[0021] As a possible implementation of the first aspect of the application, the plurality of first confidences are determined according to the plurality of initial confidences and the plurality of preset weights; and the speech processing method further comprises: before obtaining the input speech information of the user, setting the plurality of preset weights according to the scene feature.
[0022] In this way, the preset weights are set according to the scene feature, so that the preset weights can be adapted to different scenes, and the language recognition result can be obtained with the most suitable preset weights for the scene, thereby improving the language recognition ability and the speech recognition ability.
[0023] As a possible implementation of the first aspect of the application, the scene feature comprises an environment feature and / or an audio collector feature.
[0024] As a possible implementation of the first aspect of the application, the environment feature comprises one or more of an environment signal-to-noise ratio, power DC / AC information, or an environment vibration amplitude, and the audio collector feature comprises microphone arrangement information.
[0025] The environment signal-to-noise ratio, the power DC / AC information, the environment vibration amplitude, and the microphone arrangement information can all affect the language confidence, and therefore, the preset weights are adjusted according to these information, and the language recognition is performed on this basis, thereby improving the language recognition ability.
[0026] As a possible implementation of the first aspect of the application, the plurality of preset weights are set according to the scene feature, and specifically comprising: obtaining pre-collected first speech data and pre-recorded first language information of the first speech data; determining second speech data according to the first speech data and the scene feature; determining second language information of the second speech data according to the second speech data; and setting the plurality of preset weights according to the first language information and the second language information.
[0027] As a possible implementation manner of the first aspect of the present application, the second language information of the second voice data is determined according to the second voice data, specifically including: obtaining a plurality of test weight sets, each of the plurality of test weight sets including a plurality of test weights; determining a plurality of second language information according to the second voice data and the plurality of test weight sets, the plurality of second language information corresponding to the plurality of test weight sets respectively; and setting a plurality of preset weights according to the first language information and the second language information, specifically including: determining a plurality of accuracy rates of the plurality of second language information according to the first language information and the plurality of second language information; and setting the plurality of preset weights according to the test weight set corresponding to the second language information with the highest accuracy rate.
[0028] As a possible implementation manner of the first aspect of the present application, the plurality of preset weights are set, specifically including: setting the plurality of preset weights within a weight range.
[0029] As a possible implementation manner of the first aspect of the present application, the plurality of preset weights are updated, specifically including: updating the plurality of preset weights within a weight range.
[0030] If the preset weight exceeds the weight range, the recognition result will be unreliable, therefore, setting or updating the weight within the weight range can guarantee the accuracy of the recognition result as much as possible.
[0031] As a possible implementation manner of the first aspect of the present application, the weight range is determined as follows:
[0032] The plurality of test voice data sets and the first language information of the plurality of test voice data sets recorded in advance are obtained, each of the plurality of test voice data sets including a plurality of test voice data; the plurality of test weight sets are obtained, each of the plurality of test weight sets including a plurality of test weights; and the weight range is determined according to the plurality of test voice data sets, the first voice information and the plurality of test weight sets.
[0033] In this way, the weight range of the preset weight set of the multi-language is set by testing a large number of voice data sets, that is, the robustness range of the language recognition model is specified, so that the language recognition model works within this range, thereby ensuring the reliability of the language recognition result.
[0034] The second aspect of the present application provides a voice processing method, including the following contents: obtaining input voice information of a user; determining a plurality of third confidence degrees corresponding to the input voice information according to the input voice information, the plurality of third confidence degrees corresponding to a plurality of languages respectively; correcting the plurality of third confidence degrees to a plurality of fourth confidence degrees according to the scene characteristics; and determining the language of the input voice information according to the plurality of fourth confidence degrees.
[0035] According to the voice processing method, the plurality of first confidences are corrected into a plurality of second confidences according to the scene features, and the language of the input voice information is determined according to the plurality of second confidences, that is, the language of the input voice information of the user is determined on the basis of considering the scene features, so that the voice processing method can adapt to the actual scene as much as possible, the language recognition accuracy can be improved, and the voice recognition capability can be improved.
[0036] Here, the scene features can include environment features and / or audio collector features.
[0037] As a possible implementation manner of the second aspect of the application, the environment features include one or more of an environment signal-to-noise ratio, power DC-AC information, or an environment vibration amplitude, and the audio collector features include microphone arrangement information.
[0038] As a possible implementation manner of the second aspect of the application, the plurality of third confidences are corrected into a plurality of fourth confidences according to the scene features, and specifically include: setting a plurality of preset weights according to the scene features; and correcting the plurality of third confidences into the plurality of fourth confidences according to the plurality of preset weights.
[0039] As a possible implementation manner of the second aspect of the application, the plurality of preset weights are set according to the scene features, and specifically include: obtaining first voice data collected in advance and first language information of the first voice data recorded in advance; determining second voice data according to the first voice data and the scene features; determining second language information of the second voice data according to the second voice data; and setting the plurality of preset weights according to the first language information and the second language information.
[0040] As a possible implementation manner of the second aspect of the application, the second language information of the second voice data is determined according to the second voice data, and specifically includes: obtaining a plurality of test weight groups, each test weight group including a plurality of test weights; determining a plurality of second language information according to the second voice data and the plurality of test weight groups, the plurality of second language information corresponding to the plurality of test weight groups respectively; and setting the plurality of preset weights according to the first language information and the second language information, and specifically includes: determining a plurality of accuracy rates of the plurality of second language information according to the first language information and the plurality of second language information; and setting the plurality of preset weights according to a test weight group corresponding to the second language information with the highest accuracy rate.
[0041] The specific features of the second aspect of the application can be the same as or similar to those of the first aspect, and therefore the technical effects are basically the same, which will not be described here.
[0042] The third aspect of the present application provides a speech processing apparatus, comprising a processing module and a transceiver module, the transceiver module being configured to obtain input speech information of a user; the processing module being configured to determine a plurality of first confidence degrees corresponding to the input speech information according to the input speech information, the plurality of first confidence degrees corresponding to a plurality of languages respectively; and the processing module being further configured to correct the plurality of first confidence degrees into a plurality of second confidence degrees according to user features of the user, and determine a language of the input speech information according to the plurality of second confidence degrees.
[0043] As a possible implementation manner of the third aspect of the present application, the processing module is specifically configured to correct the plurality of first confidence degrees into the plurality of second confidence degrees according to the user features when the plurality of first confidence degrees are less than a first threshold.
[0044] As a possible implementation manner of the third aspect of the present application, the user features comprise one or more of historical language records and user specified languages.
[0045] As a possible implementation manner of the third aspect of the present application, the historical language records and the user specified languages are obtained according to a voiceprint feature of the input speech information.
[0046] As a possible implementation manner of the third aspect of the present application, the plurality of first confidence degrees are determined by a plurality of initial confidence degrees and a plurality of preset weights; and the processing module is further configured to update the plurality of preset weights according to the plurality of second confidence degrees.
[0047] As a possible implementation manner of the third aspect of the present application, the processing module is specifically configured to update the plurality of preset weights according to the plurality of second confidence degrees when there is a second confidence degree greater than the first threshold in the plurality of second confidence degrees.
[0048] As a possible implementation manner of the third aspect of the present application, the processing module is further configured to determine a semantic of the input speech information according to the input speech information and the language of the input speech information.
[0049] The plurality of languages can be preset.
[0050] As a possible implementation manner of the third aspect of the present application, the plurality of first confidence degrees are determined by a plurality of initial confidence degrees and a plurality of preset weights; and the processing module is further configured to set the plurality of preset weights according to scene features before obtaining the input speech information of the user.
[0051] The scene features can comprise environment features and / or audio collector features. The environment features can comprise one or more of environment signal-to-noise ratio, power AC / DC information or environment vibration amplitude, and the audio collector features can comprise microphone arrangement information.
[0052] As a possible implementation manner of the third aspect of the present application, the processing module is specifically configured to acquire the first voice data collected in advance and the first language information of the first voice data recorded in advance, determine the second voice data according to the first voice data and the scene feature, determine the second language information of the second voice data according to the second voice data, and set the plurality of preset weights according to the first language information and the second language information.
[0053] As a possible implementation manner of the third aspect of the present application, the processing module is specifically configured to acquire a plurality of test weight groups, each of the plurality of test weight groups including a plurality of test weights, determine a plurality of second language information according to the second voice data and the plurality of test weight groups, the plurality of second language information corresponding to the plurality of test weight groups respectively, determine a plurality of accuracy rates of the plurality of second language information according to the first language information and the plurality of second language information, and set the plurality of preset weights according to the test weight group corresponding to the second language information with the highest accuracy rate.
[0054] As a possible implementation manner of the third aspect of the present application, the processing module is specifically configured to set the plurality of preset weights within a weight range.
[0055] As a possible implementation manner of the third aspect of the present application, the processing module is specifically configured to set the plurality of preset weights within a weight range.
[0056] As a possible implementation manner of the third aspect of the present application, the weight range is determined in the following manner:
[0057] acquire a plurality of test voice data groups collected in advance and first language information of the plurality of test voice data groups recorded in advance, each of the plurality of test voice data groups including a plurality of test voice data; acquire a plurality of test weight groups, each of the plurality of test weight groups including a plurality of test weights; and determine a weight range according to the plurality of test voice data groups, the first voice information and the plurality of test weight groups.
[0058] The voice processing device of the third aspect can obtain the same technical effects as the voice processing method of the first aspect, which will not be described here again.
[0059] The fourth aspect of the present application provides a voice processing device, including a processing module and a transceiver module, the transceiver module being configured to acquire input voice information of a user; the processing module being configured to determine a plurality of third confidences corresponding to the input voice information according to the input voice information, the plurality of third confidences corresponding to a plurality of languages respectively, and further configured to correct the plurality of third confidences to a plurality of fourth confidences according to a scene feature, and determine a language of the input voice information according to the plurality of fourth confidences.
[0060] The scene feature can include an environmental feature and / or an audio collector feature.
[0061] The environmental features can include one or more of an environmental signal-to-noise ratio, power AC / DC information, or environmental vibration amplitude, and the audio collector features can include microphone arrangement information.
[0062] As a possible implementation manner of the fourth aspect, the processing module is specifically configured to set the plurality of preset weights according to the scene features, and correct the plurality of third confidences to the plurality of fourth confidences according to the plurality of preset weights.
[0063] As a possible implementation manner of the fourth aspect, the processing module is specifically configured to obtain pre-collected first voice data and pre-recorded first language information of the first voice data, determine second voice data according to the first voice data and the scene features, determine second language information of the second voice data according to the second voice data, and set the plurality of preset weights according to the first language information and the second language information.
[0064] As a possible implementation manner of the fourth aspect, the processing module is specifically configured to obtain a plurality of test weight groups, each test weight group including a plurality of test weights, determine a plurality of second language information according to the second voice data and the plurality of test weight groups, the plurality of second language information corresponding to the plurality of test weight groups respectively, determine a plurality of accuracy rates of the plurality of second language information according to the first language information and the plurality of second language information, and set the plurality of preset weights according to a test weight group corresponding to a second language information with a highest accuracy rate.
[0065] The speech processing apparatus of the fourth aspect can obtain the same technical effects as the speech processing method of the second aspect, which will not be described herein.
[0066] The fifth aspect of the present application provides a computing device including a processor and a memory, the memory storing computer program instructions, the computer program instructions causing the processor to execute any method described in the first aspect or the second aspect when executed by the processor.
[0067] The sixth aspect of the present application provides a computer-readable storage medium storing computer program instructions, the computer program instructions causing a computer to execute any method described in the first aspect or the second aspect when executed by the computer.
[0068] The seventh aspect of the present application provides a computer program product including computer program instructions, the computer program instructions causing a computer to execute any method described in the first aspect or the second aspect when executed by the computer.
[0069] The eighth aspect of the present application provides a system including the speech processing apparatus provided by any aspect or any possible implementation manner of the third aspect to the fourth aspect. BRIEF DESCRIPTION OF DRAWINGS
[0070] Figure 1 is a schematic diagram of one application scenario of the speech processing scheme provided by an embodiment of the present application;
[0071] Figure 2 is a schematic diagram of a speech processing system to which the speech processing scheme provided by an embodiment of the present application is applied;
[0072] Figure 3 is a flowchart of a speech processing method provided by an embodiment of the present application;
[0073] Figure 4 is a flowchart of a speech processing method provided by an embodiment of the present application;
[0074] Figure 5 is a schematic structural diagram of a speech processing apparatus provided by an embodiment of the present application;
[0075] Figure 6 is a flowchart for schematically illustrating a language recognition method provided by an embodiment of the present application;
[0076] Figure 7 is a structural schematic diagram of a language recognition apparatus provided by an embodiment of the present application;
[0077] Figure 8 is a flowchart for schematically illustrating a speech interaction method provided by an embodiment of the present application;
[0078] Figure 9 is a structural schematic diagram of a speech interaction system provided by an embodiment of the present application;
[0079] Figure 10 is a schematic diagram of one setting method of a weight range;
[0080] Figure 11 is a schematic diagram of a speech interaction system related to an embodiment of the present application;
[0081] Figure 12 is a flowchart of one flow of a speech interaction method related to an embodiment;
[0082] Figure 13 is a schematic diagram of an initialization method of a preset weight set provided by an embodiment of the present application;
[0083] Figure 14 is a schematic diagram of a part of a speech interaction process provided by an embodiment of the present application;
[0084] Figure 15 is a schematic diagram of a confidence correction method provided by an embodiment of the present application;
[0085] Figure 16 This is a schematic diagram illustrating yet another confidence level correction method provided in one embodiment of this application;
[0086] Figure 17 This is a schematic diagram illustrating an electronic control unit provided in one embodiment of this application.
[0087] It should be understood that the dimensions and shapes of the blocks in the above structural diagrams are for reference only and should not constitute an exclusive interpretation of the embodiments of the present invention. The relative positions and inclusion relationships between the blocks presented in the structural diagrams are only schematic representations of the structural relationships between the blocks, and are not intended to limit the physical connection methods of the embodiments of the present invention. Detailed Implementation
[0088] The technical solutions provided in this application will be further described below with reference to the accompanying drawings and embodiments. It should be understood that the system architecture and business scenarios provided in the embodiments of this application are mainly for illustrating possible implementations of the technical solutions of this application and should not be construed as the sole limitation on the technical solutions of this application. Those skilled in the art will recognize that the technical solutions provided in this application are equally applicable to similar technical problems as system architectures evolve and new business scenarios emerge.
[0089] It should be understood that the speech processing solutions provided in the embodiments of this application include speech processing methods, devices, and systems. Since these technical solutions solve problems based on the same or similar principles, some repetitive details may not be repeated in the following descriptions of specific embodiments. However, it should be considered that these specific embodiments have mutual references and can be combined with each other.
[0090] First refer to Figure 1 This describes an application scenario example of the voice processing solution provided in the embodiments of this application. Figure 1 The examples shown are scenarios applied to vehicles, specifically, such as Figure 1 As shown, the vehicle's infotainment system has a voice interaction function. It can receive voice commands from the driver, passengers, and other occupants through the microphone 212 (microphone array in this example) on the central control display screen 210. The system can then execute corresponding controls according to the voice commands (such as playing music, opening windows, turning on the air conditioning, or navigating). The system can also respond to the voice commands (provide feedback), for example, by displaying information on the central control display screen 210 or by emitting voice information through the speaker (not shown) on the central control display screen 210.
[0091] For example, since different occupants may ride in vehicle 200, they may issue voice commands in different languages, or even the same occupant may issue voice commands in different languages. Therefore, the vehicle's infotainment system should be able to handle voice commands in different languages. However, due to limitations in language recognition capabilities, the infotainment system may sometimes obtain incorrect language recognition results, resulting in the inability to recognize or incorrectly recognize the semantics of the voice command, and thus failing to respond correctly.
[0092] Specifically, as a language recognition solution, there is a technique that uses machine learning models for classification and recognition. However, machine learning models may learn some information that is irrelevant to the task, such as the environmental signal-to-noise ratio, audio acquisition device (sound sensor, microphone) characteristics, etc. This can lead to errors in the prediction results of machine learning models when these information changes in actual applications.
[0093] For example, in Figure 1 In the scenario shown, vehicle 200 is a convertible, and the ambient noise is relatively high (e.g., a medium-noise environment). Therefore, when the driver 300 issues the English voice command "Please play music," the vehicle's infotainment system may receive an incorrect language recognition result, thus failing to correctly recognize the voice command and respond appropriately. Additionally, if the microphone array type corresponding to the training sample data of the machine learning model is different from the microphone 212 of the vehicle's infotainment system, it may also cause the system to generate incorrect language recognition results and fail to correctly recognize the driver's voice command.
[0094] Therefore, embodiments of this application provide a speech processing method, apparatus, and system, which can improve the speech recognition capability of multilingual speech processing solutions.
[0095] The following describes a system architecture used in the speech processing methods, apparatus and systems provided in the embodiments of this application. Figure 2 This is a schematic diagram illustrating the architecture of a voice processing system used in the voice processing scheme provided in this application embodiment. For example... Figure 2 As shown, the voice processing system 180 includes a voice processing device 182, a sound sensor (microphone) 184, a speaker 186, and a display device 188.
[0096] The voice processing system 180 can be used as an in-vehicle system in smart vehicles. In addition, it can also be applied to smart homes, smart offices, smart robots, smart voice Q&A, smart voice analysis, real-time voice monitoring and analysis and other scenarios.
[0097] The sound sensor 184 is configured to acquire the input voice of the user, and the voice processing apparatus 182 obtains the input voice information of the user according to the sensor data of the sound sensor 184, processes the input voice information, and obtains the semantics of the input voice information. Then, the voice processing apparatus 182 performs corresponding control according to the semantics, for example, controls the output of the speaker 186 or the display apparatus 188. In addition, in addition to the speaker 186 and the display apparatus 188, the voice processing apparatus 182 can also be connected with other devices or mechanisms. For example, when the voice processing system 180 is applied in a car machine system, the voice processing apparatus 182 can also be connected with a window lifting system, an air conditioning system, etc., so as to realize the control of the window, the air conditioning system, etc.
[0098] The following refers to Figure 3 The voice processing method provided by an embodiment of the present application is described.
[0099] Figure 3 FIG. 1 is a flowchart of the voice processing method provided by an embodiment of the present application. The voice processing method can be executed by a vehicle, a vehicle-mounted device, a car machine, a vehicle-mounted computer, etc., or can be executed by a component such as a chip or a processor in the vehicle or the vehicle-mounted device. In addition, in addition to being applied in a vehicle, the voice processing method can also be applied in other scenes such as smart home or smart office. At this time, the voice processing method can be executed by a related device such as a control device or a processor involved in these scenes.
[0100] As shown in FIG. 1, the voice processing method includes the following contents: Figure 3
[0101] S1, obtaining the input voice information of the user. The input voice information of the user can be obtained according to the sensor data collected by the sound sensor, and can be directly obtained by using the sensor data or can be obtained by processing the sensor data. In addition, the time length of the input voice information is not particularly limited, and can correspond to a paragraph of the user or a sentence of the user. Furthermore, when processing the voice, the content spoken by the user can be divided into a plurality of input voice information, and the processing of S2-S4 described below is performed on the plurality of input voice information.
[0102] S2, determining a plurality of first confidence degrees corresponding to the input voice information according to the input voice information, and the plurality of first confidence degrees correspond to a plurality of languages respectively. Here, the plurality of languages can be pre-set. In addition, the confidence degree of the language means the probability that the input voice information belongs to the language. For example, when the plurality of first recognition confidence degrees obtained are {Chinese: 0.6; English: 0.4; Korean: 0; German: 0; Japanese: 0}, it means that the probability that the language of the input voice information is Chinese is 0.6, the probability that the language of the input voice information is English is 0.4, and the probability that the language of the input voice information is Korean, German or Japanese is 0.
[0103] In addition, different languages can refer to different language families, for example, Chinese and English belong to different languages, or refer to different small languages in the same language family, for example, Mandarin and Cantonese of Chinese also belong to different languages.
[0104] S3, correcting the plurality of first confidences to a plurality of second confidences according to user features. The user features herein are, for example, historical language records or user specified languages. The historical language records are the recognized languages of the input voice information of the user identified and recorded before the current processing period. The recognized language herein means the language to which the input voice information belongs determined by recognizing the input voice information. The user specified language refers to the type of system language set by the user, for example, according to the language commonly used by the user.
[0105] S4, determining the language of the input voice information according to the plurality of second confidences.
[0106] By using the voice processing method as above, the first confidence is corrected according to the user features, and the language of the input voice information is determined according to the second confidence obtained after the correction, so that the language of the input voice information can be determined more accurately, and the voice recognition capability can be improved.
[0107] As for the specific correction method, for example, the correction is performed according to the historical language records: assuming that the number of records of Chinese in the historical language records is relatively large, the confidence of Chinese in the plurality of first confidences obtained in the current processing period is increased, and the second confidences are obtained in this way, for example, according to the historical language records, the plurality of first confidences {Chinese: 0.6; English: 0.4; Korean: 0; German: 0; Japanese: 0} are corrected to {Chinese: 0.8; English: 0.2; Korean: 0; German: 0; Japanese: 0}.
[0108] Optionally, the plurality of first confidences can be corrected to a plurality of second confidences according to the user features when the plurality of first confidences is less than a first threshold. When the plurality of first confidences is less than the first threshold, it is difficult to determine the language of the input voice information according to the first confidences. At this time, when the plurality of first confidences is corrected to the plurality of second confidences according to the user features, the language of the input voice information is determined according to the second confidences, which can improve the recognition accuracy of the language and improve the voice recognition capability.
[0109] Optionally, the historical language records and the user specified language can be obtained by querying the voiceprint features of the input voice information. In this way, the historical language records and the user specified language can be easily obtained.
[0110] Optionally, the plurality of first confidences can be determined by a plurality of initial confidences and a plurality of preset weights, and at this time, the plurality of preset weights can be updated according to the plurality of second confidences.
[0111] Thus, according to the processing result of the current processing cycle, the preset weight is updated, so as to improve the language recognition accuracy of the subsequent processing cycle.
[0112] As a specific updating method, the updating can be performed when the second confidence values greater than the first threshold exist in the plurality of second confidence values.
[0113] When the second confidence values greater than the first threshold exist in the plurality of second confidence values, the language recognition result obtained according to the plurality of second confidence values is more reliable, and thus updating the preset weight according to the plurality of second confidence values can more reliably improve the language recognition accuracy of the subsequent processing cycle.
[0114] Optionally, after determining the language of the input voice information, the semantic of the input voice information can be determined according to the input voice information and the language of the input voice information.
[0115] In this way, the language of the input voice information can be more accurately determined, and thus the recognition accuracy of the semantic of the input voice information can be improved.
[0116] Optionally, the plurality of preset weights can be set according to the scene features before obtaining the input voice information of the user.
[0117] Since the plurality of preset weights are set according to the scene features before obtaining the input voice information of the user, the language recognition accuracy can be improved.
[0118] The scene features herein can include, for example, environmental features and / or audio collector features.
[0119] By setting the plurality of preset weights according to the environmental features and / or the audio collector features, the language recognition accuracy can be improved.
[0120] The environmental features herein can include one or more of environmental signal-to-noise ratio, power DC / AC information, or environmental vibration amplitude, and the audio collector features can include microphone arrangement information.
[0121] As a specific way of setting the plurality of preset weights, the following way can be adopted: obtaining first voice data collected in advance and first language information of the first voice data recorded in advance; determining second voice data according to the first voice data and the scene features; determining second language information of the second voice data according to the second voice data; and setting the plurality of preset weights according to the first language information and the second language information.
[0122] Further, the specific manner of determining the second language information of the second speech data according to the second speech data can be: obtaining a plurality of test weight sets, each of the plurality of test weight sets comprising a plurality of test weights; determining a plurality of second language information according to the second speech data and the plurality of test weight sets, the plurality of second language information corresponding to the plurality of test weight sets respectively; determining a plurality of accuracy rates of the plurality of second language information according to the first language information and the plurality of second language information; and setting a plurality of preset weights according to a test weight set corresponding to a second language information with a highest accuracy rate.
[0123] In addition, an adjustable range, i.e., a weight range, can be set for the plurality of preset weights, and the plurality of preset weights are set or updated within the weight range. If the preset weights exceed the weight range, the recognition result will be unreliable. Therefore, by setting the adjustable range, i.e., the weight range, the accuracy of the language recognition result can be improved.
[0124] Here, the weight range can be determined in the following manner: obtaining a plurality of test speech data sets and a first language information of the plurality of test speech data sets recorded in advance, each of the plurality of test speech data sets comprising a plurality of test speech data; obtaining a plurality of test weight sets, each of the plurality of test weight sets comprising a plurality of test weights; and determining the weight range according to the plurality of test speech data sets, the first speech information, and the plurality of test weight sets.
[0125] The following refers to Figure 4 The speech processing method provided by another embodiment of the present application is described. Figure 4 is a flowchart of the speech processing method provided by an embodiment of the present application. Similar to the above-described embodiments, the speech processing method of this embodiment can be executed by a vehicle, a vehicle-mounted device, a vehicle machine, a vehicle-mounted computer, etc., or can be executed by a component such as a chip or a processor in the vehicle or the vehicle-mounted device. In addition, some of the contents in this embodiment are the same as those in the above-described embodiments, and thus the description of these contents will not be repeated.
[0126] As Figure 4 shown, the speech processing method comprises the following contents:
[0127] S6, obtaining input speech information of a user.
[0128] S7, determining a plurality of third confidence degrees corresponding to the input speech information according to the input speech information, the plurality of third confidence degrees corresponding to a plurality of languages respectively.
[0129] S8, correcting the plurality of third confidence degrees to a plurality of fourth confidence degrees according to the scene features.
[0130] S9, determining a language of the input speech information according to the plurality of fourth confidence degrees.
[0131] By using the voice processing method, the third confidence is corrected according to the user feature, and the language of the input voice information is determined according to the fourth confidence obtained after the correction, so that the language of the input voice information can be determined more accurately, and the voice recognition capability can be improved. Here, the third confidence can be obtained in the same way as the first confidence, or in a different way. The fourth confidence can be obtained in the same way as the second confidence, or in a different way. The specific way of correction according to the scene feature can be the same as the specific way of correction according to the user feature in the above embodiment, or can be different.
[0132] The correction processing in the embodiment and the correction processing described in the above embodiment can be used in combination, i.e., the language confidence is corrected according to both the user feature and the scene feature, so that the language of the input voice information can be determined more accurately.
[0133] In the embodiment, optionally, a plurality of preset weights can be set according to the scene feature; and the plurality of third confidences are corrected into a plurality of fourth confidences according to the plurality of preset weights.
[0134] In addition, optionally, the setting way of the plurality of preset weights can be specifically: obtaining first voice data collected in advance and first language information of the first voice data recorded in advance; determining second voice data according to the first voice data and the scene feature; determining second language information of the second voice data according to the second voice data; and setting the plurality of preset weights according to the first language information and the second language information.
[0135] Optionally, as a specific implementation, a plurality of test weight groups are obtained, each test weight group including a plurality of test weights; a plurality of second language information are determined according to the second voice data and the plurality of test weight groups, the plurality of second language information corresponding to the plurality of test weight groups respectively; a plurality of accuracies of the plurality of second language information are determined according to the first language information and the plurality of second language information; and the plurality of preset weights are set according to a test weight group corresponding to second language information with the highest accuracy.
[0136] The following refers to Figure 5 A voice processing device according to an embodiment of the present application is described. Figure 5 A schematic structural diagram of the voice processing device according to an embodiment of the present application is shown. The voice processing device 190 is used to execute the voice processing method according to the embodiment described above with reference to Figure 3 The voice processing method according to the embodiment described above with reference to Figure 4 The voice processing method according to the embodiment described above with reference to Figure 5As shown, the voice processing device 190 includes a processing module 192 and a transceiver module 194. The processing module 192 can execute the contents of S2-S4 or S7-S9 described above, and the transceiver module 194 can execute the contents of S1 or S6 described above. Furthermore, the voice processing device 190 can be constructed in hardware, software, or a combination of both. The voice processing device 190 of this embodiment can achieve the same technical effects as the voice processing method described above; therefore, a repeated description of the technical effects is omitted here.
[0137] The following reference Figure 6 A language identification method provided in one embodiment of this application is described below.
[0138] Figure 6 This is a flowchart illustrating, illustratively, a language recognition method provided in one embodiment of this application. This language recognition method can be executed by a vehicle, in-vehicle device, vehicle infotainment system, in-vehicle computer, chip, processor, etc. Figure 6 As shown, in the language recognition method of this embodiment, firstly, in step S10, the user's input voice information is acquired. For example, the user's input voice data received by the microphone is acquired as the input voice information, or the microphone's input voice data is preprocessed to obtain the input voice information. In step S12, the input voice is recognized to obtain a first recognition confidence set for multiple languages, where multiple first recognition confidence scores in the first recognition confidence set correspond to multiple languages respectively. For example, the first recognition confidence set for multiple languages is {Chinese: 0.9; English: 0.1; Korean: 0; German: 0; Japanese: 0}. That is, the probability that the language of the input voice information is Chinese is 0.9, the probability that it is English is 0.1, and the probability that it is Korean, German, or Japanese is 0.
[0139] In step S14, it is determined whether there exists a first recognition confidence score greater than a threshold in the first recognition confidence score set for the multilingual languages. This threshold can be set to, for example, 0.8. When the determination result is "yes," meaning there exists a first recognition confidence score greater than the threshold (e.g., 0.9 for Chinese), a recognition result is generated and output based on this first recognition confidence score set. This recognition result can be the result indicating the recognized language (e.g., Chinese) or the first recognition confidence score set itself. Alternatively, as in other embodiments, S14 can be omitted, and step S18 described later can be performed directly.
[0140] In addition, when the judgment result in step S14 is "no", that is, there is no first recognition confidence level greater than the threshold, in step S18, the first recognition confidence level set is corrected and calculated according to the user's user characteristics to obtain the second recognition confidence level set.
[0141] As examples of the user feature, there can be mentioned a history language record of the user and a user-specified language.
[0142] The history language record refers to a record of the recognized language of the voice input by the user before the input of the above-mentioned input voice. The user-specified language refers to the kind of the system language set by the user (e.g., the system language of the voice interaction system, the system language of the operating system of the mobile phone when applied to the mobile phone, etc.). In addition, the user-specified language can be one or multiple (i.e., the user has set multiple system languages). The history language record can be obtained by querying the database with the voiceprint of the input voice, or by querying with the face information, iris information, etc. of the user. That is, the identity of the user can be determined according to the voiceprint, face information, iris information, etc., and thus the history language record of the user can be obtained from the database. Querying the history language record according to the voiceprint of the input voice can avoid the misrecognition of the language due to the misrecognition of the user (speaker) compared with the way of querying according to the face information, iris, etc. In addition, the voiceprint can be obtained according to the input voice information, while the way of querying according to the face information, iris, etc. requires obtaining the image of the user, and thus the way of querying according to the voiceprint requires less equipment and is faster.
[0143] In addition, it can be necessary to mention that the basis of the way of obtaining the user-specified language according to the voiceprint is that the voiceprint of the user is collected when the user sets the system language, and the voiceprint (or the identity of the user) is stored in the user-specified language database in association with the system language set by the user. The database mentioned here can be stored locally or on a platform with public credibility.
[0144] After the second recognition confidence set is obtained by correcting the first recognition confidence set according to the user feature, in step S20, the language recognition result is generated according to the second recognition confidence set. For example, the second recognition confidence set can be directly output as the recognition result, or when there is a second recognition confidence greater than a threshold in the second recognition confidence set, the second recognition confidence set or information indicating the recognized language is output, and when there is no second recognition confidence greater than the threshold in the second recognition confidence set, the first recognition confidence set is output as the recognition result.
[0145] By using the above method, the second recognition confidence is obtained by calculating the first recognition confidence according to the user feature, and the language recognition result is determined according to the second recognition confidence, and thus the language recognition capability can be improved.
[0146] Optionally, the language recognition method further comprises: when the plurality of second recognition confidences in the second recognition confidence set are less than the threshold, generating a language recognition result according to an automatic speech recognition confidence obtained by performing automatic speech recognition on the input speech or a natural language understanding confidence obtained by performing natural language understanding (NLU) on the input speech.
[0147] In this way, when the language of the input speech cannot be recognized according to the second recognition confidences, the language of the input speech is determined according to the automatic speech recognition confidence or the natural language understanding confidence, so that the language recognition capability can be improved. As a specific implementation, for example, the language whose automatic speech recognition confidence exceeds the threshold is taken as the recognized language of the input speech.
[0148] Optionally, in this embodiment, the first recognition confidence set can be obtained by: performing recognition on the input speech to obtain an initial confidence set; and multiplying the plurality of initial confidences in the initial confidence set by the plurality of preset weights in the preset weight set respectively to obtain the first recognition confidence set.
[0149] At this time, when the second recognition confidence set includes a second recognition confidence greater than the threshold, the preset weight set can be updated so that the preset weight of the recognized language of the plurality of languages whose second recognition confidence is greater than the threshold is increased relative to the preset weights of the other languages.
[0150] In this way, when the second recognition confidence set includes a second recognition confidence greater than the threshold, that is, when the language of the input speech can be determined, the preset weight set is updated according to the above method, so that when the subsequent input speech is processed, the updated preset weight set is used, so that the accuracy of language recognition can be improved and the language recognition capability can be improved.
[0151] The specific way of updating the preset weight set can be: performing correction calculation on the preset weight set to obtain a corrected weight set; and when the plurality of corrected weights in the corrected weight set are within the weight range, updating the preset weight set with the values of the plurality of corrected weights.
[0152] If the preset weight exceeds the weight range, the credibility of the result of language recognition obtained according to the preset weight is relatively low, and therefore, in this way, the preset weight is corrected within the weight range, so that the misrecognition rate of the language can be inhibited.
[0153] In addition, in this embodiment, the preset weight set can be preset according to the scene characteristics. Therefore, the language recognition result can be obtained with the preset weight that is most suitable for the scene as much as possible, so that the language recognition capability can be improved.
[0154] The scene features can include environmental features and / or audio collector features. The environmental features can include environmental signal-to-noise ratio, power DC-AC information, or environmental vibration amplitude, and the audio collector features can include microphone arrangement information. The microphone arrangement information refers to whether it is a single microphone or a microphone array, or whether it is a linear array, a planar array, or a stereo array when it is a microphone array.
[0155] The environmental signal-to-noise ratio, the power DC-AC information, the environmental vibration amplitude, and the microphone arrangement information can all affect the language confidence, and therefore, the preset weights are adjusted according to these information, and the language recognition is performed on this basis, so as to improve the language recognition capability.
[0156] The specific manner of setting the preset weight set can be: obtaining a plurality of test weight sets; inputting a pseudo-environmental data set into a language recognition model, the pseudo-environmental data set being obtained according to scene features and a noiseless data set; obtaining a plurality of first recognition confidence sets under the plurality of test weight sets according to an initial confidence set output by the language recognition model; calculating a prediction accuracy of the plurality of first recognition confidence sets according to language information of the pseudo-environmental data set; determining a test weight set corresponding to a first recognition confidence set with the highest prediction accuracy in the plurality of test weight sets as an optimal test weight set; and setting the preset weight set with values of the plurality of test weights in the optimal test weight set.
[0157] Optionally, when the set preset weight is within the weight range, the setting is made effective, and when the set preset weight is not within the weight range, the setting is cancelled. Alternatively, when the set preset weight is not within the weight range, the setting is still made effective, but other ways are preferred to obtain the language recognition result, for example, the recognition language of the input voice is determined according to the user-specified language, or the characteristic similarity is obtained by comparing the input voice this time with the input voice in the historical language record, and if the characteristic similarity is greater than a similarity threshold, the language of the input voice in the historical language record is determined as the recognition language of the input voice this time.
[0158] The weight range can be set in the following manner: obtaining a plurality of test data sets; obtaining a plurality of test weight sets; inputting the test data sets into a language recognition model; obtaining a plurality of first recognition confidence sets under the plurality of test weight sets according to an initial confidence set output by the language recognition model and the plurality of test weight sets; calculating a prediction accuracy of the plurality of first recognition confidence sets according to language information of the test data sets; determining a test weight set corresponding to a first recognition confidence set with the highest prediction accuracy in the plurality of test weight sets as an optimal test weight set; obtaining the optimal test weight set of the plurality of test data sets; and obtaining the weight range of a plurality of languages according to the optimal test weight set of the plurality of test data sets.
[0159] The test data set can be a pre-acquired speech data set, and the language information of the test data set is known.
[0160] In this way, the weight range of the preset weight set of the multi-language is set by testing the language recognition model with a large amount of language data sets, that is, the robustness range of the language recognition model is specified, so that the language recognition model works within the range, thereby ensuring the reliability of the language recognition result.
[0161] Figure 7 As shown in the structure schematic diagram of the language recognition device provided by an embodiment of the present application. As Figure 7 An embodiment of the present application provides a language recognition device, which is used to execute the language recognition method shown in Figure 6 The structure of the language recognition device can be known from the description of the language recognition method in Figure 6 Therefore, the language recognition device 10 is only relatively briefly described here.
[0162] As Figure 7 The language recognition device 10 includes: an input speech acquisition module 17, configured to acquire input speech of a user; a language recognition module 12, configured to recognize the input speech to obtain a first recognition confidence set, a plurality of first recognition confidences in the first recognition confidence set corresponding to a plurality of languages respectively; a language confidence correction module 16, configured to correct and calculate the first recognition confidence set according to a user feature of the user to obtain a second recognition confidence set; and a recognition result generation module 18, configured to generate a language recognition result according to the second recognition confidence set.
[0163] In this way, the second recognition confidence is obtained by calculating the first recognition confidence according to the user feature, and the language recognition result is determined according to the second recognition confidence, so that the language recognition capability can be improved.
[0164] Optionally, the language confidence correction module 16 can correct and calculate the first recognition confidence set according to the user feature of the user to obtain the second recognition confidence set when the plurality of first recognition confidences are less than a threshold value.
[0165] Optionally, the user feature includes a historical language record.
[0166] In this way, when the language of the input speech is difficult to be recognized according to the first recognition confidence, the first recognition confidence is corrected according to the historical language record of the user, and the language of the input speech is determined on this basis, so that the language recognition capability can be improved. The historical language record of the user here refers to the record of the language of the speech input by the user before the input of the input speech.
[0167] Optionally, the historical language record is queried according to a voiceprint of the input speech.
[0168] In the above manner, the historical language record is queried according to the voiceprint of the input speech, which can avoid language misrecognition caused by misrecognition of the user (speaker) (recognition of a non-speaker as a speaker) compared with a manner of querying according to face information, iris, etc.
[0169] Optionally, the user feature includes a user specified language.
[0170] In the above manner, when the language of the input speech is difficult to be recognized according to the first recognition confidence, the first recognition confidence is corrected according to the user specified language, and the language of the input speech is determined on this basis, so that the language recognition capability can be improved.
[0171] Optionally, the user specified language is queried according to a voiceprint of the input speech.
[0172] In the above manner, the user specified language is queried according to the voiceprint of the input speech, which can avoid language misrecognition caused by misrecognition of the user (speaker) (recognition of a non-speaker as a speaker) compared with a manner of querying according to face information, iris, etc.
[0173] Optionally, the recognition result generation module is further configured to: when the plurality of second recognition confidences in the second recognition confidence set are less than a threshold value, generate a language recognition result according to an automatic speech recognition (ASR) confidence obtained by performing automatic speech recognition on the input speech.
[0174] In the above manner, when the language of the input speech is difficult to be recognized according to the second recognition confidence, the language of the input speech is determined according to the automatic speech recognition confidence, so that the language recognition capability can be improved. As a specific implementation manner, for example, a language with an automatic speech recognition confidence exceeding a threshold value is taken as the recognized language of the input speech.
[0175] Optionally, the recognition result generation module is further configured to: when the plurality of second recognition confidences in the second recognition confidence set are less than a threshold value, generate a language recognition result according to a natural language understanding confidence obtained by performing natural language understanding on the input speech.
[0176] In the above manner, when the language of the input speech is difficult to be recognized according to the second recognition confidence, the language of the input speech is determined according to the automatic speech recognition confidence, so that the language recognition capability can be improved. As a specific implementation manner, for example, a language with an automatic speech recognition confidence exceeding a threshold value is taken as the recognized language of the input speech.
[0177] Optionally, the language recognition module is further configured to: recognize the input speech to obtain an initial confidence set; multiply each initial confidence in the initial confidence set by a preset weight in the preset weight set to obtain a first recognition confidence set; and the language confidence correction module is further configured to: update the preset weight set when there is a second recognition confidence greater than a threshold in the second recognition confidence set, so that the preset weight of the recognition language whose second recognition confidence is greater than the threshold in the plurality of languages is increased relative to the preset weights of the other languages.
[0178] When there is a second recognition confidence greater than a threshold in the second recognition confidence set, that is, the language of the input speech can be determined, the preset weight set is updated according to the above manner, so that when the subsequent input speech is processed, the updated preset weight set is used, thereby improving the accuracy of language recognition and the language recognition capability.
[0179] Optionally, the language confidence correction module is further configured to: perform correction calculation on the preset weight set to obtain a modified weight set; and update the preset weight set with the values of the plurality of modified weights in the modified weight set when the plurality of modified weights are within a weight range.
[0180] When the preset weight is outside the weight range, the reliability of the language recognition result obtained according to the preset weight is relatively low, and therefore, the preset weight is corrected within the weight range according to the above manner, thereby inhibiting the misrecognition rate of the language.
[0181] Optionally, the language recognition module is further configured to: recognize the input speech to obtain an initial confidence set; multiply each initial confidence in the initial confidence set by a preset weight in the preset weight set to obtain a first recognition confidence set; and the language confidence correction module is further configured to set the preset weight set according to a scene feature.
[0182] According to the above manner, the preset weight set is set according to the scene feature, so that the preset weight set can be adapted to different scenes to obtain a language recognition result that is most suitable for the scene as much as possible, thereby improving the language recognition capability.
[0183] Optionally, the scene feature includes an environment feature and / or an audio collector feature.
[0184] Optionally, the environment feature includes an environment signal-to-noise ratio, power DC / AC information, or environment vibration amplitude, and the audio collector feature includes microphone arrangement information. The microphone arrangement information refers to whether it is a single microphone or a microphone array, or whether it is a linear array, a planar array, or a stereo array when it is a microphone array.
[0185] The environment signal-to-noise ratio, power DC-AC information, environment vibration amplitude, and microphone arrangement information can affect the language confidence, and therefore, the preset weight is adjusted according to the information, and language recognition is performed based on the preset weight, so that the language recognition capability can be improved.
[0186] Optionally, the language confidence correction module is further configured to: obtain a plurality of test weight sets; input a pseudo-environment data set into the language recognition model, the pseudo-environment data set being obtained according to the scene feature and the noiseless data set; obtain a plurality of first recognition confidence sets in the case of the plurality of test weight sets according to the initial confidence set output by the language recognition model; calculate the prediction accuracy of the plurality of first recognition confidence sets according to the language information of the pseudo-environment data set; determine the test weight set corresponding to the first recognition confidence set with the highest prediction accuracy in the plurality of test weight sets as the optimal test weight set; and set the preset weight set with the values of the plurality of test weights in the optimal test weight set. In this way, the language confidence correction module can be regarded as having the preset weight setting module.
[0187] Optionally, the language confidence correction module is further configured to set a plurality of preset weights within a weight range.
[0188] Optionally, the weight range is set in the following manner: obtaining a plurality of test data sets; obtaining a plurality of test weight sets; inputting the test data sets into the language recognition model; obtaining a plurality of first recognition confidence sets in the case of the plurality of test weight sets according to the initial confidence set output by the language recognition model and the plurality of test weight sets; calculating the prediction accuracy of the plurality of first recognition confidence sets according to the language information of the test data sets; determining the test weight set corresponding to the first recognition confidence set with the highest prediction accuracy in the plurality of test weight sets as the optimal test weight set; obtaining the optimal test weight set of the plurality of test data sets; and obtaining the weight range of the plurality of languages according to the optimal test weight set of the plurality of test data sets. The function of setting the weight range can be implemented by the language recognition device 10, which can be regarded as having the weight range setting module, or by a test device for testing the language recognition device 10.
[0189] In the above manner, the weight range of the preset weight set of the plurality of languages is set by testing the language recognition model with a large number of language data sets, that is, the robustness range of the language recognition model is specified, so that the language recognition model works within the range, thereby ensuring the reliability of the language recognition result.
[0190] An embodiment of the present application provides a computing device including a processor and a memory, the memory storing program instructions which, when executed by the processor, cause the processor to perform the voice processing method and the language recognition method described above. The computing device can be a server, a personal computer, a mobile phone, a tablet computer, a wearable device, or the like.Figure 17 Learn more from the explanations provided.
[0191] One embodiment of this application provides a computer-readable storage medium storing program instructions, characterized in that, when executed by a computer, the program instructions cause the computer to perform the aforementioned speech processing method and language recognition method.
[0192] One embodiment of this application provides a computer program that, when executed by a computer, causes the computer to perform the aforementioned speech processing method and language recognition method.
[0193] Figure 8 This flowchart illustrates a voice interaction method provided in one embodiment of this application. Some steps in this voice interaction method are the same as those in the language recognition method described above. Here, the same reference numerals are used to label the same content, and the description is simplified.
[0194] like Figure 8 As shown, in this voice interaction method, firstly, in step S10, the user's input voice information is acquired. For example, the user's input voice information received by the microphone is acquired. At this time, on the one hand, in step S40, automatic speech recognition is performed on the input voice using a speech recognition model; on the other hand, in step S12, language recognition is performed on the input voice using a language recognition model. As another embodiment, automatic speech recognition and language recognition can also be performed sequentially.
[0195] In addition, in step S40, in order to recognize input speech in multiple languages, speech content recognition processing is performed on the input speech using speech recognition models of multiple different languages (five languages in this embodiment: Chinese, English, Korean, German and Japanese) to obtain text Ti in multiple different languages.
[0196] Subsequently, in step S42, multiple texts Ti are input into the text translation model, which translates these texts Ti into text Ai in the target language (e.g., Chinese).
[0197] Next, in step S44, multiple texts Ai are sequentially input into the semantic understanding model. The semantic understanding model performs semantic understanding processing on these texts Ai, thereby obtaining multiple corresponding candidate commands Oi. A candidate command means a command that has not yet been identified as to be executed.
[0198] In addition, in step S12, the input speech is recognized to obtain a first recognition confidence set of multiple languages, and the multiple first recognition confidences in the first recognition confidence set correspond to the multiple languages respectively. For example, the first recognition confidence set of multiple languages is {Chinese: 0.9; English: 0.1; Korean: 0; German: 0; Japanese: 0}.
[0199] In step S14, it is judged whether there is a first recognition confidence greater than a threshold value in the first recognition confidence set of multiple languages. The threshold value herein can be set to 0.8 for example. When the judgment result is "yes", that is, there is a first recognition confidence (for example, the first recognition confidence of Chinese 0.9) greater than the threshold value, in step S16, the language (for example, Chinese) with the first recognition confidence greater than the threshold value is determined as the recognition language of the input speech as the recognition result.
[0200] After that, in step S26, the candidate command corresponding to the recognition language (for example, Chinese) is selected as the target command to be executed from the multiple candidate commands Oi obtained in step S44 according to the recognition language (for example, Chinese), and then the processing of making the target command executed is performed. For example, when the target command is "turn on the air conditioner", the corresponding control is performed to turn on the air conditioner.
[0201] In addition, when the judgment result in step S14 is "no", that is, there is no first recognition confidence greater than the threshold value, or the multiple first recognition confidences are less than the threshold value, in step S18, the first recognition confidences are corrected according to the user characteristics. The specific content of the correction has been described in detail above, and thus will not be described here.
[0202] After that, in step S22, it is judged whether there is a second recognition confidence greater than a threshold value, and when the judgment result is "yes", in step S24, the language with the second recognition confidence greater than the threshold value is determined as the recognition language of the input speech, and then the processing in step S26 is performed.
[0203] In addition, when the judgment result in step S22 is "yes", that is, there is no second recognition confidence greater than the threshold value, or the multiple second recognition confidences are less than the threshold value, in step S28, it is judged whether there is an ASR confidence greater than a threshold value. When the judgment result is "yes", in step S30, the language with the ASR confidence greater than the threshold value is determined as the recognition language, and then the processing in step S26 is performed.
[0204] When the judgment result in step S28 is "No", that is, there is no ASR confidence score greater than the threshold, or multiple ASR confidence scores are less than the threshold, in step S32, it is determined whether there is an NLU confidence score greater than the threshold. When the judgment result is "Yes", in step S34, the language with an NLU confidence score greater than the threshold is determined as the recognition language, and then the processing in step S26 is executed.
[0205] When the judgment result in step S32 is "no", a message indicating that the speech content recognition failed can be output and the process ends.
[0206] By employing the above voice interaction method, in language recognition, a second recognition confidence level is obtained by calculating the first recognition confidence level based on user characteristics, and the language recognition result is determined based on the second recognition confidence level. In this way, the language recognition capability can be improved, thereby improving the voice interaction capability.
[0207] Figure 9 The diagram shown is a structural schematic of a voice interaction system provided in one embodiment of this application. Figure 9 As shown, the voice interaction system (or voice interaction device) 20 includes a voice recognition module 110, a language recognition module 12, a text translation module 130, a semantic understanding module 140, an input voice acquisition module 17, a language confidence correction module 16, and a control module 170. This voice interaction device is used to perform reference... Figure 8 The voice interaction method described herein is such that the specific processing flow is omitted here. Furthermore, this voice interaction system 20, like the language recognition device 10 described above, includes a language recognition module 12, an input voice acquisition module 17, and a language confidence correction module 16, which are labeled using the same reference numerals, and detailed descriptions of these are omitted here. Additionally, this voice interaction system may also include an execution device, such as a speaker or a display device.
[0208] The following is a brief explanation of the correspondence between the voice interaction system 20 and the steps of the above-mentioned voice interaction method.
[0209] Speech recognition module 110 execution Figure 8 Step S40. Language recognition module 12 executes. Figure 8 Step S12. Text translation module 130 executes. Figure 8 Step S42. Semantic understanding module 140 executes. Figure 8 Step S44. Input voice acquisition module 17 executes. Figure 8 Step S10. Language confidence correction module 16 executes. Figure 8 Step S18. Control module 170 executes. Figure 8The steps S14, S16, S22, S24, S28, S30, S32, S34 in the method for voice interaction can be executed by the voice interaction system 100. In addition, the steps S14, S16, S22, S24 can also be executed by the language confidence correction module 16.
[0210] In addition, as can be known from the above description, the voice interaction method described above essentially comprises a multi-language voice recognition method, which can recognize input voice in multiple languages; the voice interaction device described above also comprises a voice recognition device for executing the multi-language voice recognition method. Since there are many repeated contents, the voice recognition method and the voice recognition device will not be described separately in the embodiments of the present application. Figure 9 The voice interaction method described above essentially comprises a multi-language voice recognition method, which can recognize input voice in multiple languages; the voice interaction device described above also comprises a voice recognition device for executing the multi-language voice recognition method. Since there are many repeated contents, the voice recognition method and the voice recognition device will not be described separately in the embodiments of the present application. Figure 11- Figure 17 The voice interaction method described above essentially comprises a multi-language voice recognition method, which can recognize input voice in multiple languages; the voice interaction device described above also comprises a voice recognition device for executing the multi-language voice recognition method. Since there are many repeated contents, the voice recognition method and the voice recognition device will not be described separately in the embodiments of the present application.
[0211] The voice interaction system 100 and the voice interaction method executed by the voice interaction system 100 according to an embodiment of the present application will be described below. Figure 11 The voice interaction system 100 and the voice interaction method executed by the voice interaction system 100 according to an embodiment of the present application will be described below.
[0212] In the present embodiment, the voice interaction system 100 is taken as an example of being applied to a car to constitute a vehicle-mounted voice interaction system, however, the present application is not limited thereto, but can also be applied to other scenarios, such as smart home, smart robot, smart voice question and answer, smart voice analysis, real-time voice monitoring and analysis, etc. In addition, the vehicle-mounted voice interaction system also constitutes a vehicle control device. In addition, it can be understood that through the voice interaction method described above and the voice interaction system described in the present embodiment, a voice processing method, device and system are provided in the embodiments of the present application.
[0213] <System Architecture>
[0214] The voice interaction system 100 in the present embodiment can receive input voice of a user (i.e. a speaker), and execute corresponding processing in response to the content of the input voice, such as opening an air conditioner, opening a window, etc. Moreover, the voice interaction system 100 can respond to voice in multiple different languages, for example, in the present embodiment, it can respond to voice in Chinese, English, Korean, German and Japanese.
[0215] Here, the voice in different languages includes voice in different language families, such as Chinese and English, and also includes voice in different small languages under the same language family, such as Mandarin and Cantonese in Chinese.
[0216] Figure 12This is a schematic diagram illustrating a voice interaction system according to an embodiment of this application. The voice interaction system 100 includes a voice recognition module 110, a language recognition module 120, a text translation module 130, a semantic understanding module 140, a command parsing and execution module 150, and a language confidence correction module 160. In addition, the voice interaction system 100 may also include a microphone, a speaker, a camera, or a display.
[0217] Figure 12 This is a flowchart illustrating one aspect of a voice interaction method described in one embodiment. See below for reference. Figure 12 The processing flow of the voice interaction system 100 is described to give a general overview of the architecture of the voice interaction system 100.
[0218] like Figure 17 As shown, when a user emits a voice message S, the voice interaction system 100 acquires the voice message (referred to as the input voice message) through a microphone. Firstly, ① the input voice message is processed by the input voice recognition module 110, which calls multiple voice recognition sub-modules in different languages (in this embodiment, Chinese, English, Korean, German, and Japanese; obviously, it could be any number of other languages) to perform voice content recognition processing, obtaining multiple texts Ti in different languages. ② The multiple texts Ti are input to the text translation module 130, which translates these texts Ti into text Ai in the target language (e.g., Chinese). ③ The multiple texts Ai are sequentially input to the semantic understanding module 140, which performs semantic understanding processing on these texts Ai, thereby obtaining multiple corresponding candidate commands Oi.
[0219] On the other hand, ④ the user's input voice is also processed by the input language recognition module 120. The language recognition module 120 performs language recognition processing on the input voice, generates initial confidence scores for multiple languages, and multiplies each initial confidence score by a corresponding preset weight to obtain the recognition confidence scores for multiple languages.
[0220] ⑤ When there is a recognition confidence level greater than the threshold λ among the recognition confidence levels of multiple languages, the language of the input speech can be considered as the language with a recognition confidence level greater than the threshold λ. The command parsing and execution module 150 determines the candidate command Oi corresponding to the language with a recognition confidence level greater than the threshold λ among the multiple candidate commands Oi as the target command to be executed, and performs corresponding processing according to the content of the target command.
[0221] In addition, when there is no recognition confidence score greater than the threshold λ among the recognition confidence scores of multiple languages, the language confidence score correction module 160 corrects the recognition confidence scores of multiple languages based on user characteristics, etc. The specific details will be described in detail later.
[0222] The structure of each component of the voice interaction system 100 will be described below.
[0223] <Structure>
[0224] In the present embodiment, the voice recognition module 110, the language recognition module 120, the text translation module 130 and the semantic understanding module 140 respectively include algorithm models, i.e. voice recognition model, language recognition model, text translation model and semantic understanding model, by which the voice recognition processing, the language recognition processing, the text translation processing and the semantic understanding processing are respectively performed.
[0225] The voice recognition module 110 is configured to convert the voice of a person, i.e. the voice to be recognized, into text in a corresponding language, which can also be said as predicting the content of the voice or performing automatic speech recognition (ASR). Here, the voice recognition module 110 has a plurality of voice recognition sub-modules, each of which corresponds to a language and is configured to convert the voice into text in the corresponding language Ti. For example, in the present embodiment, the voice recognition module 110 has voice recognition sub-modules for Chinese, English, Korean, German and Japanese, which are respectively configured to convert the input voice into Chinese text T1, English text T2, Korean text T3, German text T4 and Japanese text T5. After the recognition is completed, the sub-modules output the text Ti as the prediction result and the confidence of the text Ti, which is referred to as ASR confidence. The ASR confidence represents the prediction probability of the text predicted by the sub-module, or the prediction probability of the content of the voice.
[0226] The text translation module 130 is configured to convert text in one natural language (source language) into text in another natural language (target language), for example, to convert English text into Chinese text. Here, the text translation module 130 has a plurality of text translation sub-modules, each of which corresponds to a language. For example, in the present embodiment, Chinese is the target language, and thus the text translation module 130 has text translation sub-modules for English, Korean, German and Japanese, which are respectively configured to translate English text, Korean text, German text and Japanese text into Chinese text Ai. In addition, when Chinese text is input into the text translation module 130, the text translation module 130 can not process the input Chinese text and output the input Chinese text as it is, because Chinese is the target language for translation. Finally, the text translation module 130 outputs five Chinese texts Ai.
[0227] The semantic understanding module 140 is configured to perform natural language understanding (NLU) on the text in the target language, which can also be said to be predicting the intent of the text, and generating a command that can be understood by a machine. For example, the text is "please play the song XX", after the semantic understanding module 140, the machine can get the intent "please play the song XX". While generating the command, the semantic understanding module 140 also generates an NLU confidence, which represents the prediction probability of the semantic understanding module 140 on the intent of the text. In addition, since the speech recognition module 110 outputs five language texts, the semantic understanding module 140 will eventually generate commands and five NLU confidences corresponding to the five languages. Furthermore, the commands output by the semantic understanding module 140 have not been determined to be executed, so they are called candidate commands.
[0228] The language identification (LID) module is configured to identify the language of the input speech of the user, that is, the speech to be recognized, which can also be said to be predicting which one of the multiple languages the input speech of the user belongs to. For example, in the present embodiment, the language identification module 120 identifies which one of Chinese, English, Korean, German and Japanese the input speech belongs to, and outputs a set of recognition confidences of multiple languages as the recognition result, which represents the prediction probability of the language identification module 120 on which language the input speech belongs to. In addition, in the language identification model of the present embodiment, the input speech is algorithmically recognized to obtain the confidences of multiple languages (this confidence is called initial confidence), and the initial confidences of multiple languages are multiplied by the corresponding preset weight values to obtain the recognition confidences of multiple languages, which are output by the language identification module 120 as the prediction result. The calculation of multiplying the initial confidences by the preset weight values can be performed by the language identification model, or can not be performed by the language identification model.
[0229] The command analysis and execution module 150 is configured to select a target command to be executed from the candidate commands output by the semantic understanding module 140 according to the output of the language identification module 120. In the embodiment, when there is a recognition confidence greater than a threshold λ (for example, set to 0.8 or above) in the recognition confidences of the multiple languages output by the language identification module 120, the command analysis and execution module 150 determines the language with the recognition confidence greater than the threshold λ as the language of the above-mentioned input speech of the user, and determines the candidate command corresponding to the language as the target command to be executed. For example, when the confidences of the multiple languages output by the language identification module 120 are {Chinese: 0.9; English: 0.1; Korean: 0; German: 0; Japanese: 0}, Chinese is determined as the language of the above-mentioned input speech of the user, and the candidate command corresponding to Chinese among the five candidate commands output by the semantic understanding module 140 is determined as the target command to be executed.
[0230] After determining the target command to be executed, the command analysis and execution module 150 performs control for enabling the target command to be executed, for example, when the determined target command to be executed is "please play the song XX", and the voice interaction system 100 has a music playing module, the command analysis and execution module 150 controls the music playing module to play the song XX. If the music playing module does not belong to the voice interaction system 100 in terms of rights and is not controlled by the command analysis and execution module 150, the determined target command to be executed can be sent to a higher controller common to the voice interaction system 100 and the music playing module, and the higher controller sends a command to the controller of the music playing module to play the song XX.
[0231] In addition, after determining the target command to be executed, the command analysis and execution module 150 can respond to the user through a speaker or a display, for example, when the determined target command to be executed is "please play the song XX", the command analysis and execution module 150 controls the speaker to output the sound "OK, playing for you" to respond to the user.
[0232] On the other hand, in the embodiment, when there is no recognition confidence greater than a threshold λ (for example, 0.8) in the set of recognition confidences of the multiple languages output by the language identification module 120, the command analysis and execution module 150 determines the target command to be executed in other ways, which are exemplified as follows.
[0233] Method one
[0234] The command analysis and execution module 150 performs correction calculation on the recognition confidence of the above-mentioned multiple languages (corresponding to the "first recognition confidence" in the present application), and the command analysis and execution module 150 performs corresponding processing according to the output of the language confidence correction module 160. For example, the correction can be performed according to user characteristics, and the user characteristics herein include historical language records of the user and user-specified languages, etc. The historical language records and the user-specified languages can be determined according to the audio features (i.e., voiceprints) to determine the user identity, and the user identity is queried in the historical language record database and the user-specified language database of the voice interaction system 100. The specific content of the correction will be described in detail later. After obtaining the corrected recognition confidence, the command analysis and execution module 150 determines the language of the input voice according to the corrected recognition confidence, and performs corresponding processing. For example, when there is a recognition confidence greater than a threshold λ in the corrected recognition confidence of the multiple languages, the language with the recognition confidence greater than the threshold λ is determined as the language of the user input voice, and the candidate command corresponding to the determined language is determined as the target command to be executed.
[0235] As a specific implementation manner of the correction calculation on the recognition confidence in the first manner, the numerical value of the recognition confidence can be directly corrected, or the preset weight can be corrected and calculated, and then the recognition confidence is calculated again according to the initial confidence set and the corrected preset weight set.
[0236] In the present embodiment, the language confidence correction module 160 has a audio feature-based adjustment module 162, a video feature-based adjustment module 163, and a comprehensive adjustment module 164, which are used to correct the language confidence in different ways.
[0237] The second manner
[0238] The command analysis and execution module 150 determines the target command to be executed according to the ASR confidence output by the voice recognition module 110 or the NLU confidence output by the semantic understanding module 140. For example, when there is an ASR confidence greater than an ASR confidence threshold (which can be set to the same value as the above-mentioned threshold λ, for example, 0.8), the language corresponding to the ASR confidence is determined as the language of the input voice, and the candidate command corresponding to the language is determined as the target command to be executed. Alternatively, when there is an NLU confidence greater than an NLU confidence threshold (which can be set to the same value as the above-mentioned threshold λ, for example, 0.8), the language corresponding to the NLU confidence is determined as the language of the input voice, and the candidate command corresponding to the language is determined as the target command to be executed.
[0239] The execution timing of the second mode can be freely set. Optionally, the second mode can be executed after the first mode, or before the first mode, or between the first mode and the second mode, or among the multiple modes listed in the description of the first mode.
[0240] The third mode
[0241] The command analysis and execution module 150 determines the language of the input speech by feature similarity. For example, the audio data of the current input speech is compared with the audio data of the historical input speech in the historical record, and the feature similarity of the two is obtained by cosine similarity, linear regression or deep learning, and when the feature similarity exceeds the threshold, the recognition language of the historical input speech is determined as the language of the current input speech.
[0242] The execution timing of the third mode can be freely set. Optionally, the third mode can be executed after or before the first mode or the second mode, or between the first mode and the second mode, or among the multiple modes listed in the description of the first mode.
[0243] The confidence correction module is described below.
[0244] The language confidence correction module 160 includes a real-time scene adaptation module 161, an audio feature-based adjustment module 162, a video feature-based adjustment module 163, and a comprehensive adjustment module 164.
[0245] The real-time scene adaptation module 161 is used to initialize the preset weight set of multiple languages according to the environmental features and the audio collector (i.e. microphone) features when the language recognition model initially contacts the scene. The initial contact scene here is, for example, when the user just purchases the voice interaction system or the vehicle. At this time, the user generally turns on the voice interaction system to make some basic settings or tests. In this embodiment, the real-time scene adaptation module 161 can use this opportunity to initialize the preset weight set. In addition, as other embodiments, the initialization of the preset weight set is not limited to being executed at the initial contact scene, but can also be executed at other appropriate times, such as when a new audio collector is replaced, or when the execution timing is selected by the user.
[0246] The video feature-based adjustment module 163 is used to correct the recognition confidence set of multiple languages according to the captured user image.
[0247] The audio feature adjustment module 162 is used to correct the recognition confidence set of multiple languages based on the user's voice information. Specifically, it can query the database of the voice interaction system 100 based on the voice information (voiceprint) to obtain the user's historical language records, and correct the recognition confidence set of multiple languages based on the historical language records.
[0248] The comprehensive adjustment module 164 is mainly used to correct the recognition confidence set of multiple languages based on the user-specified language. The user-specified language is obtained by querying the database of the voice interaction system 100 based on the voiceprint of the input speech.
[0249] When a language recognition confidence score greater than a threshold λ exists among the corrected recognition confidence scores, the confidence correction module recalculates the preset weight set for the multiple languages, increasing the preset weight of the language with a confidence score greater than the threshold λ relative to the preset weights of other languages. This results in a corrected weight set. Next, the confidence correction module determines whether each corrected weight in the corrected weight set is within its weight range. If the result indicates it is within the weight range, the preset weight set is updated with the values from the corrected confidence weight set for use by the language recognition module 120 in subsequent language recognition.
[0250] The functions of the speech recognition module 110, language recognition module 120, text translation module 130, semantic understanding module 140, command parsing and execution module 150, and confidence correction module can be implemented by the processor executing the program (software) stored in the memory, or by hardware such as LSI (Large Scale Integration) and ASIC (Application Specific Integrated Circuit).
[0251] Typically, these modules can be constructed using an electronic control unit (ECU). Optionally, a module can be constructed using one ECU, multiple ECUs, or multiple modules can be constructed using one ECU.
[0252] An ECU (Electronic Control Unit) is a control device composed of integrated circuits used to perform a series of functions, such as data analysis, processing, and transmission. For example... Figure 3 As shown in the figure, this application provides an electronic control unit (ECU), which includes a microcomputer, an input circuit, an output circuit, and an analog-to-digital (A / D) converter.
[0253] The main function of the input circuit is to pre-process the input signal (for example, the signal from the sensor), and the input signal is different, and the processing method is also different. Specifically, because the input signal has two types: analog signal and digital signal, the input circuit can include an input circuit for processing analog signals and an input circuit for processing digital signals.
[0254] The main function of the A / D converter is to convert the analog signal into a digital signal, and the analog signal is pre-processed by the corresponding input circuit and input into the A / D converter for processing and conversion into a digital signal accepted by the microcomputer.
[0255] The output circuit is a device for establishing a connection between the microcomputer and the actuator. Its function is to convert the processing result issued by the microcomputer into a control signal to drive the actuator to work. The output circuit generally uses a power transistor, and according to the instructions of the microcomputer, the electronic circuit of the actuator is controlled by being turned on or off.
[0256] The microcomputer includes a central processing unit (CPU), a memory, and an input / output (I / O) interface. The CPU is connected to the memory and the I / O interface through a bus, and information can be exchanged between them through the bus. The memory can be a read-only memory (ROM) or a random access memory (RAM) or the like. The I / O interface is a connection circuit for exchanging information between the central processing unit (CPU) and the input circuit, the output circuit or the A / D converter. Specifically, the I / O interface can be divided into a bus interface and a communication interface. The memory stores a program, and the CPU calling the program in the memory can realize the functions of the above modules, or execute the methods described with reference to Figure 4 、 Figure 6 、 Figure 8 、 Figure 12 、 Figure 14 and the like.
[0257] In addition, as described above, the voice interaction system 100 also has a microphone, a speaker, a camera or a display. The microphone is used to obtain the input voice of the user, corresponding to the voice obtaining module in the present application. The speaker is used to play sound, for example, to play the response sound "OK" for the input voice of the user. The camera is used to collect the facial image of the user and the like, and send the collected image to the command analysis and execution module 150, which can perform image recognition on the image, so that the identity of the user can be authenticated. The display is used to respond to the input voice of the user, for example, when the input voice is "play the song XX", the display shows the playing screen of the song.
[0258] <Actions and processing flow>
[0259] The voice interaction system 100 will be described in more detail in combination with the description of the actions and processing flow of the voice interaction system 100. In addition, the description of the actions and processing flow of the voice interaction system 100 is combined with the description of the voice interaction method involved in the present embodiment, and it can also be known from the following description that the voice interaction method includes a language recognition method (corresponding to the processing of the language recognition module 120, part of the processing of the command analysis and execution module 150, the processing of the confidence correction module, etc.).
[0260] As described above, the language recognition module 120 uses a language recognition model to perform language recognition. In the present embodiment, as shown in step S210 in Figure 13 When the trained language recognition model is initially contacted with the scene, the real-time scene adaptation module 161 initializes the preset weight set of multiple languages according to the environmental characteristics, the audio collector characteristics, as shown in step S210 in Figure 13 An example of the initialization method will be described below.
[0261] As shown in Figure 13 , the real-time scene adaptation module 161 generates a pseudo-environmental data set according to the environmental characteristics, the audio collector characteristics and the expert data set. Here, the environmental characteristics include, for example, the environmental signal-to-noise ratio, the microphone power source information (direct current-alternating current information) or the environmental vibration amplitude, etc. The microphone power source information can be obtained, for example, through the controller area network (CAN) signal of the vehicle. The audio collector characteristics mainly include the microphone arrangement information (single microphone or microphone array, wherein the microphone array includes linear array, planar array and stereo array). The expert data set is a batch of pre-collected multi-person, multi-language and noise-free audio data set, the content (the language of each voice data) of which is pre-recorded and known.
[0262] Additionally, randomly initialize N different confidence weight sets 1-N for multiple languages, for example, see [link to example]. Figure 13 The examples listed are confidence weight set 1 {Chinese: 0.80; English: 0.04; Korean: 0.06; Japanese: 0.05; German: 0.05}, confidence weight set 2 {Chinese: 0.21; English: 0.19; Korean: 0.22; Japanese: 0.20; German: 0.18}, and confidence weight set N {Chinese: 0.31; English: 0.09; Korean: 0.12; Japanese: 0.25; German: 0.23}.
[0263] Input the simulated environment dataset into the language recognition model to obtain an initial confidence set for multiple languages. Figure 13 In the dataset {Chinese: p1; English: p2; Korean: p3; Japanese: p4; German: p5}, the initial confidence level is multiplied by the confidence weight set 1 - confidence weight set N for each of the N languages, resulting in N recognition confidence sets. Since the content of the expert dataset (the language of each speech data point) is known, the accuracy acc of the resulting N recognition confidence sets can be calculated. The confidence weight set corresponding to the recognition confidence set with the highest accuracy is determined as the optimal confidence weight set. The values of this optimal confidence weight set are used to set the preset weight set, thus completing the initialization of the preset weight set. For example, Figure 14 Among them, the confidence weight set 2 {Chinese: 0.21; English: 0.19; Korean: 0.22; Japanese: 0.20; German: 0.18} corresponds to the highest accuracy (0.98). Therefore, the preset weight set is set to {Chinese: 0.21; English: 0.19; Korean: 0.22; Japanese: 0.20; German: 0.18}.
[0264] By employing the above-mentioned technical methods, when the language recognition model initially encounters a scene, a preset weight set is initialized based on environmental features and audio acquisition device features. This adjusts the recognition confidence of subsequent user-input speech, enabling the voice interaction system 100 to adapt to different scenarios and achieve language recognition with optimal accuracy, thereby improving the reliability of the recognition results. In other words, these technical methods can suppress the problem of "low reliability of recognition results due to the limited adaptability of the trained language recognition model to different scenarios."
[0265] In this embodiment, the preset weight set is initialized based on both environmental features and audio collector features. However, in other embodiments, the preset weight set may be initialized based on only one of the environmental features and audio collector features.
[0266] After the preset weight set is initialized, Figure 14In step S212, it is determined whether the preset weight of each language in the preset weight set is within the weight range. The weight range is preset, and the specific value can be determined by testing, which will be described later.
[0267] When the preset weight of each language in the preset weight set set according to the tentative environment data set is within the weight range, the confidence of the recognition result of the language recognition model in this environment (the environment features and the audio collector features) is higher. When the confidence weight set according to the tentative environment data set is not within the weight range (no in step S212), the confidence of the result of the language recognition model in this environment is lower. At this time, in the embodiment, as shown in steps S214, S217, etc., the language of the input speech of the user can be determined by the historical language record and the user specified language. Figure 14
[0268] Specifically, in step S214, it is determined whether there is a historical language record. When there is a historical language record, the input speech in the historical record is compared with the input speech of the user this time to obtain a feature similarity, so as to determine the language of the input speech of the user this time.
[0269] When there is no historical language record, in step S217, the voiceprint is queried to determine whether the user has specified a language. When there is a user specified language, the user specified language is used to determine the recognition language of the input speech of the user. The user specified language can be one or multiple. When there is only one user specified language, the user specified language is determined as the recognition language of the input speech. When there are multiple user specified languages, for example, the language with the highest frequency can be determined as the recognition language of the input speech. For example, when the user specified language queried is {Chinese: 3 times; English: 1 time; German: 1 time}, Chinese is determined as the recognition language of the input speech. When there is no user specified language, in step S219, the recognition result of the language recognition model is used to determine the language of the input speech.
[0270] In step S214, step S214 is executed before step S217, but the execution order of the process of determining the language by the historical language record and the process of determining the language by the user specified language is not limited. Figure 15
[0271] By using the above technical means, when the preset weight set is not within the weight range, the language of the input speech of the user is predicted according to the historical language record or the user specified language, so as to improve the reliability of the language prediction of the input speech of the voice interaction system 100.
[0272] Additionally, when the weight values of each language in the preset weight set set according to the simulated environment dataset are within the weight range ("Yes" in step S212), when user input speech is detected, in step S200, the language recognition model is used to identify the language of the input speech. In step S221, it is determined whether there is a recognition confidence score greater than the threshold λ in the multilingual recognition confidence score set obtained according to the language recognition model. If there is ("Yes" in step S221), in step S222, the user's identity is determined by voiceprint. As another embodiment, the user's identity can also be determined by face recognition or iris recognition. Then, in step S223, the user's historical language record and the current dialogue round language record are updated (i.e., the current language is added to the record). Then, in step S225, the multilingual recognition confidence score set is output as the language recognition result to the command parsing and execution module 150. Here, the current dialogue turn refers to one cycle of continuously listening to (receiving) the user's input voice, such as the period from the opening to the closing of a language recognition system or a voice interaction system.
[0273] When there is no recognition confidence score greater than the threshold λ in the multilingual recognition confidence score set ("No" in step S221), the language confidence score correction module 160 calls the audio feature-based adjustment module 162 or the video feature-based adjustment module 163 to correct the multilingual recognition confidence score set. In this embodiment, the audio feature-based adjustment module 162 is called first for correction. When there is no recognition confidence score greater than the threshold λ in the multilingual recognition confidence score set corrected by the audio feature-based adjustment module 162, the video feature-based adjustment module 163 is called for correction. In other embodiments, the video feature-based adjustment module 163 can be called first.
[0274] The following reference Figure 15 The processing performed by the audio feature adjustment module 162 in this embodiment will be described.
[0275] like Figure 15 As shown, in step S231, the user's identity is determined through voiceprint; in step S232, the user's historical language records are queried. If historical language records exist, the distribution of each language in the historical language records is calculated. For example, refer to... Figure 15In the right-hand section, when the historical language records are {Chinese: 8 entries; English: 1 entry; Korean: 0 entries; Japanese: 1 entry; German: 0 entries}, the distribution of each language is calculated based on these historical language records (which can also be called normalization according to weighted numerical forms) as {Chinese: 0.8; English: 0.1; Korean: 0.0; Japanese: 0.1; German: 0.0}. The time range for the historical language records can be freely set, such as the current dialogue round, a few days, several months, or longer. Here, as one approach, all historical language records at all time points and the historical language records of the current dialogue round (hereinafter referred to as the current dialogue round language records) can be pre-stored. The processing in step S236 below is performed separately based on both the historical language records at all time points and the historical language records of the current dialogue round. When the results obtained from the two conflict, the result obtained from the historical language records of the current dialogue round takes precedence, considering that the reliability of the result obtained from the current dialogue round language records is relatively higher. Alternatively, different weight values can be assigned to the two for the calculation in step S236 below.
[0276] Then, in step S236, the multilingual recognition confidence set is corrected and calculated using the distribution of each language in the historical language record, resulting in a corrected multilingual recognition confidence set (corresponding to the second recognition confidence set in this application). For example, refer to... Figure 14 In the right-hand part of the diagram, with the initial confidence set for multiple languages being {Chinese: 0.7; English: 0.1; Korean: 0.1; Japanese: 0.05; German: 0.05}, the pre-set confidence weights being {Chinese: 0.25; English: 0.25; Korean: 0.25; Japanese: 0.25; German: 0.25}, and the language distribution being {Chinese: 0.8; English: 0.1; Korean: 0.0; Japanese: 0.1; German: 0.0}, the calculated corrected recognition confidence set (after normalization) is {Chinese: 0.973; English: 0.017; Korean: 0.000; Japanese: 0.010; German: 0.000}.
[0277] Next, in step S237, it is determined whether there is a corrected confidence score greater than the threshold λ in the corrected multilingual recognition confidence score set. If there is ("Yes" in step S237), on the one hand, in step S239, the corrected multilingual recognition confidence score set is output as the language recognition result to the command parsing and execution module 150. In addition, the historical language record and the current dialogue round language record are updated (see [link to specific processing method]). Figure 15In steps S222 and S223), on the other hand, in step S238, a process for adjusting the confidence weights is performed. Specifically, adjustments are made based on the corrected set of recognition confidence scores, increasing the confidence weight of languages with recognition confidence scores greater than a threshold λ relative to the confidence weights of other languages. For example, referring to... Figure 16 The old confidence weight set {Chinese: 0.25; English: 0.25; Korean: 0.25; Japanese: 0.25; German: 0.25} is modified into a new confidence weight set (or modified weight set) {Chinese: 0.29; English: 0.24; Korean: 0.24; Japanese: 0.24; German: 0.24}.
[0278] In step S271, it is determined whether each corrected weight in the corrected weight set is within the weight range. If the determination result is within the weight range, in step S272, the preset weight set is updated with the values in the corrected confidence weight set for use by the language recognition module 120 in subsequent language recognition, and then the process ends. If the determination result is not within the weight range, in step S271, the preset weight set is not updated, and the process ends.
[0279] In addition, if the judgment result in step S235 is "no" or the judgment result in step S237 is "no", the language confidence correction module 160 calls the comprehensive adjustment module 164 for processing.
[0280] The following reference Figure 16 The processing of the comprehensive adjustment module 164 is explained.
[0281] like Figure 16 As shown, in step S251, based on the output of the speech recognition module 110, it is determined whether there is an ASR confidence level greater than the threshold λ. If so, in step S252, the comprehensive adjustment module 164 outputs the set of recognition confidence levels for multiple languages as the language recognition result to the command parsing and execution module 150. In this case, the command parsing and execution module 150 can determine the language corresponding to the ASR confidence level greater than the threshold λ as the language of the input speech (which can be called the recognition language), and determine the candidate command corresponding to that language as the target command to be executed.
[0282] When the result of the determination in step S251 is "No", based on the output of the semantic understanding module 140, it is determined whether there is an NLU confidence greater than the threshold λ. When there is an NLU confidence greater than the threshold λ, in step S254, the comprehensive adjustment module 164 outputs the set of recognition confidences of multiple languages as the language recognition result to the command analysis and execution module 150. In this case, the command analysis and execution module 150 determines the language corresponding to the NLU confidence greater than the threshold λ as the language of the input speech, and determines the candidate command corresponding to the language as the target command to be executed.
[0283] When the result of the determination in step S253 is "No", in step S256, it is determined whether there is a user specified language based on the voiceprint (user identity). The user specified language here is the kind of system language of the voice interaction system 100 set by the user. When there is a user specified language, the set of recognition confidences of multiple languages is calculated to correct the recognition confidence of the user specified language to be greater than the recognition confidences of other languages, so as to obtain the corrected set of recognition confidences of multiple languages (corresponding to the second recognition confidence in the present application). In addition, there can be multiple user specified languages (there are multiple languages in the history record of system languages stored in the database), for example, referring to the right part of FIG. 13, when the user specified language, i.e., the language set by the user, includes Chinese, English and German, the old set of recognition confidences {Chinese: 0.75; English: 0.12; Korean: 0.11; Japanese: 0.01; German: 0.01} is corrected to the new set of recognition confidences {Chinese: 0.95; English: 0.32; Korean: 0.11; Japanese: 0.01; German: 0.21}. Figure 14
[0284] Here, the user specified language is an example of the user operation record. As other examples, the language of the song played by the user history can also be given.
[0285] In addition, optionally, after the recognition language of the input speech is determined based on the ASR confidence or the NLU confidence, the set of preset weights of multiple languages can be updated for the language recognition of the input speech thereafter. The manner of updating is the same as described with reference to FIG. 12, which is not repeated here. Figure 10
[0286] Afterwards, in step S259, it is judged whether there is a recognition confidence greater than the threshold λ in the set of recognition confidences of multiple languages after the correction in step S258, and when there is, the set of recognition confidences of multiple languages after the correction is outputted to the command analysis and recognition module as the language recognition result in step S261. After that, in step S262, the set of pre-set confidence weights of multiple languages is adjusted. The method of adjustment is the same as the method described above, and will not be described again here. In addition, as described above, when each of the correction weights in the set of correction weights is within the weight range, the set of pre-set confidence weights of multiple languages is updated.
[0287] In addition, when the result of the judgment in step S256 is "No", i.e., there is no user-specified language, or when the result of the judgment in step S259 is "No", i.e., there is no recognition confidence greater than the threshold λ, it is judged in step S264 whether there is a user's historical language record according to the user's identity. When the result of the judgment is that there is a user's historical language record, the input speech of the user this time is compared with the input speech in the historical language record in step S256 to obtain a feature similarity, and the language most similar to the unknown input speech is found according to the feature similarity.
[0288] In addition, when the result of the judgment in step S264 is "No", i.e., there is no user's historical language record, the recognition language confidence of multiple users is directly outputted to the command analysis and execution module 150 by the comprehensive adjustment module 164 as the language recognition result. In this case, the command analysis and execution module 150 can consider that the language of the input speech cannot be recognized, and can feed back this matter to the user, for example, by playing the speech.
[0289] With the above embodiment, the set of recognition confidences of multiple languages is adjusted according to the user's features including the historical language record or the user-specified language, etc., so that the prediction accuracy of the voice interaction system 100 for the input speech can be improved, and the degree of trust of the user for the intelligence of the voice interaction system 100 can be improved.
[0290] In the above description, the "weight range" is mentioned, i.e., when the pre-set weight set is initialized by the real-time scene adaptation module 161, it is judged whether the initialized pre-set weight is within the weight range, and when the pre-set weight is updated by the audio feature-based adjustment module 162, the video feature-based adjustment module 163 and the comprehensive adjustment module 164, it is also judged whether the pre-set weight is within the weight range. It can be seen that the "weight range" reflects the robustness range possessed by the model itself.
[0291] In the embodiment, a method for setting the "weight range" is also provided, which is implemented, for example, in the test stage before the voice interaction system 100 is shipped, and in addition, can also be implemented in the offline inspection stage after the shipment.
[0292] The method mainly comprises the following steps:
[0293] ① Collecting language data sets data1, data2, …, data n under different scenarios, the content of the language data set is known in advance;
[0294] ② Inputting the language data set data i under a scenario into a language recognition model to test the language recognition model, and obtaining the optimal confidence weight (that is, the confidence weight corresponding to the highest recognition accuracy) t i of each language under the scenario (i∈[1,n]);
[0295] ③ Inputting all language data sets into the model to execute the above ②, so as to obtain the best confidence weight set T k of each language (k∈[1,m]) = {t 1k , t 2k , …, t nk};
[0296] ④ Obtaining the best confidence weight range F k of each language = [a k , b k ], wherein a k = min(T k ) and b k = max(T k ).
[0297] Note: n is the number of language data sets, and m is the number of languages.
[0298] Here, the language data sets data1, data2, …, data n correspond to the test data set in the present application.
[0299] For the above method, the present embodiment gives an implementation scheme as shown in Figure 10 , but is not limited to this scheme. The scheme will be described below with reference to . First, randomly initialize k different confidence weight sets, which correspond to the multiple test weight sets in the present application; then, input the language data set data i under a scenario into a language recognition model to obtain an output, and respectively multiply the output with the k confidence weight sets (the set here can be understood as a matrix), and then normalize to obtain the corrected language confidence set of multiple languages; then, according to the language confidence set and the known language data set data iThe content gets the accuracy of each multilingual confidence weight set (acc in the figure), so the confidence weight set with high accuracy (for example, {Chinese: 0.21; English: 0.19; Korean: 0.22; Japanese: 0.20; German: 0.18} in the figure) is the optimal confidence weight set for the language data set data i (i.e., the optimal confidence weight set of scenario i).
[0300] Finally, repeating the above processing for the language data set of each scenario, the confidence weight set of each single language can be obtained, and the confidence weight range of each single language can be obtained, for example, the confidence weight range of Chinese [c a ,c b ], the confidence weight range of English [e a ,e b ], the confidence weight range of Korean [h a ,h b ], the confidence weight range of Japanese [r a ,r b ], and the confidence weight range of German [d a ,d b ] shown in the figure.
[0301] By using the above technical means, before language recognition is performed by using the language recognition model, the language recognition model is tested by using a large amount of language data set, so as to set the weight range of the preset weight set of the multilingual language, that is, the robustness range of the language recognition model is specified, so that the language recognition model works in this range, thereby ensuring the reliability of the language recognition result.
[0302] Note that the above is only the preferred embodiment of the present application and the technical principle used. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and those skilled in the art can make various obvious changes, readjustments and substitutions without departing from the scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and all belong to the protection scope of the present application.
Claims
1. A speech processing method, characterized in that, include: Obtain the user's voice input information; Based on the input voice information, multiple first confidence levels corresponding to the input voice information are determined. The multiple first confidence levels correspond to multiple languages respectively. The multiple first confidence levels are determined by multiple initial confidence levels and multiple preset weights. The multiple first confidence levels are adjusted to multiple second confidence levels based on the user's user characteristics; the user characteristics include historical language records; When there is a second confidence level greater than the first threshold among the plurality of second confidence levels, the language corresponding to the second confidence level greater than the first threshold is determined as the language of the input voice information, and the plurality of preset weights are updated according to the plurality of second confidence levels; When there is no second confidence level greater than the first threshold among the plurality of second confidence levels, automatic speech recognition (ASR) is performed on the input speech information to determine the plurality of ASR confidence levels corresponding to the input speech information. When there is an ASR confidence level greater than the first threshold among the plurality of ASR confidence levels, the language corresponding to the ASR confidence level greater than the first threshold is determined as the language of the input speech information. When there is no ASR confidence greater than the first threshold among the multiple ASR confidences, Natural Language Understanding (NLU) is performed on the input speech information to determine multiple NLU confidences corresponding to the input speech information. When there is an NLU confidence greater than the first threshold among the multiple NLU confidences, the language corresponding to the NLU confidence greater than the first threshold is determined as the language of the input speech information. If none of the multiple NLU confidence scores are greater than the first threshold, output information indicating that speech content recognition has failed.
2. The speech processing method according to claim 1, characterized in that, The step of correcting the plurality of first confidence scores to a plurality of second confidence scores based on the user's user characteristics specifically includes: When the plurality of first confidence levels are less than a first threshold, the plurality of first confidence levels are corrected to the plurality of second confidence levels based on the user characteristics.
3. The speech processing method according to claim 1 or 2, characterized in that, The user characteristics include one or more of the user-specified languages.
4. The speech processing method according to claim 3, characterized in that, The historical language record and the user-specified language are obtained by querying the voiceprint features of the input voice information.
5. The speech processing method according to claim 1, 2, or 4, characterized in that, Also includes: The semantics of the input voice information are determined based on the input voice information and the language of the input voice information.
6. The speech processing method according to claim 1, 2, or 4, characterized in that, The multiple languages are pre-defined.
7. The speech processing method according to claim 1, 2, or 4, characterized in that, The plurality of first confidence levels are determined by a plurality of initial confidence levels and a plurality of preset weights; The voice processing method further includes setting multiple preset weights based on scene characteristics before acquiring the user's input voice information.
8. The speech processing method according to claim 7, characterized in that, The scene features include environmental features and / or audio acquisition device features.
9. The speech processing method according to claim 8, characterized in that, The environmental features include one or more of the following: environmental signal-to-noise ratio, power supply DC / AC information, or environmental vibration amplitude; the audio acquisition device features include microphone arrangement information.
10. The speech processing method according to claim 7, characterized in that, The step of setting the multiple preset weights based on scene characteristics specifically includes: Acquire the first pre-collected voice data and the first language information pre-recorded from the first voice data; The second voice data is determined based on the first voice data and the scene features; The second language information of the second voice data is determined based on the second voice data; The multiple preset weights are set based on the first language information and the second language information.
11. The speech processing method according to claim 10, characterized in that, The step of determining the second language information of the second speech data based on the second speech data specifically includes: Obtain multiple test weight reconfigurations, any one of which includes multiple test weights; Multiple pieces of second language information are determined based on the second voice data and the multiple test weight recombinations, and the multiple pieces of second language information correspond to the multiple test weight recombinations respectively; The step of setting the multiple preset weights based on the first language information and the second language information specifically includes: Based on the first language information and the plurality of second language information, determine multiple accuracy rates of the plurality of second language information; The multiple preset weights are set based on the test weight reorganization corresponding to the second language information with the highest accuracy.
12. The speech processing method according to claim 7, characterized in that, Setting the plurality of preset weights specifically includes: setting the plurality of preset weights within a weight range.
13. The speech processing method according to claim 1, characterized in that, The updating of the plurality of preset weights specifically includes: updating the plurality of preset weights within a weight range.
14. The speech processing method according to claim 12 or 13, characterized in that, The weight range is determined as follows: Acquire multiple pre-collected test speech data sets and pre-recorded first language information of the multiple test speech data sets, wherein any one of the multiple test speech data sets includes multiple test speech data sets; Obtain multiple test weight reconfigurations, any one of which includes multiple test weights; The weight range is determined based on the multiple test speech data sets, the first language information, and the multiple test weight sets.
15. A speech processing method, characterized in that, include: Obtain the user's voice input information; Based on the input voice information, multiple third confidence levels are determined corresponding to the input voice information. The multiple third confidence levels correspond to multiple languages respectively. The multiple third confidence levels are determined by multiple initial confidence levels and multiple preset weights. The multiple third confidence levels are corrected into multiple fourth confidence levels based on scene features; the scene features include environmental features and / or audio collector features; When there is a fourth confidence level greater than the first threshold among the plurality of fourth confidence levels, the language corresponding to the fourth confidence level greater than the first threshold is determined as the language of the input voice information, and the plurality of preset weights are updated according to the plurality of fourth confidence levels. When there is no fourth confidence level greater than the first threshold among the plurality of fourth confidence levels, automatic speech recognition (ASR) is performed on the input speech information to determine the plurality of ASR confidence levels corresponding to the input speech information. When there is an ASR confidence level greater than the first threshold among the plurality of ASR confidence levels, the language corresponding to the ASR confidence level greater than the first threshold is determined as the language of the input speech information. When there is no ASR confidence greater than the first threshold among the multiple ASR confidences, Natural Language Understanding (NLU) is performed on the input speech information to determine multiple NLU confidences corresponding to the input speech information. When there is an NLU confidence greater than the first threshold among the multiple NLU confidences, the language corresponding to the NLU confidence greater than the first threshold is determined as the language of the input speech information. If none of the multiple NLU confidence scores are greater than the first threshold, output information indicating that speech content recognition has failed.
16. The speech processing method according to claim 15, characterized in that, The environmental features include one or more of the following: environmental signal-to-noise ratio, power supply DC / AC information, or environmental vibration amplitude; the audio acquisition device features include microphone arrangement information.
17. The speech processing method according to any one of claims 15-16, characterized in that, The step of correcting the multiple third confidence levels to multiple fourth confidence levels based on scene features specifically includes: Multiple preset weights are set according to the scene characteristics; The plurality of third confidence levels are corrected to the plurality of fourth confidence levels based on the plurality of preset weights.
18. The speech processing method according to claim 17, characterized in that, The step of setting the multiple preset weights based on scene characteristics specifically includes: Acquire the first pre-collected voice data and the first language information pre-recorded from the first voice data; The second voice data is determined based on the first voice data and the scene features; The second language information of the second voice data is determined based on the second voice data; The multiple preset weights are set based on the first language information and the second language information.
19. The speech processing method according to claim 18, characterized in that, The step of determining the second language information of the second speech data based on the second speech data specifically includes: Obtain multiple test weight reconfigurations, wherein the test weight reconfigurations include multiple test weights; Multiple pieces of second language information are determined based on the second voice data and the multiple test weight recombinations, and the multiple pieces of second language information correspond to the multiple test weight recombinations respectively; The step of setting the multiple preset weights based on the first language information and the second language information specifically includes: Based on the first language information and the plurality of second language information, determine multiple accuracy rates of the plurality of second language information; The multiple preset weights are set based on the test weight reorganization corresponding to the second language information with the highest accuracy.
20. A voice processing device, characterized in that, Includes processing modules and transceiver modules. The transceiver module is used to acquire the user's input voice information; The processing module is used to determine multiple first confidence levels corresponding to the input voice information based on the input voice information. The multiple first confidence levels correspond to multiple languages respectively, and the multiple first confidence levels are determined by multiple initial confidence levels and multiple preset weights. The processing module is further configured to correct the plurality of first confidence scores to a plurality of second confidence scores based on the user's user characteristics, and determine the language of the input voice information based on the plurality of second confidence scores, wherein the user characteristics include historical language records; The processing module is further configured to, when there is a second confidence level greater than the first threshold among the plurality of second confidence levels, determine the language corresponding to the second confidence level greater than the first threshold as the language of the input voice information, and update the plurality of preset weights according to the plurality of second confidence levels; If none of the multiple second confidence levels is greater than the first threshold, perform Automatic Speech Recognition (ASR) on the input speech information to determine multiple ASR confidence levels corresponding to the input speech information. If any of the multiple ASR confidence levels is greater than the first threshold, determine the language corresponding to the ASR confidence level greater than the first threshold as the language of the input speech information. If none of the multiple ASR confidence levels is greater than the first threshold, perform Natural Language Understanding (NLU) on the input speech information to determine multiple NLU confidence levels corresponding to the input speech information. If any of the multiple NLU confidence levels is greater than the first threshold, determine the language corresponding to the NLU confidence level greater than the first threshold as the language of the input speech information. If none of the multiple NLU confidence levels is greater than the first threshold, output a message indicating speech content recognition failure.
21. The speech processing apparatus according to claim 20, characterized in that, The processing module is specifically used to, when the plurality of first confidence levels are less than a first threshold, correct the plurality of first confidence levels to the plurality of second confidence levels based on the user characteristics.
22. The speech processing apparatus according to claim 20 or 21, characterized in that, The user characteristics include one or more of the user-specified languages.
23. The speech processing apparatus according to claim 22, characterized in that, The historical language record and the user-specified language are obtained by querying the voiceprint features of the input voice information.
24. The speech processing apparatus according to claim 20, 21 or 23, characterized in that, The processing module is further configured to determine the semantics of the input voice information based on the input voice information and the language of the input voice information.
25. The speech processing apparatus according to claim 20, 21 or 23, characterized in that, The multiple languages are pre-defined.
26. The speech processing apparatus according to claim 20, 21 or 23, characterized in that, The plurality of first confidence levels are determined by a plurality of initial confidence levels and a plurality of preset weights; The processing module is also used to set the multiple preset weights according to scene characteristics before acquiring the user's input voice information.
27. The speech processing apparatus according to claim 26, characterized in that, The scene features include environmental features and / or audio acquisition device features.
28. The speech processing apparatus according to claim 27, characterized in that, The environmental features include one or more of the following: environmental signal-to-noise ratio, power supply DC / AC information, or environmental vibration amplitude; the audio acquisition device features include microphone arrangement information.
29. The speech processing apparatus according to claim 26, characterized in that, The processing module is specifically used to: acquire pre-collected first voice data and pre-recorded first language information of the first voice data; determine second voice data based on the first voice data and the scene features; determine second language information of the second voice data based on the second voice data; and set the multiple preset weights based on the first language information and the second language information.
30. The speech processing apparatus according to claim 29, characterized in that, The processing module is specifically used to: acquire multiple test weight reassemblies, any one of which includes multiple test weights; determine multiple second language information based on the second speech data and the multiple test weight reassemblies, the multiple second language information corresponding to the multiple test weight reassemblies respectively; determine multiple accuracy rates of the multiple second language information based on the first language information and the multiple second language information; and set the multiple preset weights based on the test weight reassembly corresponding to the second language information with the highest accuracy rate.
31. The speech processing apparatus according to claim 26, characterized in that, The processing module is specifically used to set the plurality of preset weights within a weight range.
32. The speech processing apparatus according to claim 20, characterized in that, The processing module is specifically used to update the plurality of preset weights within the weight range.
33. The speech processing apparatus according to claim 31 or 32, characterized in that, The weight range is determined as follows: Acquire multiple pre-collected test speech data sets and pre-recorded first language information of the multiple test speech data sets, wherein any one of the multiple test speech data sets includes multiple test speech data sets; Obtain multiple test weight reconfigurations, any one of which includes multiple test weights; The weight range is determined based on the multiple test speech data sets, the first language information, and the multiple test weight sets.
34. A voice processing device, characterized in that, Includes processing modules and transceiver modules. The transceiver module is used to acquire the user's input voice information; The processing module is used to determine multiple third confidence levels corresponding to the input voice information based on the input voice information. The multiple third confidence levels correspond to multiple languages respectively, and the multiple third confidence levels are determined by multiple initial confidence levels and multiple preset weights. The processing module is further configured to correct the plurality of third confidence scores to a plurality of fourth confidence scores based on scene features, and determine the language of the input speech information based on the plurality of fourth confidence scores, wherein the scene features include environmental features and / or audio collector features; The processing module is further configured to, when there is a fourth confidence level greater than the first threshold among the plurality of fourth confidence levels, determine the language corresponding to the fourth confidence level greater than the first threshold as the language of the input voice information, and update the plurality of preset weights according to the plurality of fourth confidence levels; When there is no fourth confidence level greater than the first threshold among the plurality of fourth confidence levels, automatic speech recognition (ASR) is performed on the input speech information to determine the plurality of ASR confidence levels corresponding to the input speech information. When there is an ASR confidence level greater than the first threshold among the plurality of ASR confidence levels, the language corresponding to the ASR confidence level greater than the first threshold is determined as the language of the input speech information. When there is no ASR confidence greater than the first threshold among the multiple ASR confidences, Natural Language Understanding (NLU) is performed on the input speech information to determine multiple NLU confidences corresponding to the input speech information. When there is an NLU confidence greater than the first threshold among the multiple NLU confidences, the language corresponding to the NLU confidence greater than the first threshold is determined as the language of the input speech information. If none of the multiple NLU confidence scores are greater than the first threshold, output information indicating that speech content recognition has failed.
35. The speech processing apparatus according to claim 34, characterized in that, The environmental features include one or more of the following: environmental signal-to-noise ratio, power supply DC / AC information, or environmental vibration amplitude; the audio acquisition device features include microphone arrangement information.
36. The speech processing apparatus according to any one of claims 34-35, characterized in that, The processing module is specifically used to set multiple preset weights according to the scene features, and to correct the multiple third confidence levels to the multiple fourth confidence levels according to the multiple preset weights.
37. The speech processing apparatus according to claim 36, characterized in that, The processing module is specifically used to: acquire pre-collected first voice data and pre-recorded first language information of the first voice data; determine second voice data based on the first voice data and the scene features; determine second language information of the second voice data based on the second voice data; and set the multiple preset weights based on the first language information and the second language information.
38. The speech processing apparatus according to claim 37, characterized in that, The processing module is specifically used to: acquire multiple test weight reassemblies, the test weight reassemblies including multiple test weights; determine multiple second language information based on the second speech data and the multiple test weight reassemblies, the multiple second language information corresponding to the multiple test weight reassemblies respectively; determine multiple accuracy rates of the multiple second language information based on the first language information and the multiple second language information; and set the multiple preset weights based on the test weight reassembly corresponding to the second language information with the highest accuracy rate.
39. A computing device, characterized in that, It includes a processor and a memory, the memory storing computer program instructions that, when executed by the processor, cause the processor to perform the method of any one of claims 1-19.
40. A computer-readable storage medium, characterized in that, The device stores computer program instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1-19.
41. A computer program product, characterized in that, It includes computer program instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1-19.
Citation Information
Patent Citations
Speech translation method and device
CN109522564A
Voice recognition method and device
CN110970018A