Language identification method, device, electronic device and medium

By obtaining the scores of wake-up words and control command audio data and using the language association weighting function to weight calculation, the problem of low language recognition accuracy in the prior art is solved, and higher language recognition accuracy and user experience are achieved.

CN115148188BActive Publication Date: 2025-05-16HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210564948.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-23
Publication Date
2025-05-16
Estimated Expiration
2042-05-23

AI Technical Summary

Technical Problem

In the prior art, the accuracy of language recognition is low, making it difficult to effectively identify the language type to which the voice data belongs in a multilingual mixed environment.

Method used

By obtaining the scores of the respective audio data corresponding to the wake-up word and the control command, and using the language association weighting function to weight calculation, the target score is determined to identify the language.

Benefits of technology

It improves the accuracy of language recognition and enhances the user's voice interaction experience in multilingual environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115148188B_ABST
    Figure CN115148188B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a language recognition method, device, electronic device and medium, and in particular to the field of speech recognition technology; wherein the method comprises: obtaining first scores of first audio data belonging to different candidate languages, and second scores of second audio data belonging to different candidate languages, the first audio data being audio data corresponding to a wake-up word of a target control device, and the second audio data being audio data corresponding to a control command of the target control device; determining target scores of the second audio data belonging to different candidate languages ​​according to a language association weight function of the first audio data and the second audio data, the first score and the second score; determining the target language corresponding to the second audio data based on the target score. The disclosed embodiment can perform language recognition on the audio data corresponding to the control command of the target control device, and the accuracy of language recognition is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of speech recognition technology, and in particular to a language recognition method, device, electronic device and medium. Background Art

[0002] Language recognition is the process of using a computer to identify the language to which the voice data belongs. In work and daily life, the phenomenon of speaking multiple languages ​​is becoming more and more common, which makes language recognition difficult. Especially in the process of far-field voice interaction, after the user successfully wakes up the device with a specific wake-up word, he or she can interact with the device according to the corresponding control command.

[0003] In the prior art, language recognition is mainly divided into three processes: first, feature extraction is performed based on the speech signal, then a language recognition model is established, and finally, the language of the test speech is determined. Traditional language recognition systems include language recognition based on machine learning such as Hidden Markov Model (HMM), language recognition based on phoneme recognizers, and language recognition based on underlying acoustic features. However, language recognition in the prior art is limited to the lack of acoustic research and modeling, resulting in the accuracy of language recognition needs to be improved. Summary of the invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a language recognition method, device, electronic device and medium, which can recognize the language of audio data corresponding to the control command of the target control device, and the accuracy of language recognition is relatively high.

[0005] In order to achieve the above objectives, the technical solutions provided by the embodiments of the present disclosure are as follows:

[0006] In a first aspect, the present disclosure provides a language recognition method, the method comprising:

[0007] Obtaining first scores of first audio data belonging to different candidate languages, and second scores of second audio data belonging to different candidate languages, wherein the first audio data is audio data corresponding to a wake-up word of a target control device, and the second audio data is audio data corresponding to a control command of the target control device;

[0008] Determining target scores for the second audio data to belong to different candidate languages, respectively, based on the language association weight function of the first audio data and the second audio data, the first score, and the second score;

[0009] A target language corresponding to the second audio data is determined based on the target score.

[0010] As an optional implementation of the embodiment of the present disclosure, determining target scores of the second audio data belonging to different candidate languages ​​respectively according to the language association weight function of the first audio data and the second audio data, the first score, and the second score includes:

[0011] Determining language association weight coefficients of the first audio data and the second audio data according to an end time corresponding to the first audio data, a start time corresponding to the second audio data, and the language association weight function;

[0012] Based on the language association weight coefficient, the first score, and the second score, target scores for the second audio data belonging to different candidate languages ​​are determined.

[0013] As an optional implementation of the embodiment of the present disclosure, before determining the language association weight coefficient of the first audio data and the second audio data according to the end time corresponding to the first audio data, the start time corresponding to the second audio data, and the language association weight function, the method further includes:

[0014] Based on an endpoint detection method, an end time corresponding to the first audio data and a start time corresponding to the second audio data are determined.

[0015] As an optional implementation of the embodiment of the present disclosure, determining the target scores of the second audio data belonging to different candidate languages ​​based on the language association weight coefficient, the first score, and the second score includes:

[0016] multiplying the first score by the language association weight coefficient to obtain a corresponding product;

[0017] The product is added to the second score to obtain target scores for the second audio data belonging to different candidate languages.

[0018] As an optional implementation of the embodiment of the present disclosure, obtaining first scores of the first audio data belonging to different candidate languages ​​includes:

[0019] Preprocessing the first audio data to obtain a processed audio signal;

[0020] Extracting features of the audio signal to obtain corresponding first Mel-frequency cepstral coefficient features;

[0021] The first Mel-frequency cepstral coefficient feature is input into a mixed language recognition model to obtain first scores of the first audio data belonging to different candidate languages.

[0022] As an optional implementation of the embodiment of the present disclosure, the mixed language recognition model is trained in the following way:

[0023] Acquire a training set sample, wherein the training set sample includes multiple audio samples in different languages;

[0024] Extracting features from the audio sample to obtain corresponding second Mel-frequency cepstral coefficient features;

[0025] The second Mel-frequency cepstral coefficient feature is input into a mixed-language language recognition model for training until the mixed-language language recognition model converges.

[0026] As an optional implementation manner of the embodiment of the present disclosure, before determining the target scores of the second audio data belonging to different candidate languages ​​respectively according to the language association weight function of the first audio data and the second audio data, the first score and the second score, the method further includes:

[0027] Determine the weight prediction order and weight prediction coefficient;

[0028] Based on the weight prediction order and the weight prediction coefficient, a language association weight function of the first audio data and the second audio data is determined.

[0029] In a second aspect, the present disclosure provides a language recognition device, the device comprising:

[0030] a prediction score determination module, used to obtain first scores of first audio data belonging to different candidate languages, and second scores of second audio data belonging to different candidate languages, wherein the first audio data is audio data corresponding to a wake-up word of a target control device, and the second audio data is audio data corresponding to a control command of the target control device;

[0031] a target score determination module, configured to determine target scores for the second audio data belonging to different candidate languages ​​according to a language association weight function of the first audio data and the second audio data, the first score, and the second score;

[0032] A target language determination module is used to determine a target language corresponding to the second audio data based on the target score.

[0033] As an optional implementation of the embodiment of the present disclosure, the target score determination module includes:

[0034] a coefficient determination unit, configured to determine a language association weight coefficient of the first audio data and the second audio data according to an end time corresponding to the first audio data, a start time corresponding to the second audio data, and the language association weight function;

[0035] A score determination unit is used to determine target scores for the second audio data belonging to different candidate languages ​​based on the language association weight coefficient, the first score and the second score.

[0036] As an optional implementation of the embodiment of the present disclosure, the device further includes: a time determination module, configured to:

[0037] Before determining the language association weight coefficients of the first audio data and the second audio data according to the end time corresponding to the first audio data, the start time corresponding to the second audio data and the language association weight function, the end time corresponding to the first audio data and the start time corresponding to the second audio data are determined based on an endpoint detection method.

[0038] As an optional implementation of the embodiment of the present disclosure, the score determination unit is specifically used to:

[0039] multiplying the first score by the language association weight coefficient to obtain a corresponding product;

[0040] The product is added to the second score to obtain target scores for the second audio data belonging to different candidate languages.

[0041] As an optional implementation of the embodiment of the present disclosure, the prediction score determination module includes:

[0042] A first score determination unit is configured to:

[0043] Preprocessing the first audio data to obtain a processed audio signal;

[0044] Extracting features of the audio signal to obtain corresponding first Mel-frequency cepstral coefficient features;

[0045] Inputting the first Mel-frequency cepstral coefficient feature into a mixed language recognition model to obtain first scores indicating that the first audio data belongs to different candidate languages;

[0046] A second score determination unit is configured to:

[0047] Second scores of the second audio data belonging to different candidate languages ​​are obtained.

[0048] As an optional implementation of the embodiment of the present disclosure, the mixed language recognition model is trained in the following way:

[0049] Acquire a training set sample, wherein the training set sample includes multiple audio samples in different languages;

[0050] Extracting features from the audio sample to obtain corresponding second Mel-frequency cepstral coefficient features;

[0051] The second Mel-frequency cepstral coefficient feature is input into a mixed-language language recognition model for training until the mixed-language language recognition model converges.

[0052] As an optional implementation of the embodiment of the present disclosure, the device further includes: a function determination module, configured to:

[0053] Before determining target scores for the second audio data to belong to different candidate languages ​​respectively based on the language association weight function of the first audio data and the second audio data, the first score, and the second score, determining a weight prediction order and a weight prediction coefficient;

[0054] Based on the weight prediction order and the weight prediction coefficient, a language association weight function of the first audio data and the second audio data is determined.

[0055] In a third aspect, the present disclosure further provides an electronic device, including:

[0056] one or more processors;

[0057] a storage device for storing one or more programs,

[0058] When the one or more programs are executed by the one or more processors, the one or more processors implement any one of the language recognition methods described in the embodiments of the present disclosure.

[0059] In a fourth aspect, the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any one of the language recognition methods described in the embodiments of the present disclosure.

[0060] The technical solution provided by the embodiments of the present disclosure has the following advantages over the prior art: first, first scores of the first audio data belonging to different candidate languages ​​and second scores of the second audio data belonging to different candidate languages ​​are obtained, the first audio data is the audio data corresponding to the wake-up word of the target control device, and the second audio data is the audio data corresponding to the control command of the target control device, and then the target scores of the second audio data belonging to different candidate languages ​​are determined according to the language association weight function of the first audio data and the second audio data, the first score and the second score, and finally the target language corresponding to the second audio data is determined based on the target score. Through the above method, the language of the audio data corresponding to the control command of the target control device can be recognized, and the accuracy of language recognition is high, which is conducive to improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0062] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0063] Figure 1 A schematic diagram of an application scenario of a language recognition method in an embodiment of the present disclosure;

[0064] Figure 2A A hardware configuration block diagram of an electronic device according to one or more embodiments of the present disclosure;

[0065] Figure 2B A schematic diagram of software configuration of an electronic device according to one or more embodiments of the present disclosure;

[0066] Figure 2C A schematic diagram showing an icon control interface of an application included in a control device according to one or more embodiments of the present disclosure;

[0067] Figure 3A A flow chart of a language recognition method provided in an embodiment of the present disclosure;

[0068] Figure 3B A schematic diagram of the principle of a language recognition method provided by an embodiment of the present disclosure;

[0069] Figure 4A A flowchart of another language recognition method provided by an embodiment of the present disclosure;

[0070] Figure 4B A schematic diagram of another language recognition method provided by an embodiment of the present disclosure;

[0071] Figure 4C A schematic diagram of a first audio data and a second audio data having no time difference provided by an embodiment of the present disclosure;

[0072] Figure 4D A schematic diagram of a time difference between first audio data and second audio data provided by an embodiment of the present disclosure;

[0073] Figure 5A A flowchart of another language recognition method provided by an embodiment of the present disclosure;

[0074] Figure 5B A schematic diagram of the principle of determining the end time corresponding to the first audio data provided by an embodiment of the present disclosure;

[0075] Figure 5C A schematic diagram of an endpoint detection method based on short-time average amplitude and zero-crossing rate provided in an embodiment of the present disclosure;

[0076] Fig. 6A A schematic diagram of a principle for determining a first score provided in an embodiment of the present disclosure;

[0077] Figure 6B A schematic diagram of the training process of the mixed language recognition model provided in the embodiment of the present disclosure;

[0078] Fig. 7A A schematic diagram of the principle of determining a language association weight function in an embodiment of the present disclosure;

[0079] Figure 7B A schematic diagram of a language association weight function in an embodiment of the present disclosure;

[0080] Fig. 8A is a structural diagram of a language recognition device provided by an embodiment of the present disclosure;

[0081] Figure 8B is a schematic diagram of the structure of a target score determination module in the language recognition device according to an embodiment of the present disclosure;

[0082] Fig. 9 It is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0083] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0084] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0085] The terms "first" and "second" in the present disclosure are used to distinguish different objects rather than to describe a specific order of objects. For example, a first score and a second score are used to distinguish different scores rather than to describe a specific order of scores.

[0086] With the continuous development of science and technology, various control devices are becoming more and more intelligent and can interact with users through voice, which has brought great convenience to people's lives. Language recognition is the process of analyzing the audio obtained by the control device and determining the language type according to the characteristics of the audio. The development of automatic language recognition has solved the problem of large-scale cross-language speech recognition to a certain extent. On the one hand, for interaction objects using different languages, both parties have a preliminary judgment on languages ​​that are not native to them. On the other hand, for interaction objects that need to use the same language, the control device or both parties need to use the language type interface to perform speech synthesis broadcast and display feedback in the corresponding language. With the continuous development of intelligent voice, people have more and more demands for various voice interaction scenarios, and the requirements are getting higher and higher. The accuracy of automatic language recognition is becoming more and more important in the interaction. Therefore, language recognition has important research value. Especially in the far-field voice interaction process between the user and the control device, the control device can be successfully awakened by the audio data corresponding to the wake-up word of the control device, and then the control command can be issued to realize the corresponding function.

[0087] In the prior art, language recognition of audio data corresponding to control commands of control devices is usually performed through a traditional language recognition system. However, this method is limited by the shortcomings of acoustic research and modeling, resulting in the accuracy of language recognition needs to be improved. In the case of inaccurate language recognition, the user experience will be reduced.

[0088] It can be seen from the above that the accuracy of the existing language recognition method is not high. Therefore, a language recognition method with higher recognition accuracy is needed.

[0089] Figure 1 FIG. 1 is a schematic diagram of an application scenario of a language recognition method in an embodiment of the present disclosure. Figure 1As shown, assuming that the control devices in the smart home scenario include a smart speaker 100, a smart washing machine 101 and a smart display device 102, and assuming that the language of the audio data corresponding to the control command of a certain control device is to be recognized, the first score of the audio data corresponding to the wake-up word of the control device belonging to different candidate languages ​​and the second score of the audio data corresponding to the control command of the control device belonging to different candidate languages ​​can be obtained first, and then the target scores of the audio data corresponding to the control command belonging to different candidate languages ​​are determined according to the language association weight function of the audio data corresponding to the wake-up word and the audio data corresponding to the control command, the first score and the second score. Finally, based on the target scores, the candidate language corresponding to the target score with the highest score in the target scores is determined as the target language to which the audio data corresponding to the control command of the control device belongs.

[0090] In the above process, two parameters, namely the first score corresponding to the audio data of the wake-up word and the language association weight function of the audio data corresponding to the wake-up word and the audio data corresponding to the control command, are added to the language recognition process. Since the wake-up word is usually easy to recognize and the language between the wake-up word and the control command is associated, the accuracy of language recognition can be effectively improved.

[0091] It should be noted that the smart home scene is one of the application scenes, and this embodiment does not impose any specific restrictions on this. The smart home scene may include a variety of control devices. Figure 1 This is just an example description and does not impose any specific limitation on the type and number of control devices.

[0092] The language recognition method provided by the embodiments of the present disclosure may be implemented based on an electronic device or a functional module or functional entity in the electronic device.

[0093] The electronic device may be a personal computer (PC), a server, a mobile phone, a tablet computer, a notebook computer, a mainframe computer, etc., which is not specifically limited in the embodiments of the present disclosure.

[0094] For example, Figure 2A FIG. 1 is a hardware configuration block diagram of an electronic device according to one or more embodiments of the present disclosure. Figure 2AAs shown, the electronic device includes: a tuner 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and at least one of a user interface 280. Among them, the controller 250 includes a central processing unit, a video processor, an audio processor, a graphics processor, a RAM, a ROM, and a first interface to an nth interface for input / output. The display 260 can be at least one of a liquid crystal display, an OLED display, a touch display, and a projection display, and can also be a projection device and a projection screen. The tuner 210 receives broadcast television signals by wired or wireless reception, and demodulates audio and video signals from multiple wireless or wired broadcast television signals, such as and EPG audio and video data signals. The communicator 220 is a component for communicating with an external device or server according to various communication protocol types. For example: the communicator can include other network communication protocol chips or near field communication protocol chips such as Wifi modules, Bluetooth modules, wired Ethernet modules, and at least one of infrared receivers. The electronic device can establish the transmission and reception of control signals and data signals with the server 203 or the local control device 205 through the communicator 220. The detector 230 is used to collect signals of the external environment or the external interaction. The controller 250 and the tuner-demodulator 210 can be located in different split devices, that is, the tuner-demodulator 210 can also be in an external device of the main device where the controller 250 is located, such as an external set-top box. The user interface 280 can be used to receive control signals of a control device (such as an infrared remote controller, etc.).

[0095] In some embodiments, the controller 250 controls the operation of the electronic device and responds to the user's operation through various software control programs stored in the memory. The controller 250 controls the overall operation of the electronic device. The user can input a user command in a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input command through the graphical user interface (GUI). Alternatively, the user can input a user command by inputting a specific sound or gesture, and the user input interface recognizes the sound or gesture through a sensor to receive the user input command.

[0096] In some embodiments, "user interface" is a medium interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The commonly used form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed in a graphical manner. It can be an interface element such as an icon, window, and control displayed on the display screen of an electronic device, where the control can include at least one of the visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, widgets, etc.

[0097] Figure 2B FIG. 1 is a schematic diagram of software configuration of an electronic device according to one or more embodiments of the present disclosure, such as Figure 2B As shown in the figure, the system is divided into four layers, from top to bottom, namely the application layer (Applications) layer (referred to as "application layer"), the application framework layer (Application Framework) layer (referred to as "framework layer"), the Android runtime (Android runtime) and system library layer (referred to as "system runtime library layer"), and the kernel layer.

[0098] In some embodiments, at least one application is running in the application layer, and these applications may be window programs, system settings programs, clock programs, etc. provided by the operating system, or applications developed by third-party developers. In specific implementations, the applications in the application layer include but are not limited to the above examples.

[0099] In some embodiments, the system runtime layer provides support for the upper layer, namely the framework layer. When the framework layer is used, the Android operating system will run the C / C++ library contained in the system runtime layer to implement the functions to be implemented by the framework layer.

[0100] In some embodiments, the kernel layer is a layer between hardware and software, and includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0101] Figure 2C The following is a schematic diagram showing an icon control interface of an application program included in a control device (mainly a smart playback device, such as a smart TV, a digital cinema system, or an audio and video server, etc.) according to one or more embodiments of the present disclosure. Figure 2CAs shown in , the application layer includes at least one application that can display a corresponding icon control in the display, such as: a live TV application icon control, a video on demand VOD application icon control, a media center application icon control, an application center icon control, a game application icon control, etc. The live TV application can provide live TV through different signal sources. The video on demand VOD application can provide videos from different storage sources. Unlike the live TV application, the video on demand provides video display from certain storage sources. The media center application can provide applications for playing various multimedia content. The application center can provide storage for various applications.

[0102] The language recognition method provided in the embodiment of the present application can be implemented based on the above-mentioned electronic device.

[0103] The language recognition process provided by the embodiment of the present disclosure first obtains first scores of first audio data belonging to different candidate languages ​​and second scores of second audio data belonging to different candidate languages, the first audio data is audio data corresponding to the wake-up word of the target control device, and the second audio data is audio data corresponding to the control command of the target control device, and then determines the target scores of the second audio data belonging to different candidate languages ​​according to the language association weight function of the first audio data and the second audio data, the first score and the second score, and finally determines the target language corresponding to the second audio data based on the target score. Through the above method, the language of the audio data corresponding to the control command of the target control device can be recognized, and the accuracy of language recognition is high, which is beneficial to improving the user experience.

[0104] In order to explain this scheme in more detail, the following will be combined in an exemplary manner Figure 3A To explain, it is understandable that Figure 3A The steps involved may include more steps or fewer steps in actual implementation, and the order of these steps may also be different, so as to implement the language recognition method provided in the embodiment of the present application.

[0105] Figure 3A A flow chart of a language recognition method provided in an embodiment of the present disclosure is provided. Figure 3B A schematic diagram of the principle of a language recognition method provided in an embodiment of the present disclosure. This embodiment can be applied to the case where the language of the audio data corresponding to the control command of the control device is recognized. The method of this embodiment can be executed by a language recognition device, which can be implemented in hardware and / or software and can be configured in an electronic device.

[0106] like Figure 3A As shown, the method specifically comprises the following steps:

[0107] S310, obtaining first scores of the first audio data belonging to different candidate languages, and second scores of the second audio data belonging to different candidate languages.

[0108] Among them, the first audio data is the audio data corresponding to the wake-up word of the target control device. The wake-up word can be the identification information of the target control device, such as the name of the control device. The second audio data is the audio data corresponding to the control command of the target control device. The control command can be the user's control intention for the target control device, that is, what the target control device wants to do, such as playing XX music, playing XX TV series, or turning on the switch of XX control device. The target control device can be any control device that can interact with the user by voice. The candidate language can be any language, for example, it can be a language of different countries such as Chinese, English or German, and can also be Mandarin, dialects and other languages. The first score is used to measure the possibility that the first audio data belongs to a candidate language. The second score is used to measure the possibility that the second audio data belongs to a candidate language. The number of the first score is the same as the number of candidate languages. Preferably, the number of the second score is the same as the number of the first score, and the candidate language corresponding to the first score is also the same as the candidate language corresponding to the second score.

[0109] In order to perform language recognition on the audio data corresponding to the control command of the target control device, in this embodiment, first scores of the first audio data belonging to different candidate languages ​​and second scores of the second audio data belonging to different candidate languages ​​are first obtained. Specifically, the first score and the second score can be obtained by a corresponding language recognition method, for example, language recognition can be realized by using a recurrent neural network, a naive Bayes classification method or a multi-class logistic regression method, etc. This embodiment does not specifically limit the language recognition method.

[0110] S320: Determine target scores for the second audio data belonging to different candidate languages ​​according to the language association weight function of the first audio data and the second audio data, the first score, and the second score.

[0111] The language association weight function may be a pre-set expression characterizing the language association relationship between the first audio data and the second audio data, or may be determined according to specific circumstances, which is not specifically limited in this embodiment.

[0112] After obtaining the first score and the first score, the first score of the first audio data belonging to a candidate language, the second score of the second audio data belonging to the same candidate language, and the language association weight function are weighted and summed through the corresponding weighted summation method, so that the target score of the second audio data belonging to different candidate languages ​​can be determined, and the target score is related to the first score and the second score.

[0113] Exemplarily, assuming that the first score of the first audio data belonging to candidate language 1 is A1, and the second score of the second audio data belonging to candidate language 1 is B1; the first score of the first audio data belonging to candidate language 2 is A2, and the second score of the second audio data belonging to candidate language 2 is B2; the first score of the first audio data belonging to candidate language 3 is A3, and the second score of the second audio data belonging to candidate language 3 is B3, then according to the language association weight function, the first score and the corresponding second score, it can be calculated that the target score of the second audio data belonging to candidate language 1 is C1, the target score of the second audio data belonging to candidate language 2 is C2, and the target score of the second audio data belonging to candidate language 3 is C3.

[0114] S330: Determine a target language corresponding to the second audio data based on the target score.

[0115] The target language is the language to which the second audio data belongs.

[0116] Since the first score and the second score correspond to different candidate languages, there may be multiple first scores and second scores, and there may also be multiple target scores. Therefore, after obtaining the target scores, the candidate language corresponding to the target score with the highest score among the target scores is the target language corresponding to the second audio data.

[0117] Exemplarily, assuming that the target score of the second audio data belonging to candidate language 1 is C1, the target score of the second audio data belonging to candidate language 2 is C2, and the target score of the second audio data belonging to candidate language 3 is C3, compare the sizes of C1, C2 and C3, assuming that C3 is the largest, then candidate language 3 corresponding to C3 is the target language corresponding to the second audio data.

[0118] In this embodiment, through the above S310-S330, the language of the audio data corresponding to the control command of the target control device can be recognized. Since the wake-up word is usually easy to recognize and the language between the wake-up word and the control command is related, the accuracy of language recognition and the user experience can be effectively improved.

[0119] Figure 4A A flowchart of another language recognition method provided by an embodiment of the present disclosure is provided. Figure 4B A schematic diagram of another language recognition method provided by an embodiment of the present disclosure. This embodiment is a further expansion and optimization based on the above embodiment. Optionally, this embodiment mainly describes the process of determining the target scores of the second audio data belonging to different candidate languages.

[0120] like Figure 4A As shown, the method specifically comprises the following steps:

[0121] S410, obtaining first scores of the first audio data belonging to different candidate languages, and second scores of the second audio data belonging to different candidate languages.

[0122] S420: Determine language association weight coefficients of the first audio data and the second audio data according to the end time corresponding to the first audio data, the start time corresponding to the second audio data, and the language association weight function.

[0123] The end time can be understood as the end time of the voice signal of the first audio data, and the start time can be understood as the start time of the voice signal of the second audio data.

[0124] Normally, according to the user's objective habits in language expression, the language used in the wake-up word and the control command has a certain degree of correlation, and the language association weight function is used to characterize the correlation. When the user and the control device perform voice interaction, the time interval (also called time difference) between the wake-up word and the control command is usually relatively short. This correlation decreases as the time interval between the two increases, and reaches the lowest threshold time after the control device wakes up. Therefore, the language association weight function is related to the time interval between the wake-up word and the control command. The threshold time can be a set value or it can be determined according to the specific situation. This embodiment does not limit this.

[0125] Since the language association weight function is a function related to the time interval between the wake-up word and the control command, the time interval between the two can be determined based on the end time corresponding to the first audio data and the start time corresponding to the second audio data. By substituting the time interval into the language association weight function, the language association weight coefficient of the first audio data and the second audio data can be determined.

[0126] S430: Determine target scores for the second audio data belonging to different candidate languages ​​based on the language association weight coefficient, the first score, and the second score.

[0127] After obtaining the language association weight coefficients of the first audio data and the second audio data, the target scores of the second audio data belonging to different candidate languages ​​can be calculated based on the language association weight coefficients, the first score and the corresponding second score.

[0128] It should be noted that: usually the first score, the second score and the target score are for the same candidate language, that is, based on the first score, the second score and the language association weight coefficient of the same candidate language, the target score of the second audio data belonging to the same candidate language is determined.

[0129] S440: Determine a target language corresponding to the second audio data based on the target score.

[0130] In this embodiment, determining the target score by the above method is simple and efficient, which is conducive to the subsequent determination of the target language corresponding to the second audio data.

[0131] In some embodiments, optionally, determining target scores of the second audio data belonging to different candidate languages ​​based on the language association weight coefficient, the first score, and the second score includes:

[0132] multiplying the first score by the language association weight coefficient to obtain a corresponding product;

[0133] The product is added to the second score to obtain target scores for the second audio data belonging to different candidate languages.

[0134] For example, assuming that the first score of the first audio data belonging to candidate language 1 is A1, the second score of the second audio data belonging to candidate language 1 is B1, and the language association weight coefficient is a, then according to the above method, the target score y1 of the second audio data belonging to candidate language 1 can be determined as:

[0135] y1=A1*a+B1 (1)

[0136] In this embodiment, the above method can associate the language corresponding to the wake-up word with the language corresponding to the control command, which is beneficial to improving the accuracy of language recognition.

[0137] For example, Figure 4C A schematic diagram of a first audio data and a second audio data having no time difference provided by an embodiment of the present disclosure. Figure 4C As shown, there is no time difference between the first audio data and the second audio data, that is, the language association weight coefficient of the first audio data and the second audio data is relatively large.

[0138] For example, Figure 4D A schematic diagram of a time difference between first audio data and second audio data provided by an embodiment of the present disclosure. Figure 4D As shown: there is a time difference between the first audio data and the second audio data, and as the time difference increases, the language association weight coefficient of the first audio data and the second audio data will decrease.

[0139] Figure 5A A flow chart of another language recognition method provided by the embodiment of the present disclosure. This embodiment is further expanded and optimized on the basis of the above embodiment. Optionally, this embodiment mainly describes the process of determining the end time corresponding to the first audio data and the start time corresponding to the second audio data.

[0140] like Figure 5A As shown, the method specifically comprises the following steps:

[0141] S510, obtaining first scores of the first audio data belonging to different candidate languages, and second scores of the second audio data belonging to different candidate languages.

[0142] S520: Determine an end time corresponding to the first audio data and a start time corresponding to the second audio data based on an endpoint detection method.

[0143] Among them, the endpoint detection method can distinguish the start time and end time of the speech signal.

[0144] There are many endpoint detection methods. This embodiment mainly describes a dual-threshold endpoint detection method based on short-time average amplitude and zero-crossing rate. The process is as follows:

[0145] 1. Preprocessing the speech signal to be detected, wherein the preprocessing may include: quantization, sampling, pre-filtering, pre-emphasis, and frame windowing processing, etc. The preprocessing process is a commonly used signal processing method and will not be described in detail here;

[0146] 2. Draw the short-time average amplitude curve and zero-crossing rate curve corresponding to the voice signal to be detected, and set a high threshold value (A H ) and the lower threshold (A L ), set a high threshold value for the zero crossing rate (B H ) and the lower threshold (B L ), if either the short-time average amplitude or the zero-crossing rate exceeds the corresponding high threshold value, the audio signal is judged to be a speech part, and the initial starting position and the initial ending position of the speech signal to be detected can be determined according to the time corresponding to the time when the corresponding high threshold value is exceeded;

[0147] 3. At the endpoints where the initial starting position and the initial ending position are determined, the calculation is performed frame by frame towards the forward and backward sides respectively. If any one of the short-time average amplitude and the zero-crossing rate exceeds the corresponding lower threshold value, the frame signal is judged to be a speech signal, and the other frame signals are judged to be silent segments. The endpoint detection is completed, and the starting position and the ending position of the speech signal to be detected can be determined.

[0148] Based on the above endpoint detection method, the first audio data and the second audio data are respectively used as voice signals to be detected, and the end time corresponding to the first audio data and the start time corresponding to the second audio data can be determined.

[0149] S530: Determine language association weight coefficients of the first audio data and the second audio data according to the end time corresponding to the first audio data, the start time corresponding to the second audio data, and the language association weight function.

[0150] S540: Determine target scores for the second audio data belonging to different candidate languages ​​based on the language association weight coefficient, the first score, and the second score.

[0151] S550: Determine a target language corresponding to the second audio data based on the target score.

[0152] For example, Figure 5B A schematic diagram of the principle of determining the end time corresponding to the first audio data provided by an embodiment of the present disclosure, Figure 5B The corresponding steps have been described in the above embodiments, and will not be described again here to avoid repetition.

[0153] For example, Figure 5C A schematic diagram of an endpoint detection method based on short-time average amplitude and zero-crossing rate provided in an embodiment of the present disclosure, such as Figure 5C As shown, through the above endpoint detection method, it can be determined that the initial starting position corresponding to the voice signal to be detected is h1, the initial ending position is h2, the starting position of the voice signal to be detected is h3, and the ending position is h4.

[0154] In some embodiments, optionally, obtaining first scores of the first audio data belonging to different candidate languages ​​includes:

[0155] Preprocessing the first audio data to obtain a processed audio signal;

[0156] Extracting features of the audio signal to obtain corresponding first Mel-frequency cepstral coefficient features;

[0157] The first Mel-frequency cepstral coefficient feature is input into a mixed language recognition model to obtain first scores of the first audio data belonging to different candidate languages.

[0158] Among them, the mixed language language recognition model can be a Gaussian mixed language wake-up language recognition model that can recognize wake-up words in different languages. The model can calculate the likelihood scores that the first audio data belongs to different candidate languages, that is, the first scores.

[0159] Specifically, the first audio data is preprocessed, that is, the first audio data is quantized, sampled, pre-filtered, pre-emphasized, and framed and windowed to obtain a processed audio signal; the audio signal is feature extracted, that is, first, the audio signal is fast Fourier transformed to obtain a corresponding spectrum signal, then the spectrum signal is passed through a Mel filter group to obtain a Mel spectrum, and finally, the Mel spectrum is cepstrum analyzed to obtain a corresponding first Mel Frequency Cepstrum Coefficient (MFCC) feature; the first MFCC feature is input into a mixed language recognition model to obtain first scores of the first audio data belonging to different candidate languages.

[0160] For example, Fig. 6A A schematic diagram of a principle for determining a first score provided in an embodiment of the present disclosure, Fig. 6A The corresponding steps have been described in the above embodiments, and will not be described again here to avoid repetition.

[0161] In this embodiment, the above method can determine the first scores of the first audio data belonging to different candidate languages, which is convenient for subsequent determination of the target score.

[0162] In some embodiments, optionally, obtaining second scores of the second audio data belonging to different candidate languages ​​includes:

[0163] Preprocessing the first audio data to obtain a processed target audio signal;

[0164] The target audio signal is input into a preset language recognition model to obtain second scores indicating that the second audio data belongs to different candidate languages.

[0165] The preset language recognition model may be a support vector machine (SVM) multi-classification confidence scoring model, or other models, which is not limited in this embodiment.

[0166] In this embodiment, the above method can be used to determine the second scores of the second audio data belonging to different candidate languages, which is convenient for subsequent determination of the target score.

[0167] In some embodiments, optionally, the mixed language recognition model is trained in the following manner:

[0168] Acquire a training set sample, wherein the training set sample includes multiple audio samples in different languages;

[0169] Extracting features from the audio sample to obtain corresponding second Mel-frequency cepstral coefficient features;

[0170] The second Mel-frequency cepstral coefficient feature is input into a mixed-language language recognition model for training until the mixed-language language recognition model converges.

[0171] Among them, the training set samples can be understood as a large number of pre-collected audio samples in different languages ​​corresponding to the wake-up words of different control devices, such as Chinese audio samples, English audio samples and German audio samples corresponding to the wake-up words of a certain control device, among which the Chinese audio samples can include Mandarin audio samples and dialect audio samples, etc.

[0172] A training set sample is obtained, and the audio samples of multiple different languages ​​contained in the training set sample are preprocessed and feature extracted (as described in the above embodiment), so as to obtain the second MFCC feature corresponding to each audio sample, and the second MFCC feature is input into the mixed language recognition model for training until the accuracy of the model meets the requirements, and it is determined that the mixed language recognition model converges.

[0173] Figure 6B A schematic diagram of the training process of the mixed language recognition model provided in the embodiment of the present disclosure is shown in FIG. Figure 6B As shown, the figure takes audio sample 1, audio sample 2 and audio sample 3 as examples for explanation, and these three audio samples belong to different languages ​​respectively, but the audio sample may include multiple ones, which is not limited in this embodiment. Feature 1 is obtained after feature extraction of audio sample 1, feature 2 is obtained after feature extraction of audio sample 2, and feature 3 is obtained after feature extraction of audio sample 3. The specific model training process has been described in the above embodiment, and will not be repeated here to avoid repetition.

[0174] In this embodiment, the mixed language recognition model is trained by the above method. Since the wake-up word is simple and contains fewer features, the above model training process is fast and efficient.

[0175] In some embodiments, optionally, before determining the target scores of the second audio data belonging to different candidate languages ​​respectively according to the language association weight function of the first audio data and the second audio data, the first score and the second score, the method further includes:

[0176] Determine the weight prediction order and weight prediction coefficient;

[0177] Based on the weight prediction order and the weight prediction coefficient, a language association weight function of the first audio data and the second audio data is determined.

[0178] Among them, the larger the value of the weight prediction order, the more accurate the language association weight function. Correspondingly, as the weight prediction order increases, the amount of calculation will also increase. Therefore, the weight prediction order needs to be an appropriate value, which can usually be determined based on simulation experimental methods.

[0179] Specifically, the weight prediction order and the weight prediction coefficient are determined by a corresponding simulation experiment method, and then based on the weight prediction order and the weight prediction coefficient, the language association weight function of the first audio data and the second audio data can be determined.

[0180] Exemplarily, the language association weight function can be expressed by the following formula:

[0181]

[0182] Among them, a i represents the weight prediction coefficient, n represents the weight prediction order (i.e. the highest order), t represents the independent variable (i.e. the time difference), Y(t) represents the language association weight function, and n is a positive integer.

[0183] For example, Fig. 7A FIG. 1 is a schematic diagram of a principle for determining a language association weight function in an embodiment of the present disclosure. Fig. 7A The corresponding steps have been described in the above embodiments, and will not be described again here to avoid repetition.

[0184] In this embodiment, the language association weight function is determined by the above method, which can well characterize the language association relationship between the first audio data and the second audio data, and is beneficial to the subsequent language recognition process.

[0185] For example, Figure 7B Schematic diagram of a language association weight function in the embodiment of the present disclosure. Taking the weight prediction order as 1, the time difference between the start time corresponding to the second sample and the end time corresponding to the first sample as 0 seconds, the association weight coefficient as b, and the time difference between the start time corresponding to the second sample and the end time corresponding to the first sample as a seconds, the association weight coefficient as c as an example, the language association weight function Y(t) can be drawn, as shown in Figure 7B The straight line corresponding to the first order in ; if the weight prediction order is multi-order, the language association weight function Y(t) drawn is as follows Figure 7B The corresponding curves of the multiple orders are shown in .

[0186] The first sample is an audio sample of a wake-up word of a control device, and the second sample is an audio sample of a control command of the control device.

[0187] Fig. 8A: is a schematic diagram of the structure of a language recognition device provided by an embodiment of the present disclosure. The device is configured in an electronic device and can implement the language recognition method described in any embodiment of the present application. The device specifically includes the following:

[0188] A prediction score determination module 801 is used to obtain first scores of first audio data belonging to different candidate languages, and second scores of second audio data belonging to different candidate languages, wherein the first audio data is audio data corresponding to a wake-up word of a target control device, and the second audio data is audio data corresponding to a control command of the target control device;

[0189] A target score determination module 802, configured to determine target scores for the second audio data belonging to different candidate languages ​​according to the language association weight function of the first audio data and the second audio data, the first score, and the second score;

[0190] The target language determination module 803 is configured to determine the target language corresponding to the second audio data based on the target score.

[0191] Figure 8B is a schematic diagram of the structure of a target score determination module in the language recognition device according to an embodiment of the present disclosure. Figure 8B As shown, the target score determination module 802 includes:

[0192] A coefficient determination unit 8021 is used to determine the language association weight coefficient of the first audio data and the second audio data according to the end time corresponding to the first audio data, the start time corresponding to the second audio data and the language association weight function;

[0193] The score determination unit 8022 is used to determine target scores for the second audio data belonging to different candidate languages ​​based on the language association weight coefficient, the first score and the second score.

[0194] As an optional implementation of the embodiment of the present disclosure, the device further includes: a time determination module, configured to:

[0195] Before determining the language association weight coefficients of the first audio data and the second audio data according to the end time corresponding to the first audio data, the start time corresponding to the second audio data and the language association weight function, the end time corresponding to the first audio data and the start time corresponding to the second audio data are determined based on an endpoint detection method.

[0196] As an optional implementation of the embodiment of the present disclosure, the score determination unit 8022 is specifically configured to:

[0197] multiplying the first score by the language association weight coefficient to obtain a corresponding product;

[0198] The product is added to the second score to obtain target scores for the second audio data belonging to different candidate languages.

[0199] As an optional implementation of the embodiment of the present disclosure, the prediction score determination module 801 includes:

[0200] A first score determination unit is configured to:

[0201] Preprocessing the first audio data to obtain a processed audio signal;

[0202] Extracting features of the audio signal to obtain corresponding first Mel-frequency cepstral coefficient features;

[0203] Inputting the first Mel-frequency cepstral coefficient feature into a mixed language recognition model to obtain first scores indicating that the first audio data belongs to different candidate languages;

[0204] A second score determination unit is configured to:

[0205] Second scores of the second audio data belonging to different candidate languages ​​are obtained.

[0206] As an optional implementation of the embodiment of the present disclosure, the mixed language recognition model is trained in the following way:

[0207] Acquire a training set sample, wherein the training set sample includes multiple audio samples in different languages;

[0208] Extracting features from the audio sample to obtain corresponding second Mel-frequency cepstral coefficient features;

[0209] The second Mel-frequency cepstral coefficient feature is input into a mixed-language language recognition model for training until the mixed-language language recognition model converges.

[0210] As an optional implementation of the embodiment of the present disclosure, the device further includes: a function determination module, configured to:

[0211] Before determining target scores for the second audio data to belong to different candidate languages ​​respectively based on the language association weight function of the first audio data and the second audio data, the first score, and the second score, determining a weight prediction order and a weight prediction coefficient;

[0212] Based on the weight prediction order and the weight prediction coefficient, a language association weight function of the first audio data and the second audio data is determined.

[0213] The language recognition device provided in the embodiments of the present disclosure can execute the language recognition method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method. To avoid repetition, they will not be described here.

[0214] An embodiment of the present disclosure provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the language recognition methods described in the embodiments of the present disclosure.

[0215] Fig. 9 Schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Fig. 9 As shown, the electronic device includes a processor 910 and a storage device 920; the number of processors 910 in the electronic device can be one or more. Fig. 9 A processor 910 is taken as an example; the processor 910 and the storage device 920 in the electronic device can be connected via a bus or other means. Fig. 9 The example of connecting through bus is taken in the following.

[0216] The storage device 920 is a computer-readable storage medium that can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the language recognition method in the embodiment of the present disclosure. The processor 910 executes various functional applications and data processing of the electronic device by running the software programs, instructions and modules stored in the storage device 920, that is, implements the language recognition method provided in the embodiment of the present disclosure.

[0217] The storage device 920 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the storage device 920 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the storage device 920 may further include a memory remotely arranged relative to the processor 910, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0218] An electronic device provided in this embodiment can be used to execute the language recognition method provided in any of the above embodiments, and has corresponding functions and beneficial effects.

[0219] The embodiments of the present disclosure provide a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned language recognition method are implemented and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0220] The computer readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0221] For the convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above discussion in some embodiments is not intended to be exhaustive or limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.

Claims

1. A language recognition method, characterized in that: The method comprises: Obtaining first scores of first audio data belonging to different candidate languages, and second scores of second audio data belonging to different candidate languages, wherein the first audio data is audio data corresponding to a wake-up word of a target control device, and the second audio data is audio data corresponding to a control command of the target control device; Determining target scores for the second audio data to belong to different candidate languages, respectively, based on the language association weight function of the first audio data and the second audio data, the first score, and the second score; Determining a target language corresponding to the second audio data based on the target score; The determining, based on the language association weight function of the first audio data and the second audio data, the first score, and the second score, target scores for the second audio data belonging to different candidate languages ​​respectively includes: Determining language association weight coefficients of the first audio data and the second audio data according to an end time corresponding to the first audio data, a start time corresponding to the second audio data, and the language association weight function; Based on the language association weight coefficient, the first score, and the second score, target scores for the second audio data belonging to different candidate languages ​​are determined.

2. The method according to claim 1, characterized in that Before determining the language association weight coefficients of the first audio data and the second audio data according to the end time corresponding to the first audio data, the start time corresponding to the second audio data, and the language association weight function, the method further includes: Based on an endpoint detection method, an end time corresponding to the first audio data and a start time corresponding to the second audio data are determined.

3. The method according to claim 1, characterized in that The determining, based on the language association weight coefficient, the first score, and the second score, target scores for the second audio data belonging to different candidate languages ​​respectively includes: multiplying the first score by the language association weight coefficient to obtain a corresponding product; The product is added to the second score to obtain target scores for the second audio data belonging to different candidate languages.

4. The method according to claim 1, characterized in that: The obtaining of first scores of the first audio data belonging to different candidate languages ​​includes: Preprocessing the first audio data to obtain a processed audio signal; Extracting features of the audio signal to obtain corresponding first Mel-frequency cepstral coefficient features; The first Mel-frequency cepstral coefficient feature is input into a mixed language recognition model to obtain first scores of the first audio data belonging to different candidate languages.

5. The method according to claim 4, characterized in that The mixed language recognition model is trained in the following way: Acquire a training set sample, wherein the training set sample includes multiple audio samples in different languages; Extracting features from the audio sample to obtain corresponding second Mel-frequency cepstral coefficient features; The second Mel-frequency cepstral coefficient feature is input into a mixed-language language recognition model for training until the mixed-language language recognition model converges.

6. The method according to any one of claims 1 to 5, characterized in that: Before determining the target scores of the second audio data belonging to different candidate languages ​​respectively according to the language association weight function of the first audio data and the second audio data, the first score and the second score, the method further includes: Determine the weight prediction order and weight prediction coefficient; Based on the weight prediction order and the weight prediction coefficient, a language association weight function of the first audio data and the second audio data is determined.

7. A language recognition device, characterized in that: The device comprises: a prediction score determination module, used to obtain first scores of first audio data belonging to different candidate languages, and second scores of second audio data belonging to different candidate languages, wherein the first audio data is audio data corresponding to a wake-up word of a target control device, and the second audio data is audio data corresponding to a control command of the target control device; a target score determination module, configured to determine target scores for the second audio data belonging to different candidate languages ​​according to a language association weight function of the first audio data and the second audio data, the first score, and the second score; a target language determination module, configured to determine a target language corresponding to the second audio data based on the target score; The target score determination module includes: a coefficient determination unit, configured to determine a language association weight coefficient of the first audio data and the second audio data according to an end time corresponding to the first audio data, a start time corresponding to the second audio data, and the language association weight function; A score determination unit is used to determine target scores for the second audio data belonging to different candidate languages ​​based on the language association weight coefficient, the first score and the second score.

8. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Language recognition mode selection method, device and household appliance

    CN109360564A

  • Speech recognition method, device and system

    CN109817220A