Multi-dialect voice wake-up method, device and equipment and readable storage medium

By preprocessing and decomposing the original audio data into pronunciation sequences, combining the dialect language recognition model, the language results of the multi-dial voice wake-up system are determined and fill-in words are filtered, and the accuracy and false wake-up problems of the multi-dial wake-up system are solved, achieving higher wake-up accuracy and lower false wake-up rate.

CN120260566APending Publication Date: 2025-07-04AISPEECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510401641.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing multi-dial voice wake-up system does not have the wake-up effect when dealing with dialects with unique pronunciation rules, vocabulary and grammatical structures, and is prone to failure or incorrect wake-up.

Method used

By obtaining the original audio data, preprocessing and decomposing it into pronunciation sequences, the results of the first and second languages are determined in combination with the dialect language recognition model, and filler words are filtered to make wake-up decisions.

Benefits of technology

It improves the accuracy of multi-dialect awakening, reduces the false awakening rate, and enhances the recognition ability of different dialects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260566A_ABST
    Figure CN120260566A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-dialect voice wake-up method, device and equipment and a readable storage medium, and relates to the technical field of big data. Comprising the following steps: acquiring original audio data, and preprocessing the original audio data to obtain initial audio data; decomposing the initial audio data into a pronunciation sequence, and determining a first language result of the initial audio data based on the pronunciation sequence; inputting the initial audio data into a dialect language recognition model, and determining a second language result of the initial audio data; both the first language result and the second language result comprise a dialect type and a wake-up possibility of the initial audio data; based on the first language result and the second language result, filtering filling words in the initial audio data to obtain a target wake-up audio; and executing a wake-up decision based on the target wake-up audio. The multi-dialect voice wake-up accuracy is improved, and the false wake-up rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and more particularly to a multi-dialect voice wake-up method, device, equipment and readable storage medium. Background Art

[0002] With the wide application of intelligent voice interaction technology globally, accurate recognition and wake-up of different languages, especially dialects, have become key factors in enhancing the user experience. Current multi-dialect voice recognition systems mainly rely on large-scale datasets and algorithms such as deep neural networks (DNN), convolutional neural networks (CNN), and long short-term memory networks (LSTM) to process the speech features of different dialects.

[0003] Although certain progress has been made in current multi-dialect voice recognition technology, most current voice wake-up systems are mainly built based on Mandarin or a few common languages. For many dialects with unique pronunciation rules, vocabulary, and grammatical structures, their wake-up effects are often unsatisfactory. The differences between dialects are huge, including aspects such as tones, initial consonants, final vowels, and tone sandhi, which make it difficult for traditional single-model voice wake-up technology to adapt to multi-dialect environments and prone to wake-up failures or false wake-ups.

[0004] Therefore, there is an urgent need for a multi-dialect voice wake-up method that can improve the accuracy of dialect wake-up and reduce the false wake-up rate. Summary of the Invention

[0005] The purpose of the present invention is to provide a multi-dialect voice wake-up method, device, equipment and readable storage medium. By determining the first language result according to the pronunciation sequence and simultaneously determining the second language result according to the dialect language recognition model, and then making a wake-up judgment based on the two language results. Compared with a single judgment method, this method can increase the accuracy of dialect wake-up and reduce misjudgments caused by a single judgment method. In addition, filler words are filtered later to further avoid false wake-ups caused by similar sounds, that is, it increases the accuracy of multi-dialect wake-up and reduces the false wake-up rate.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a multi-dialect voice wake-up method, which includes:

[0008] Obtain the original audio data, and preprocess the original audio data to obtain the initial audio data;

[0009] Decompose the initial audio data into a pronunciation sequence, and determine the first language result of the initial audio data based on the pronunciation sequence;

[0010] Input the initial audio data into the dialect language recognition model to determine the second language result of the initial audio data; both the first language result and the second language result include the dialect type and wake-up possibility of the initial audio data;

[0011] Based on the first language result and the second language result, filter out the filler words in the initial audio data to obtain the target wake-up audio;

[0012] Based on the target wake-up audio, perform a wake-up decision.

[0013] In some embodiments, based on the first language result and the second language result, filtering out the filler words in the initial audio data to obtain the target wake-up audio includes:

[0014] Determine the wake-up probability of the initial audio data according to the first language result and the second language result;

[0015] If the wake-up probability is greater than a preset probability threshold, filter out the filler words in the initial audio data to obtain the target wake-up audio.

[0016] In some embodiments, filtering out the filler words in the initial audio data to obtain the target wake-up audio includes:

[0017] Establish a filler word audio library;

[0018] Based on the filler word audio library, filter out the filler words in the initial audio data to obtain the target wake-up audio.

[0019] In some embodiments, determining the wake-up probability of the initial audio data according to the first language result and the second language result includes:

[0020] Determine the type similarity of the dialect types in the first language result and the second language result;

[0021] Determine the wake-up possibilities of the first language result and the second language result;

[0022] Based on the type similarity and the wake-up possibilities, determine the wake-up probability of the initial audio data.

[0023] In some embodiments, decomposing the initial audio data into a pronunciation sequence and determining the first language result of the initial audio data based on the pronunciation sequence includes:

[0024] According to the pronunciation characteristics of different dialects, establish a pronunciation unit library;

[0025] Decompose the initial audio data into pronunciation sequences, and compare the pronunciation sequences with the pronunciation unit library to determine the first language result of the initial audio data.

[0026] In some embodiments, comparing the pronunciation sequences with the pronunciation unit library to determine the first language result of the initial audio data includes:

[0027] Calculate the matching degree between the pronunciation sequences and the standard pronunciation units in the pronunciation unit library;

[0028] Determine the first language result of the initial audio data according to the matching degree.

[0029] In a second aspect, the present invention also provides a multi-dialect voice wake-up device, which includes:

[0030] An audio acquisition module, configured to acquire original audio data and preprocess the original audio data to obtain initial audio data;

[0031] A first result module, configured to decompose the initial audio data into pronunciation sequences and determine the first language result of the initial audio data based on the pronunciation sequences;

[0032] A second result module, configured to input the initial audio data into a dialect language recognition model to determine the second language result of the initial audio data; both the first language result and the second language result include the dialect type and wake-up possibility of the initial audio data;

[0033] An audio filtering module, configured to filter out filler words in the initial audio data based on the first language result and the second language result to obtain a target wake-up audio;

[0034] A decision execution module, configured to execute a wake-up decision based on the target wake-up audio.

[0035] In a third aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the multi-dialect voice wake-up method provided in the first aspect is implemented.

[0036] In a fourth aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the multi-dialect voice wake-up method provided in the first aspect is implemented.

[0037] In a fifth aspect, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the multi-dialect voice wake-up method provided in the first aspect is implemented.

[0038] The beneficial effects of the present invention are as follows:

[0039] In the multi-dialect voice wake-up method in this application, first, the original audio data is obtained, and the original audio data is preprocessed to obtain the initial audio data; then the initial audio data is decomposed into pronunciation sequences, and based on the pronunciation sequences, the first language result of the initial audio data is determined; then the initial audio data is input into the dialect language recognition model to determine the second language result of the initial audio data; both the first language result and the second language result include the dialect type and wake-up possibility of the initial audio data; also based on the first language result and the second language result, the filler words in the initial audio data are filtered to obtain the target wake-up audio; finally, based on the target wake-up audio, a wake-up decision is made. By determining the first language result according to the pronunciation sequence and simultaneously determining the second language result according to the dialect language recognition model, and then making a wake-up judgment according to the two language results, compared with a single judgment method, this method can increase the accuracy of dialect wake-up and reduce the misjudgment caused by a single judgment method. In addition, the filler words are filtered later, further avoiding the false wake-up caused by similar sounds, that is, increasing the accuracy of multi-dialect wake-up and reducing the false wake-up rate.

[0040] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly and implement it according to the content of the description, the following takes the preferred embodiments of the present invention and combines the drawings to describe in detail as follows. Brief Description of the Drawings

[0041] Figure 1 It is a schematic flowchart of a multi-dialect voice wake-up method shown in an embodiment of the present invention;

[0042] Figure 2 It is a schematic flowchart of a method for determining the wake-up probability of initial audio data shown in an embodiment of the present invention;

[0043] Figure 3 It is a schematic flowchart of another multi-dialect voice wake-up method shown in an embodiment of the present invention;

[0044] Figure 4 It is a structural diagram of a multi-dialect voice wake-up device shown in an embodiment of the present invention;

[0045] Figure 5 It is a schematic structural diagram of an electronic device provided in an embodiment of this application. Detailed Description of the Embodiments

[0046] The technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work shall fall within the protection scope of the present invention.

[0047] It should be noted that the references to "one embodiment", "embodiment", "example embodiment", etc. in this specification mean that the described embodiment may include specific features, structures or characteristics. However, not every embodiment must include these specific features, structures or characteristics. In addition, such expressions do not refer to the same embodiment. Further, when combining specific features, structures or characteristics in conjunction with an embodiment, it has been shown that it is within the knowledge of those skilled in the art to combine such features, structures or characteristics into other embodiments whether or not explicitly described.

[0048] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0049] In some embodiments, as Figure 1 shown, a flow schematic diagram of a multi-dialect voice wake-up method is provided:

[0050] S101, obtain the original audio data, and preprocess the original audio data to obtain the initial audio data.

[0051] Specifically, the voice signal in the environment can be collected through a microphone, which is the original audio data, and then the original audio data is subjected to noise reduction, filtering and amplification processing, and the initial audio data is obtained.

[0052] S102, decompose the initial audio data into pronunciation sequences, and determine the first language result of the initial audio data based on the pronunciation sequences.

[0053] Among them, the first language result includes the first dialect type and the first wake-up possibility of the initial audio data.

[0054] Specifically, an audio decomposition model can be used to decompose the initial audio data into pronunciation sequences, and then compare the pronunciation sequences with the pronunciation sequences of various dialects, and the first dialect type of the initial audio data can be determined according to the comparison result. The semantic information of the initial audio can also be determined according to the pronunciation sequences, and then the first wake-up possibility of the initial audio data can be determined according to the semantic information.

[0055] Optionally, the method for determining the first language result may also be: establishing a pronunciation unit library according to the pronunciation characteristics of different dialects; decomposing the initial audio data into a pronunciation sequence, and comparing the pronunciation sequence with the pronunciation unit library to determine the first language result of the initial audio data.

[0056] Specifically, a large amount of dialect information needs to be collected, such as the initials, finals, tone combinations of various different dialects, etc., and according to the pronunciation characteristics of these dialects, a pronunciation unit library is pre-created. After decomposing the initial audio data into a pronunciation sequence, since the pronunciation unit library contains the pronunciation sequences of various dialects, the pronunciation sequence of the initial audio data can be compared with the pronunciation unit library, and then the first language result of the initial audio data can be determined.

[0057] Optionally, it may also be to first calculate the matching degree between the pronunciation sequence and the standard pronunciation unit in the pronunciation unit library; according to the matching degree, determine the first language result of the initial audio data.

[0058] Exemplarily, when the matching degree between the pronunciation sequence of the initial audio information and the standard pronunciation unit in the pronunciation unit library is greater than the preset matching degree threshold, it is determined that the initial audio information is the dialect corresponding to the standard pronunciation unit, that is, the first language result of the initial audio data is determined.

[0059] S103, input the initial audio data into the dialect language recognition model to determine the second language result of the initial audio data.

[0060] Among them, the second language result includes the second dialect type and the second wake-up possibility of the initial audio data; the dialect language recognition model can comprehensively judge the dialect language to which the speech belongs from aspects such as the overall characteristics of the speech, such as intonation contour, vocabulary usage frequency, grammatical structure characteristics, etc. By performing deep feature extraction and classification on the speech signal, it can accurately identify which dialect category it belongs to, such as Cantonese, Minnan dialect, Sichuan dialect, etc.

[0061] Specifically, the initial audio data can be directly input into the dialect language recognition model, and the output result of the dialect language recognition model is the second language result of the initial audio data.

[0062] S104, based on the first language result and the second language result, filter out the filler words in the initial audio data to obtain the target wake-up audio.

[0063] Among them, filler words are usually some words or sound segments with approximate sounds that appear but do not have actual wake-up significance.

[0064] Specifically, if the first dialect type and the second dialect type are the same, and both the first wake-up possibility and the second wake-up possibility are greater than the preset possibility threshold, directly filter out the filler words from the initial audio data to obtain the target wake-up audio.

[0065] S105, Perform a wake-up decision based on the target wake-up audio.

[0066] Exemplarily, when the wake-up object is a car head unit, the target wake-up audio can be directly input into the car head unit, and the car head unit can be woken up by the target wake-up audio.

[0067] In the multi-dialect voice wake-up method in the above embodiments, first obtain the original audio data, and preprocess the original audio data to obtain the initial audio data; then decompose the initial audio data into pronunciation sequences, and determine the first language result of the initial audio data based on the pronunciation sequences; then input the initial audio data into the dialect language recognition model to determine the second language result of the initial audio data; both the first language result and the second language result include the dialect type and wake-up possibility of the initial audio data; and then filter out the filler words in the initial audio data based on the first language result and the second language result to obtain the target wake-up audio; finally, perform a wake-up decision based on the target wake-up audio. By determining the first language result according to the pronunciation sequence, and at the same time determining the second language result according to the dialect language recognition model, and then making a wake-up judgment according to the two language results, compared with a single judgment method, this method can increase the accuracy of dialect wake-up and reduce the misjudgment caused by a single judgment method. In addition, filler words are filtered later, further avoiding mis-wake-up caused by approximate sounds, that is, increasing the multi-dialect wake-up accuracy and reducing the mis-wake-up rate.

[0068] In another embodiment, as Figure 2 shown, the specific process of determining the target wake-up audio is elaborated in detail. The specific method includes:

[0069] S201, Determine the wake-up probability of the initial audio data according to the first language result and the second language result.

[0070] Specifically, the first language result and the second language result can be input into the wake-up probability calculation model, and the wake-up probability calculation model can output the wake-up probability of the initial audio data.

[0071] Optionally, the method for determining the wake-up probability of the initial audio data can also be: determine the type similarity of the dialect types in the first language result and the second language result; determine the wake-up possibilities of the first language result and the second language result; and determine the wake-up probability of the initial audio data based on the type similarity and the wake-up possibilities.

[0072] Exemplarily, if the first language result is Cantonese and the second language result is also Cantonese, the type similarity of the dialect types in the first language result and the second language result is 100%. If the wake-up probability of the first language result is 50% and the wake-up probability of the first language result is 60%, then the average wake-up probability of the first language result and the second language result is 55%. At this time, the wake-up probability of the initial audio data can be the product of the type similarity and the wake-up probability, that is, 55%.

[0073] S202. If the wake-up probability is greater than the preset probability threshold, filter the filler words in the initial audio data to obtain the target wake-up audio.

[0074] Specifically, a probability threshold can be set in advance, and then the wake-up probability of the initial audio data is compared with the preset probability threshold. When the wake-up probability is greater than the preset probability threshold, it indicates that the wake-up probability is relatively high. At this time, the initial audio data can be input into the filler word filtering model to filter out the useless filler words in the initial audio model and obtain the target wake-up audio.

[0075] Optionally, another method to obtain the target wake-up audio can be: establish a filler word audio library; based on the filler word audio library, filter the filler words in the initial audio data to obtain the target wake-up audio.

[0076] Specifically, based on different dialect types, establish a filler word audio library, establish the maximum filler word segment of non-wake-up words within the wake-up word area, and filter all words in the initial audio data that exceed the set threshold, that is, complete the filtering of the filler words and obtain the target wake-up audio.

[0077] Optionally, if the wake-up probability is less than or equal to the preset probability threshold, it indicates that the initial audio data cannot be woken up, and at this time, the process does not continue.

[0078] In the method of the above embodiments, first determine the wake-up probability of the initial audio data according to the first language result and the second language result; if the wake-up probability is greater than the preset probability threshold, filter the filler words in the initial audio data to obtain the target wake-up audio. Judging the wake-up through the results of two languages greatly increases the accuracy of wake-up compared with judging through a single result. Since the filler words in the initial audio data are filtered, the situation of false wake-up caused by some detailed pronunciations is avoided, and the false wake-up rate is reduced.

[0079] To more comprehensively demonstrate the present solution, an optional way of the multi-dialect voice wake-up method is given in this embodiment, as Figure 3 shown:[[]]END]]

[0080] S301. Obtain the original audio data and preprocess the original audio data to obtain the initial audio data.

[0081] S302. Establish a pronunciation unit library according to the pronunciation characteristics of different dialects.

[0082] S303. Decompose the initial audio data into a pronunciation sequence, and calculate the matching degree between the pronunciation sequence and the standard pronunciation units in the pronunciation unit library.

[0083] S304. Determine the first language result of the initial audio data according to the matching degree.

[0084] S305. Input the initial audio data into the dialect language recognition model to determine the second language result of the initial audio data.

[0085] Wherein, both the first language result and the second language result include the dialect type and the wake-up possibility of the initial audio data.

[0086] S306. Determine the type similarity of the dialect types in the first language result and the second language result.

[0087] S307. Determine the wake-up possibility of the first language result and the second language result.

[0088] S308. Based on the type similarity and the wake-up possibility, determine the wake-up probability of the initial audio data.

[0089] S309. If the wake-up probability is greater than the preset probability threshold, establish a filler word audio library.

[0090] S310. Based on the filler word audio library, filter the filler words in the initial audio data to obtain the target wake-up audio.

[0091] S311. Based on the target wake-up audio, perform a wake-up decision.

[0092] For the specific processes of the above S301 - S311, reference can be made to the description of the method embodiments above. Their implementation principles and technical effects are similar, and will not be elaborated here.

[0093] Based on the same inventive concept, an embodiment of the present application also provides a multi-dialect voice wake-up device for implementing the above-mentioned multi-dialect voice wake-up method. The implementation solution provided by this device to solve the problem is similar to the implementation solution recorded in the above method. Therefore, the specific limitations in one or more embodiments of the multi-dialect voice wake-up device provided below can refer to the limitations on the multi-dialect voice wake-up method in the above text, and will not be elaborated here.

[0094] In one embodiment, as Figure 4 shown, a multi-dialect voice wake-up device is provided, and this device includes:

[0095] An audio acquisition module 40, configured to acquire original audio data and preprocess the original audio data to obtain initial audio data;

[0096] A first result module 41, configured to decompose the initial audio data into pronunciation sequences and determine a first language result of the initial audio data based on the pronunciation sequences;

[0097] A second result module 42, configured to input the initial audio data into a dialect language recognition model to determine a second language result of the initial audio data; both the first language result and the second language result include the dialect type and wake-up possibility of the initial audio data;

[0098] An audio filtering module 43, configured to filter filler words in the initial audio data based on the first language result and the second language result to obtain a target wake-up audio;

[0099] A decision execution module 44, configured to execute a wake-up decision based on the target wake-up audio.

[0100] In another embodiment, the above-mentioned Figure 4 audio filtering module 43 is specifically configured to: determine the wake-up probability of the initial audio data according to the first language result and the second language result; if the wake-up probability is greater than a preset probability threshold, filter filler words in the initial audio data to obtain a target wake-up audio.

[0101] Specifically, determine the type similarity of the dialect types in the first language result and the second language result; determine the wake-up possibility of the first language result and the second language result; based on the type similarity and the wake-up possibility, determine the wake-up probability of the initial audio data. Establish a filler word audio library; based on the filler word audio library, filter filler words in the initial audio data to obtain a target wake-up audio.

[0102] In another embodiment, the above-mentioned Figure 4 first result module 41 is specifically configured to: establish a pronunciation unit library according to the pronunciation characteristics of different dialects; decompose the initial audio data into pronunciation sequences and compare the pronunciation sequences with the pronunciation unit library to determine the first language result of the initial audio data.

[0103] Specifically, calculate the matching degree between the pronunciation sequence and the standard pronunciation unit in the pronunciation unit library; according to the matching degree, determine the first language result of the initial audio data

[0104] The embodiments of the present application further provide an electronic device. In some embodiments, refer to Figure 5As shown in the figure, the electronic device 700 includes an input unit 710, a memory 720, a processor 730, and an output unit 740. The memory 720 stores program instructions that can run on the processor 730. The processor 730 can execute the multi-dialect voice wake-up method and / or technical solution based on the foregoing embodiments by invoking the program instructions. The electronic device 700 can be a mobile terminal device such as a mobile phone or a computer.

[0105] In addition, an embodiment of the present application further provides a computer-readable storage medium for storing a computer program for executing the multi-dialect voice wake-up method. For example, computer program instructions, when executed by a computer, can invoke or provide the method and / or technical solution according to the present application through the operation of the computer. The program instructions for invoking the method of the present application may be stored in a fixed or removable storage medium, and / or transmitted and / or stored in a storage medium running according to the program instructions through a data stream in a broadcast or other signal-bearing medium.

[0106] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program code executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps of them can be fabricated into a single integrated circuit module to be implemented. In this way, the present application is not limited to any specific combination of hardware and software.

[0107] The technical features of the above embodiments can be arbitrarily integrated. For the sake of brevity of description, not all possible integrations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the integration of these technical features, it should be considered as the scope described in this specification.

[0108] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.

Claims

1. A multi-dialect voice wake-up method, characterized in that, The method includes: Obtaining original audio data, and preprocessing the original audio data to obtain initial audio data; Decomposing the initial audio data into pronunciation sequences, and determining a first language result of the initial audio data based on the pronunciation sequences; Inputting the initial audio data into a dialect language recognition model to determine a second language result of the initial audio data; both the first language result and the second language result include the dialect type and wake-up possibility of the initial audio data; Filtering filler words in the initial audio data based on the first language result and the second language result to obtain a target wake-up audio; Performing a wake-up decision based on the target wake-up audio.

2. The multi-dialect voice wake-up method according to claim 1, wherein Filtering filler words in the initial audio data based on the first language result and the second language result to obtain a target wake-up audio, including: Determining a wake-up probability of the initial audio data according to the first language result and the second language result; If the wake-up probability is greater than a preset probability threshold, filtering filler words in the initial audio data to obtain a target wake-up audio.

3. The multi-dialect voice wake-up method according to claim 2, wherein Filtering filler words in the initial audio data to obtain a target wake-up audio, including: Establishing a filler word audio library; Filtering filler words in the initial audio data based on the filler word audio library to obtain a target wake-up audio.

4. The multi-dialect voice wake-up method according to claim 2, wherein, Determining a wake-up probability of the initial audio data according to the first language result and the second language result, including: Determining a type similarity of the dialect types in the first language result and the second language result; Determining the wake-up possibilities of the first language result and the second language result; Determining the wake-up probability of the initial audio data based on the type similarity and the wake-up possibilities.

5. The multi-dialect voice wake-up method according to claim 1, characterized in that Decomposing the initial audio data into pronunciation sequences, and determining a first language result of the initial audio data based on the pronunciation sequences, including: Establishing a pronunciation unit library according to the pronunciation characteristics of different dialects; Decomposing the initial audio data into pronunciation sequences, and comparing the pronunciation sequences with the pronunciation unit library to determine a first language result of the initial audio data.

6. The multi-dialect voice wake-up method according to claim 5, wherein Comparing the pronunciation sequences with the pronunciation unit library to determine a first language result of the initial audio data, including: Calculating a matching degree between the pronunciation sequences and standard pronunciation units in the pronunciation unit library; Determining a first language result of the initial audio data according to the matching degree.

7. A multi-dialect voice wake-up device, characterized in that, The device includes: An audio acquisition module, configured to obtain original audio data, and preprocess the original audio data to obtain initial audio data; A first result module, configured to decompose the initial audio data into pronunciation sequences, and determine a first language result of the initial audio data based on the pronunciation sequences; A second result module, configured to input the initial audio data into a dialect language recognition model to determine a second language result of the initial audio data; both the first language result and the second language result include the dialect type and wake-up possibility of the initial audio data; An audio filtering module, configured to filter filler words in the initial audio data based on the first language result and the second language result, to obtain a target wake-up audio; A decision execution module, configured to execute a wake-up decision based on the target wake-up audio.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the multi-dialect voice wake-up method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the multi-dialect voice wake-up method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the multi-dialect voice wake-up method according to any one of claims 1 to 6 is implemented.