Speech recognition method, device, mobile terminal and computer-readable storage medium

By using sensor data to update historical location information in the positioning blind spot and obtaining the target speech recognition model, the problem of low speech recognition accuracy in the positioning blind spot of the mobile terminal is solved, and accurate positioning and efficient speech recognition without GPS or base station signals are achieved.

CN111798839BActive Publication Date: 2025-07-18CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202010734647.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-27
Publication Date
2025-07-18
Estimated Expiration
2040-07-27

AI Technical Summary

Technical Problem

Mobile terminals cannot locate location information when positioning blind spots, resulting in low accuracy of speech recognition and poor user experience.

Method used

It is determined whether the mobile terminal is in a positioning blind spot by intervaling the first preset time, and updates the historical position information based on the sensor data at the interval of the second preset time when positioning the blind spot, acquires the target speech recognition model to recognize the speech data, and converts it into standard Mandarin text.

Benefits of technology

When GPS, base station or mobile network cannot be located, the location of the mobile terminal can still be accurately positioned, improving the accuracy of voice recognition and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111798839B_ABST
    Figure CN111798839B_ABST
Patent Text Reader

Abstract

This application relates to artificial intelligence and speech processing, and provides a speech recognition method, device, mobile terminal, and computer-readable storage medium. The method includes: determining whether the mobile terminal is in a positioning blind area at intervals of a first preset time; when it is determined that the mobile terminal is in the positioning blind area, updating the historical location information of the mobile terminal based on the sensor data of the mobile terminal at intervals of a second preset time; when obtaining the speech data of the user, if the mobile terminal is still in the positioning blind area, obtaining a target speech recognition model according to the updated historical location information; and recognizing the speech data according to the target speech recognition model to obtain the standard Mandarin text corresponding to the speech data. This application can solve the problem that when the mobile terminal is located in the positioning blind area, the location information of the mobile terminal cannot be located, and thus the accuracy of speech recognition cannot be guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology in artificial intelligence, and particularly to a speech recognition method, device, mobile terminal, and computer-readable storage medium. Background Art

[0002] With the rapid development of speech recognition technology, speech recognition technology has gradually been applied to mobile terminals, enabling mobile terminals to recognize users' speech data and obtain text data. However, when recognizing users' speech data, it is affected and interfered by users' accents and dialects, and the accuracy of speech recognition is relatively low. Currently, the location information of the mobile terminal can be located through a GPS positioning device, a base station, or a mobile network, and a dialect recognition model is matched according to the location information. Then, through the dialect recognition model and the mapping relationship between the dialect and standard Mandarin, the speech data is processed to obtain a standard Mandarin text to improve the accuracy of speech recognition. However, in some cases, the location information of the mobile terminal cannot be located through a GPS positioning device, a base station, or a mobile network, making it impossible to match a dialect recognition model based on the user's location information, resulting in a relatively low accuracy of speech recognition and a poor user experience. Summary of the Invention

[0003] The main purpose of this application is to provide a speech recognition method, device, mobile terminal, and computer-readable storage medium, aiming to solve the problem that when the mobile terminal is in a positioning blind area, the location information of the mobile terminal cannot be located, and thus the accuracy of speech recognition cannot be guaranteed.

[0004] In a first aspect, this application provides a speech recognition method, including:

[0005] Determining whether the mobile terminal is in a positioning blind area at intervals of a first preset time;

[0006] When it is determined that the mobile terminal is in a positioning blind area, updating the historical location information of the mobile terminal based on the sensor data of the mobile terminal at intervals of a second preset time;

[0007] When the speech data of the user is obtained, if the mobile terminal is still in a positioning blind area, obtaining a target speech recognition model according to the updated historical location information;

[0008] Recognizing the speech data according to the target speech recognition model to obtain a standard Mandarin text corresponding to the speech data.

[0009] In a second aspect, this application further provides a speech recognition device, and the speech recognition device includes:

[0010] A determination module, configured to determine whether the mobile terminal is in a positioning blind area at intervals of a first preset time;

[0011] A location update module, configured to update the historical location information of the mobile terminal based on the sensor data of the mobile terminal at intervals of a second preset time when it is determined that the mobile terminal is in a positioning blind area;

[0012] An acquisition module, configured to, when acquiring voice data of a user, if the mobile terminal is still in a positioning blind area, acquire a target voice recognition model according to the updated historical location information;

[0013] A voice recognition module, configured to recognize the voice data according to the target voice recognition model to obtain a standard Mandarin text corresponding to the voice data.

[0014] In a third aspect, the present application further provides a mobile terminal, which includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the voice recognition method described above are implemented.

[0015] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the voice recognition method described above are implemented.

[0016] The present application provides a voice recognition method, apparatus, mobile terminal, and computer-readable storage medium. The present application determines whether a mobile terminal is in a positioning blind area at intervals of a first preset time, and when it is determined that the mobile terminal is in a positioning blind area, updates the historical location information of the mobile terminal based on the sensor data of the mobile terminal at intervals of a second preset time. Then, when acquiring voice data of a user, if the mobile terminal is still in a positioning blind area, a target voice recognition model is acquired according to the updated historical location information, and the voice data is recognized according to the target voice recognition model to obtain a standard Mandarin text corresponding to the voice data. The above technical solution can still locate the location information of the mobile terminal when the location information of the mobile terminal cannot be located according to a GPS positioning device, a base station, or a mobile network, so as to match an accurate voice recognition model, improve the accuracy of voice recognition, and greatly improve the user experience. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1Flow chart of a speech recognition method provided by an embodiment of the present application;

[0019] Figure 2 is Figure 1 Sub-step flow chart of the speech recognition method in;

[0020] Figure 3 Schematic block diagram of a speech recognition device provided by an embodiment of the present application;

[0021] Figure 4 is Figure 3 Schematic block diagram of a sub-module of the speech recognition device in;

[0022] Figure 5 Schematic block diagram of the structure of a mobile terminal provided by an embodiment of the present application.

[0023] The realization, functional features and advantages of the purpose of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0025] The flow charts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, combined or partially merged, so the actual execution order may be changed according to the actual situation.

[0026] The embodiments of the present application provide a speech recognition method, device, mobile terminal and computer-readable storage medium. Among them, the speech recognition method can be applied to terminal devices, and the terminal devices can be electronic devices such as mobile phones, tablet computers, notebook computers, desktop computers, personal digital assistants and wearable devices.

[0027] Next, some embodiments of the present application will be described in detail with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0028] Please refer to Figure 1 , Figure 1 Flow chart of a speech recognition method provided by an embodiment of the present application.

[0029] As Figure 1 shown, the speech recognition method includes steps S101 to S104.

[0030] Step S101: Determine whether the mobile terminal is in a positioning blind area at intervals of a first preset time.

[0031] Among them, the mobile terminal can locate the position information of the mobile terminal through technologies such as the Global Positioning System (GPS), base station positioning technology, and network positioning technology. The positioning blind area refers to an area where the mobile terminal cannot be located through GPS positioning technology, base station positioning technology, and / or network positioning technology, etc. The first preset time can be set based on actual situations, and this application does not make specific limitations. For example, the first preset time is 10 seconds or 30 seconds.

[0032] In an embodiment, the mobile terminal attempts to locate the mobile terminal at intervals of the first preset time through a GPS positioning device, a base station positioning program, and / or the network signal strength of the mobile terminal. When the mobile terminal cannot be located through the GPS positioning device, the base station positioning program, and / or the network signal strength of the mobile terminal, it is determined that the mobile terminal is in a positioning blind area. That is, the positioning blind area refers to a spatial area where the mobile terminal cannot be located through the GPS positioning device, the base station positioning program, and / or the network signal strength of the mobile terminal. If the mobile terminal can be located through the GPS positioning device, or can be located through the base station positioning program, or can be located through the network, it can be determined that the mobile terminal is not in a positioning blind area.

[0033] Exemplarily, the mobile terminal attempts to locate the mobile terminal at intervals of the first preset time through the GPS positioning device. If the positioning fails, it attempts to locate the mobile terminal through the base station positioning program. If it fails again, it obtains the network signal strength of the mobile terminal and determines whether the network signal strength of the mobile terminal is zero. If it is determined that the network signal strength of the mobile terminal is zero, it can be determined that the mobile terminal is in a positioning blind area.

[0034] Step S102: When it is determined that the mobile terminal is in a positioning blind area, update the historical position information of the mobile terminal based on the sensor data of the mobile terminal at intervals of a second preset time.

[0035] Among them, the mobile terminal includes an acceleration sensor, a direction sensor, a barometric pressure sensor, etc. The historical location information includes the location information of the mobile terminal determined by a GPS positioning device, a base station positioning program, and / or the network signal strength of the mobile terminal when the mobile terminal is not in a positioning blind area. For example, when the mobile terminal locates itself at intervals of a first preset time through a GPS positioning device, a base station positioning program, and / or the network signal strength of the mobile terminal, when the location information of the mobile terminal is located, the located location information is stored, so as to facilitate obtaining the stored location information and obtaining the historical location information when it is determined that the mobile terminal is in a positioning blind area later. The second preset time can be set based on actual situations, and the present application does not make specific limitations thereon. For example, the second preset time is 15 seconds or 20 seconds.

[0036] In one embodiment, the sensor data includes the acceleration of the mobile terminal output by the acceleration sensor of the mobile terminal and the moving direction of the mobile terminal output by the direction sensor. The manner of updating the historical location information of the mobile terminal based on the sensor data of the mobile terminal can be: determining the moving distance of the mobile terminal according to the acceleration and the second preset time; and updating the historical location information of the mobile terminal according to the moving distance and the moving direction. Among them, the moving distance of the mobile terminal can be obtained by performing a second integral calculation on the acceleration at the second preset time of the mobile terminal, that is, first integrating the acceleration within the second preset time of movement to obtain the moving speed, that is where v is the moving speed, a is the acceleration, and t is the second preset time. Then, integrate the moving speed within the second preset time of movement to obtain the moving distance, that is where d is the moving distance, v is the moving speed, and t is the second preset time.

[0037] In one embodiment, the manner of updating the historical location information of the mobile terminal according to the moving distance and the moving direction can be: marking the historical location information of the mobile terminal in a preset offline map to obtain the historical location point of the mobile terminal; determining the current location point of the mobile terminal on the preset offline map according to the moving distance, the moving direction, and the historical location point; obtaining the location information of the current location point on the preset offline map, and using the location information as the updated historical location information to update the historical location information of the mobile terminal.

[0038] Step S103, when the voice data of the user is obtained, if the mobile terminal is still in a positioning blind area, obtain a target voice recognition model according to the updated historical location information.

[0039] When the voice data of the user is obtained, the mobile terminal is attempted to be located through the GPS positioning device, the base station positioning program, and / or the network signal strength of the mobile terminal. When the mobile terminal cannot be located through the GPS positioning device, the base station positioning program, and / or the network signal strength of the mobile terminal, it is determined that the mobile terminal is still in the positioning blind area. Then, according to the updated historical location information, the target speech recognition model is obtained. When the mobile terminal can be located through the GPS positioning device, the base station positioning program, and / or the network signal strength of the mobile terminal, it is determined that the mobile terminal has moved from the positioning blind area to the locatable area. Therefore, the current positioning information obtained by positioning the mobile terminal through the GPS positioning device, the base station positioning program, and / or the network signal strength of the mobile terminal is obtained, and the target speech recognition model is obtained based on the current positioning information.

[0040] In one embodiment, the method for obtaining the target speech recognition model according to the updated historical location information may be: obtaining the regional code of the region where the updated historical location information is located; obtaining the speech recognition model bound to the regional code to obtain the target speech recognition model. Among them, different speech recognition models are established in advance based on different regions, and the binding relationship between the speech recognition model and the regional code is established, and the binding relationship between the speech recognition model and the regional code is stored in the memory of the mobile terminal. Thus, based on the regional code and the binding relationship between the speech recognition model and the regional code, the speech recognition model bound to the regional code can be obtained, and the regional code can be determined according to the divided cities.

[0041] In one embodiment, as Figure 2 shown, step S103 includes sub-steps S1031 to S1032.

[0042] Sub-step S1031: Determine the dialect type to which the voice data belongs.

[0043] Among them, the dialect types include Mandarin, Jin dialect, Wu dialect, Xiang dialect, Cantonese, Gan dialect, Hui dialect, Min dialect, Hakka, and Pinghua, etc. Mandarin can be further divided into Beijing Mandarin, Northeast Mandarin, Jilu Mandarin, Jianghuai Mandarin, Southwest Mandarin, JiaoLiao Mandarin, Central Plains Mandarin, and Lanyin Mandarin, etc.

[0044] In one embodiment, Mel cepstral features, fundamental frequency contour features, duration features, and energy features are extracted from the speech data; the Mel cepstral features are input into a preset pronunciation type recognition model to obtain the pronunciation type of each syllable segment; the fundamental frequency contour features, duration features, and energy features are input into a preset tone recognition model to obtain the tone of each syllable segment; and the dialect type to which the speech data belongs is determined according to the pronunciation type and the tone. Among them, the preset pronunciation type recognition model is obtained by training a neural network model using the Mel cepstral features and the pronunciation types corresponding to the Mel cepstral features as sample data, and the preset tone recognition model is obtained by training a neural network model using the fundamental frequency contour features, duration features, and energy features and the tones corresponding to the fundamental frequency contour features, duration features, and energy features as sample data.

[0045] In one embodiment, the manner of determining the dialect type to which the speech data belongs according to the pronunciation type and the tone may be: obtaining a pre-stored mapping relationship table between the pronunciation type, the tone, and the dialect type, and determining the dialect type to which the speech data belongs according to the mapping relationship table, the pronunciation type, and the tone, that is, querying the mapping relationship table based on the pronunciation type and the tone to obtain the dialect type corresponding to the pronunciation type and the tone, and using the dialect type corresponding to the pronunciation type and the tone as the dialect type to which the speech data belongs. Among them, the pre-stored mapping relationship table between the pronunciation type, the tone, and the dialect type can be set according to the actual situation, and the present application does not make specific limitations thereon.

[0046] In one embodiment, Mel cepstral features, fundamental frequency contour features, duration features, and energy features of each syllable segment are extracted from the speech data; the Mel cepstral features of each syllable segment are input into a preset pronunciation type recognition model to obtain the pronunciation type of each syllable segment; the pronunciation type of each syllable segment is input into a first preset dialect type recognition model to obtain the first probability corresponding to each dialect type to which the speech data belongs; the fundamental frequency contour features, duration features, and energy features of each syllable segment are input into a second preset dialect type recognition model to obtain the second probability corresponding to each dialect type to which the speech data belongs; and the dialect type to which the speech data belongs is determined according to the first probability and the second probability corresponding to each dialect type to which the speech data belongs. The dialect type to which the speech data belongs can be accurately determined through the Mel cepstral features, fundamental frequency contour features, duration features, and energy features of each syllable segment.

[0047] In one embodiment, the preset pronunciation type recognition model is trained by using the pronunciation type and the mel cepstrum features for a Gaussian mixture model. The first preset dialect type recognition model is obtained by training a three-layer neural network with the pronunciation type, the positional relationship of the pronunciation type, and the probability of the dialect type to which the syllable segment belongs as sample data. The second preset dialect type recognition model is trained by using the fundamental frequency contour features, the duration features, the energy features, and the probability of the dialect type to which the tone corresponding to the fundamental frequency contour features, the duration features, and the energy features belongs for a Gaussian mixture model. The fundamental frequency contour features, the duration features, and the energy features can better describe the features such as the tone pattern and the persistence of the tone.

[0048] In one embodiment, the three-layer neural network includes an observation layer, a hidden layer, and an output layer. The observation layer is the mel cepstrum features of the syllable segment, the hidden layer is the pronunciation type segment corresponding to the mel cepstrum features, and it is agreed that from top to bottom, it corresponds to the pronunciation types under the initial consonant, the onset and the nucleus, and the coda. The output layer is the dialect classification, so as to output the first probability corresponding to each dialect type to which the voice data belongs respectively.

[0049] Among them, the syllable segment is the three syllable segments obtained by dividing each syllable according to the initial consonant and the final. The pronunciation type corresponding to the position of the first syllable segment is stop, fricative, affricate, nasal, and lateral. The pronunciation type corresponding to the position of the second syllable segment is open syllable, front vowel syllable, rounded vowel syllable, and retroflex vowel syllable. The pronunciation type corresponding to the position of the third syllable segment is stop, fricative, and nasal. Modern phonology believes that tone, initial consonant, and final are the basic elements that make up a Chinese syllable. If the tone is not considered, the phoneme composition of a Chinese syllable is a four-position structure. Among them, the initial consonant occupies the first position, and the final is further divided into the onset, the nucleus, and the coda, occupying the second, third, and fourth positions. According to the articulation method, the initial consonants can be divided into five types of pronunciation, namely stop, fricative, affricate, nasal, and lateral. In the final, according to the combination of the onset and the nucleus, it can be divided into four types of pronunciation, namely open syllable, front vowel syllable, rounded vowel syllable, and retroflex vowel syllable. And the coda in the final can be divided into three types of pronunciation, namely stop, fricative, and nasal. Thus, a Chinese character syllable is composed of 3 pronunciation types, and the differences in Chinese dialects can be summarized as the frequencies of different pronunciation types and the order in which different pronunciation types appear in the syllable.

[0050] In one embodiment, the manner of determining the dialect type to which the voice data belongs according to the first probability and the second probability corresponding to each dialect type to which the voice data belongs respectively may be: according to the first probability and the second probability corresponding to each dialect type to which the voice data belongs respectively, calculate the average probability corresponding to each dialect type to which the voice data belongs respectively, and use the average probability corresponding to each dialect type to which the voice data belongs respectively as the target probability corresponding to each dialect type to which the voice data belongs respectively; use the dialect type with the maximum target probability as the dialect type to which the voice data belongs.

[0051] Sub-step S1032: Obtain a target speech recognition model according to the dialect type and the updated historical location information.

[0052] In an embodiment, obtain a first regional code bound to the dialect type, and obtain a second regional code of the region where the updated historical location information is located; according to the first regional code and the second regional code, obtain the target speech recognition model, that is, determine whether the first regional code is the same as the second regional code. When it is determined that the first regional code is the same as the second regional code, obtain the speech recognition model bound to the second regional code to obtain the target speech recognition model. When it is determined that the first regional code is different from the second regional code, obtain the speech recognition model bound to the first regional code to obtain the target speech recognition model. Among them, the method for obtaining the second regional code of the region where the updated historical location information is located can be: determine the latitude and longitude range where the updated historical location information is located, and obtain the regional code bound to the latitude and longitude range, and use the regional code bound to the latitude and longitude range as the second regional code.

[0053] Step S104: Recognize the speech data according to the target speech recognition model to obtain the standard Mandarin text corresponding to the speech data.

[0054] After determining the target speech recognition model, first convert the speech data into a text with dialect according to the target speech recognition model, query the corresponding relationship between the dialect text and the standard Mandarin text, obtain the standard Mandarin characters corresponding to each dialect character in the dialect text, and replace each dialect character in the dialect text with the corresponding standard Mandarin character, so as to obtain the standard Mandarin text corresponding to the speech data. Among them, the corresponding relationship between the dialect text and the standard Mandarin text is pre-stored in the storage of the mobile terminal, and the corresponding relationship between the dialect text and the standard Mandarin text can be set according to the actual situation, and the present application does not make specific limitations on this.

[0055] The voice recognition method provided by the above embodiments determines whether the mobile terminal is in a positioning blind area at intervals of a first preset time, and when it is determined that the mobile terminal is in a positioning blind area, updates the historical position information of the mobile terminal based on the sensor data of the mobile terminal at intervals of a second preset time. Then, when the voice data of the user is obtained, if the mobile terminal is still in the positioning blind area, a target voice recognition model is obtained according to the updated historical position information, and the voice data is recognized according to the target voice recognition model to obtain the standard Mandarin text corresponding to the voice data. The above technical solution can still locate the position information of the mobile terminal when the position information of the mobile terminal cannot be located according to the GPS positioning device, the base station or the mobile network, so as to match an accurate voice recognition model, improve the accuracy of voice recognition, and greatly improve the user experience.

[0056] Please refer to Figure 3 , Figure 3 which is a schematic block diagram of a voice recognition device provided by an embodiment of the present application.

[0057] As Figure 3 shown, the voice recognition device 200 includes: a determination module 210, a position update module 220, an acquisition module 230, and a voice recognition module 240, where:

[0058] The determination module 210 is configured to determine whether the mobile terminal is in a positioning blind area at intervals of a first preset time;

[0059] The position update module 220 is configured to, when it is determined that the mobile terminal is in a positioning blind area, update the historical position information of the mobile terminal based on the sensor data of the mobile terminal at intervals of a second preset time;

[0060] The acquisition module 230 is configured to, when the voice data of the user is obtained, if the mobile terminal is still in the positioning blind area, obtain a target voice recognition model according to the updated historical position information;

[0061] The voice recognition module 240 is configured to recognize the voice data according to the target voice recognition model to obtain the standard Mandarin text corresponding to the voice data.

[0062] In one embodiment, the sensor data includes the acceleration of the mobile terminal output by the acceleration sensor of the mobile terminal and the moving direction of the mobile terminal output by the direction sensor; the position update module 220 is further configured to:

[0063] Determine the moving distance of the mobile terminal according to the acceleration and the second preset time;

[0064] Update the historical position information of the mobile terminal according to the moving distance and the moving direction.

[0065] In one embodiment, the location update module 220 is further configured to:

[0066] Mark the historical location information of the mobile terminal in a preset offline map to obtain the historical location points of the mobile terminal;

[0067] Determine the current location point of the mobile terminal on the preset offline map according to the moving distance, moving direction, and historical location points;

[0068] Obtain the location information of the current location point on the preset offline map, and use the location information as the updated historical location information.

[0069] In one embodiment, the obtaining module 230 is further configured to:

[0070] Obtain the regional code of the region where the updated historical location information is located;

[0071] Obtain the speech recognition model bound to the regional code to obtain the target speech recognition model.

[0072] In one embodiment, as Figure 4 shown, the obtaining module 230 includes:

[0073] A determination sub-module 231, configured to determine the dialect type to which the voice data belongs;

[0074] An obtaining sub-module 232, configured to obtain the target speech recognition model according to the dialect type and the updated historical location information.

[0075] In one embodiment, the determination sub-module 231 is further configured to:

[0076] Extract Mel cepstrum features, fundamental frequency contour features, duration features, and energy features from the voice data;

[0077] Input the Mel cepstrum features into a preset pronunciation type recognition model to obtain the pronunciation type of each syllable segment;

[0078] Input the fundamental frequency contour features, duration features, and energy features into a preset tone recognition model to obtain the tone of each syllable segment;

[0079] Determine the dialect type to which the voice data belongs according to the pronunciation type and tone.

[0080] In one embodiment, the obtaining sub-module 232 is further configured to:

[0081] Obtain the first regional code bound to the dialect type, and obtain the second regional code of the region where the updated historical location information is located;

[0082] Obtain a target speech recognition model according to the first regional code and the second regional code.

[0083] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module and unit can refer to the corresponding processes in the foregoing embodiments of the speech recognition method, and will not be elaborated herein.

[0084] The device provided in the foregoing embodiment can be implemented in the form of a computer program, and the computer program can run on a Figure 5 mobile terminal as shown.

[0085] Please refer to Figure 5 , Figure 5 which is a schematic block diagram of the structure of a mobile terminal provided in an embodiment of the present application. The mobile terminal can be a server or a terminal.

[0086] As Figure 5 shown, the mobile terminal includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.

[0087] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any speech recognition method.

[0088] The processor is used to provide computing and control capabilities to support the operation of the entire mobile terminal.

[0089] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any data leakage reminder method.

[0090] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 5 the structure shown in

[0091] It should be understood that the processor may be a Central Processing Unit (CPU), and the processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0092] Those skilled in the art can understand that Figure 5 The structure shown in [the figure] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the mobile terminal to which the solution of this application is applied. The specific mobile terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0093] It should be understood that the processor may be a Central Processing Unit (CPU), and the processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0094] Among them, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0095] Determine whether the mobile terminal is in a positioning blind area at intervals of a first preset time;

[0096] When it is determined that the mobile terminal is in a positioning blind area, update the historical location information of the mobile terminal based on the sensor data of the mobile terminal at intervals of a second preset time;

[0097] When voice data of the user is obtained, if the mobile terminal is still in a positioning blind area, obtain a target speech recognition model according to the updated historical location information;

[0098] Perform speech recognition on the speech data according to the target speech recognition model to obtain the standard Mandarin text corresponding to the speech data.

[0099] In one embodiment, the sensor data includes the acceleration of the mobile terminal output by the acceleration sensor of the mobile terminal and the moving direction of the mobile terminal output by the direction sensor; updating the historical location information of the mobile terminal based on the sensor data of the mobile terminal includes:

[0100] Determine the moving distance of the mobile terminal according to the acceleration and the second preset time;

[0101] Update the historical location information of the mobile terminal according to the moving distance and the moving direction.

[0102] In one embodiment, updating the historical location information of the mobile terminal according to the moving distance and the moving direction includes:

[0103] Mark the historical location information of the mobile terminal in the preset offline map to obtain the historical location point of the mobile terminal;

[0104] Determine the current location point of the mobile terminal on the preset offline map according to the moving distance, the moving direction and the historical location point;

[0105] Obtain the location information of the current location point on the preset offline map, and use the location information as the updated historical location information.

[0106] In one embodiment, obtaining the target speech recognition model according to the updated historical location information includes:

[0107] Obtain the regional code of the region where the updated historical location information is located;

[0108] Obtain the speech recognition model bound to the regional code to obtain the target speech recognition model.

[0109] In one embodiment, obtaining the target speech recognition model according to the updated historical location information includes:

[0110] Determine the dialect type to which the speech data belongs;

[0111] Obtain the target speech recognition model according to the dialect type and the updated historical location information.

[0112] In one embodiment, obtaining the target speech recognition model according to the dialect type and the updated historical location information includes:

[0113] Obtain a first regional code bound to the dialect type, and obtain a second regional code of the region where the updated historical location information is located;

[0114] Obtain a target speech recognition model according to the first regional code and the second regional code.

[0115] In one embodiment, determining the dialect type to which the speech data belongs includes:

[0116] Extract Mel cepstrum features, fundamental frequency contour features, duration features, and energy features from the speech data;

[0117] Input the Mel cepstrum features into a preset pronunciation type recognition model to obtain the pronunciation type of each syllable segment;

[0118] Input the fundamental frequency contour features, duration features, and energy features into a preset tone recognition model to obtain the tone of each syllable segment;

[0119] Determine the dialect type to which the speech data belongs according to the pronunciation type and the tone.

[0120] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described mobile terminal can refer to the corresponding process in the foregoing speech recognition method embodiments, and will not be elaborated herein.

[0121] From the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a mobile terminal (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0122] This application embodiment also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. The computer program includes program instructions. The method implemented when the program instructions are executed can refer to various embodiments of the speech recognition method of this application.

[0123] Among them, the computer-readable storage medium may be the internal storage unit of the mobile terminal described in the foregoing embodiments, such as the hard disk or memory of the mobile terminal. The computer-readable storage medium may also be an external storage device of the mobile terminal, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the mobile terminal.

[0124] Furthermore, the computer-readable storage medium may mainly include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the blockchain node, etc.

[0125] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, essentially a decentralized database, is a series of data blocks generated by using cryptographic methods. Each data block contains information on a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain may include a blockchain underlying platform, a platform product service layer, an application service layer, etc.

[0126] It should be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0127] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations. It should be noted that in this article, the term "comprise", "include" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.

[0128] The serial numbers of the embodiments of the present application above are for description only and do not represent the superiority or inferiority of the embodiments. The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A speech recognition method, characterized in that, Applied to a mobile terminal, the method includes: Determining whether the mobile terminal is in a positioning blind area at intervals of a first preset time; When it is determined that the mobile terminal is in a positioning blind area, updating the historical location information of the mobile terminal based on the sensor data of the mobile terminal at intervals of a second preset time; When voice data of a user is obtained, attempting to locate the mobile terminal through the GPS positioning device, base station positioning program and network signal strength of the mobile terminal. When the mobile terminal cannot be located through the GPS positioning device, the base station positioning program and the network signal strength, determining that the mobile terminal is still in a positioning blind area, and then obtaining a target speech recognition model according to the updated historical location information; Recognizing the voice data according to the target speech recognition model to obtain a standard Mandarin text corresponding to the voice data; Among them, the determining whether the mobile terminal is in a positioning blind area at intervals of a first preset time includes: attempting to locate the mobile terminal through the GPS positioning device, base station positioning program and network signal strength of the mobile terminal at intervals of a first preset time; when the mobile terminal cannot be located through the GPS positioning device, the base station positioning program and the network signal strength, determining that the mobile terminal is in a positioning blind area; The obtaining a target speech recognition model according to the updated historical location information includes: determining the dialect type to which the voice data belongs; obtaining a first regional code bound to the dialect type, and obtaining a second regional code of the region where the updated historical location information is located; determining whether the first regional code is the same as the second regional code; when it is determined that the first regional code is the same as the second regional code, obtaining a speech recognition model bound to the second regional code to obtain a target speech recognition model; when it is determined that the first regional code is different from the second regional code, obtaining a speech recognition model bound to the first regional code to obtain a target speech recognition model.

2. The speech recognition method according to claim 1, wherein The sensor data includes the acceleration of the mobile terminal output by the acceleration sensor of the mobile terminal and the moving direction of the mobile terminal output by the direction sensor. The updating the historical location information of the mobile terminal based on the sensor data of the mobile terminal includes: Determining the moving distance of the mobile terminal according to the acceleration and the second preset time; Updating the historical location information of the mobile terminal according to the moving distance and the moving direction.

3. The voice recognition method according to claim 2, wherein The updating the historical location information of the mobile terminal according to the moving distance and the moving direction includes: Marking the historical location information of the mobile terminal in a preset offline map to obtain a historical location point of the mobile terminal; Determining a current location point of the mobile terminal on the preset offline map according to the moving distance, the moving direction and the historical location point; Obtaining the location information of the current location point on the preset offline map and using the location information as the updated historical location information.

4. The voice recognition method according to claim 1, wherein The determining the dialect type to which the voice data belongs includes: Extract Mel cepstral features, fundamental frequency contour features, duration features, and energy features from the speech data; Input the Mel cepstral features into a preset pronunciation type recognition model to obtain the pronunciation type of each syllable segment; Input the fundamental frequency contour features, duration features, and energy features into a preset tone recognition model to obtain the tone of each syllable segment; Determine the dialect type to which the speech data belongs according to the pronunciation type and tone.

5. A voice recognition device, characterized in that, The speech recognition device includes: A determination module for determining whether the mobile terminal is in a positioning blind area at intervals of a first preset time; A position update module for, when it is determined that the mobile terminal is in a positioning blind area, updating the historical position information of the mobile terminal based on the sensor data of the mobile terminal at intervals of a second preset time; An acquisition module for, when acquiring the speech data of the user, attempting to locate the mobile terminal through the GPS positioning device, base station positioning program, and network signal strength of the mobile terminal. When it is impossible to locate the mobile terminal through the GPS positioning device, the base station positioning program, and the network signal strength, determine that the mobile terminal is still in a positioning blind area, and then obtain a target speech recognition model according to the updated historical position information; A speech recognition module for recognizing the speech data according to the target speech recognition model to obtain the standard Mandarin text corresponding to the speech data; Among them, the determination module is further configured to attempt to locate the mobile terminal through the GPS positioning device, base station positioning program, and network signal strength of the mobile terminal at intervals of a first preset time; when it is impossible to locate the mobile terminal through the GPS positioning device, the base station positioning program, and the network signal strength, determine that the mobile terminal is in a positioning blind area; The acquisition module includes: A determination sub-module for determining the dialect type to which the speech data belongs; An acquisition sub-module for acquiring a first regional code bound to the dialect type and acquiring a second regional code of the region where the updated historical position information is located; determining whether the first regional code is the same as the second regional code; when it is determined that the first regional code is the same as the second regional code, acquiring a speech recognition model bound to the second regional code to obtain a target speech recognition model; when it is determined that the first regional code is different from the second regional code, acquiring a speech recognition model bound to the first regional code to obtain a target speech recognition model.

6. A mobile terminal, characterized in that, The mobile terminal includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the speech recognition method according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the steps of the speech recognition method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Acoustic model adaptation using geographic information

    CN104575493A

  • Language translation method and system based on geographical location information

    CN105912532A

  • Method and device for identifying dialect types

    CN108877769A

  • Voice control method and terminal device

    CN109509473A

  • Positioning method and device

    CN109756837A