Speech recognition method, apparatus, device, vehicle and medium

By combining acoustic and linguistic features with the client's hot word library, the problem of inaccurate recognition caused by similar sounds in voice navigation has been solved, achieving higher recognition accuracy.

CN119068869BActive Publication Date: 2026-05-26BEIJING CO WHEELS TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CO WHEELS TECH CO LTD
Filing Date
2023-05-31
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, when voice navigation commands contain similar sounds, the navigation location cannot be accurately identified, resulting in low recognition accuracy.

Method used

By combining acoustic and linguistic features and using the client's preset hot word library, the target hot word score for voice navigation information is determined. The final recognition result is determined by integrating the acoustic score, linguistic score, and hot word score.

Benefits of technology

It improves the recognition accuracy of voice navigation commands, effectively removes the influence of similar sounds, and enhances the precision of recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119068869B_ABST
    Figure CN119068869B_ABST
Patent Text Reader

Abstract

This disclosure relates to a speech recognition method, apparatus, device, vehicle, and medium. The method includes: performing speech recognition on voice navigation information to obtain at least one candidate recognition result; determining an acoustic score for each candidate recognition result and a language score for each candidate recognition result; identifying target hot words contained in each candidate recognition result according to a preset hot word library and obtaining a target hot word score for the predetermined target hot words; determining a reference score for each candidate recognition result based on the acoustic score, language score, and target hot word score, and determining the target recognition result based on the reference score. In the embodiments of this disclosure, based on the speech recognition results combining language and acoustics, client-side hot words are introduced to determine the final recognition result. Since client-side hot words reflect personalized search habits, the influence of similar sounds can be removed, further improving the accuracy of the recognition results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence, and more particularly to a speech recognition method, apparatus, device, vehicle, and medium. Background Technology

[0002] Voice recognition technology, such as voice navigation commands, accounts for a very high proportion of in-vehicle voice interaction. In voice navigation command recognition, the accuracy of navigation point of interest (POI) recognition has a crucial impact on the success rate of voice navigation initiation.

[0003] In related technologies, a correspondence is established between regions and language recognition models. The language recognition model can be used to distinguish between similar pronunciations in the corresponding region and other regions due to accent issues. The region to which the client belongs is determined based on the latitude and longitude of the initiated voice navigation. The language recognition model of the region is determined based on the above correspondence. Navigation points are identified based on the language recognition model of the region.

[0004] However, the aforementioned technologies can only distinguish between similar pronunciations in the corresponding region and other regions due to accent issues. When the voice navigation command contains similar sounds, it cannot accurately identify the results. For example, for a region that includes two locations, "Yindu Building" and "Yingdu Building", it obviously cannot accurately identify the voice navigation command containing "yindudasha". Summary of the Invention

[0005] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this disclosure provides a speech recognition method, apparatus, device, vehicle and medium to solve the technical problem in the prior art that the recognition results cannot remove the influence of similar sounds, resulting in low recognition accuracy.

[0006] This disclosure provides a speech recognition method, the method comprising: responding to received voice navigation information from a client, performing speech recognition on the voice navigation information to obtain at least one candidate recognition result; determining an acoustic score for each candidate recognition result, and determining a language score for each candidate recognition result; identifying target hot words contained in each candidate recognition result according to a preset hot word library of the client, and obtaining a predetermined target hot word score for the target hot words, wherein the target hot word score is determined based on the total number of times the target hot word is added to the preset hot word library; determining a reference score for each candidate recognition result based on the acoustic score, the language score, and the target hot word score; and determining a target recognition result from the at least one candidate recognition result based on the reference score.

[0007] This disclosure also provides a speech recognition device, comprising: a recognition module, configured to respond to received voice navigation information from a client, and perform speech recognition on the voice navigation information to obtain at least one candidate recognition result; a first determining module, configured to determine an acoustic score for each candidate recognition result; a second determining module, configured to determine a language score for each candidate recognition result; an acquisition module, configured to identify target hot words contained in each candidate recognition result according to a preset hot word library of the client, and acquire a predetermined target hot word score for the target hot words, wherein the target hot word score is determined based on the total number of times the target hot word is added to the preset hot word library; a third determining module, configured to determine a reference score for each candidate recognition result based on the acoustic score, the language score, and the target hot word score; and a fourth determining module, configured to determine a target recognition result among the at least one candidate recognition result based on the reference score.

[0008] This disclosure also provides an electronic device, which includes: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the speech recognition method provided in this disclosure.

[0009] This disclosure also provides a vehicle that includes the voice recognition device described in the above embodiments, or an electronic device as described in the above embodiments.

[0010] This disclosure also provides a computer-readable storage medium storing a computer program for performing the speech recognition method provided in this disclosure.

[0011] The technical solution provided in this disclosure has the following advantages compared with the prior art:

[0012] The speech recognition scheme provided in this disclosure responds to received voice navigation information from a client, performs speech recognition on the voice navigation information to obtain at least one candidate recognition result, determines the acoustic score and the language score of each candidate recognition result, and then identifies the target hot words contained in each candidate recognition result according to the client's preset hot word library, and obtains the target hot word score of the predetermined target hot word. The target hot word score is determined based on the total number of times the target hot word is added to the client's preset hot word library. A reference score is determined for each candidate recognition result based on the acoustic score, language score, and target hot word score. The target recognition result is then determined from at least one candidate recognition result based on the reference score. In this technical solution, based on the speech recognition results obtained by combining language and acoustics, the client's hot words are introduced to determine the final recognition result. Since the client's hot words reflect personalized search habits, the influence of similar sounds can be removed, further improving the accuracy of the recognition results. Attached Figure Description

[0013] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0014] Figure 1 A flowchart illustrating a speech recognition method provided in an embodiment of this disclosure;

[0015] Figure 2 A schematic diagram of a speech recognition functional module provided in an embodiment of this disclosure;

[0016] Figure 3 This is a schematic diagram of the word segmentation node structure corresponding to the word segmentation language score provided in an embodiment of the present disclosure;

[0017] Figure 4 This is a schematic diagram of the structure of a speech recognition device provided in an embodiment of the present disclosure;

[0018] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0019] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0020] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0021] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0022] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0023] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0024] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0025] To address the aforementioned issues, this disclosure provides a speech recognition method that combines acoustics and language to determine the target recognition result, thereby initially ensuring recognition accuracy. Acoustics refers to physical quantities representing the acoustic characteristics of speech, and is also a general term for the acoustic performance of various sound elements, such as the energy concentration area, formant frequency, formant intensity, and bandwidth representing timbre, and duration, fundamental frequency, and average speech power representing the prosodic characteristics of speech. Language reflects the word formation features of speech. Furthermore, the speech recognition method incorporates client-side hot words to determine the final recognition result. Since client-side hot words reflect the client's personalized search habits, recognition enhancement can be performed. For example, it ensures accurate recognition of navigation voice commands even in the presence of similar sounds or accents.

[0026] The method will be described below with reference to specific embodiments.

[0027] Figure 1This is a flowchart illustrating a speech recognition method provided in an embodiment of the present disclosure. The method can be executed by a speech recognition device, which can be implemented using software and / or hardware, and is generally integrated into an electronic device. Figure 1 As shown, the method includes:

[0028] Step 101: In response to receiving the voice navigation information from the client, perform voice recognition on the voice navigation information to obtain at least one candidate recognition result.

[0029] In different application scenarios, the way to receive voice navigation information is different. In some possible embodiments, voice navigation information can be received through the vehicle microphone.

[0030] In one embodiment of this disclosure, in response to receiving voice navigation information from a client, voice recognition is performed on the voice navigation information to obtain at least one candidate recognition result. The candidate recognition result may be considered to include the recognition result of the navigation location. The at least one candidate recognition result may include near-homophones, words with similar accents, etc. For example, when the candidate voice navigation information contains "yingdudasha", the at least one candidate recognition result may include "Yindu Building" and "Yingdu Building".

[0031] It should be noted that the methods for obtaining at least one candidate recognition result through voice navigation information speech recognition differ in different application scenarios. Examples are illustrated below:

[0032] In some possible examples, speech recognition can be performed on navigation voice information, a preset speech table can be queried, which contains recognition results and the corresponding speech, and the recognition results that match the speech in the preset speech table with a preset matching degree threshold can be identified as candidate recognition results.

[0033] Step 102: Determine the acoustic score for each candidate recognition result and the language score for each candidate recognition result.

[0034] It is easy to understand that language scores can determine the correlation between characters in candidate recognition results based on linguistic perspectives such as word formation, while acoustic scores can determine the confidence level of candidate recognition results from acoustic perspectives such as syllables.

[0035] In the embodiments of this disclosure, an acoustic score and a language score are determined for each candidate recognition result, so as to combine the acoustic score and the language score to determine the corresponding target recognition result. Thus, the target recognition result is determined by combining acoustic characteristics and language characteristics, which ensures the accuracy of the target recognition result.

[0036] It should be noted that the method for determining the acoustic score of each candidate recognition result differs in different application scenarios. Examples are provided below:

[0037] In some possible embodiments, candidate acoustic feature information can be extracted for each candidate recognition result. The candidate acoustic feature information can be any information representing the candidate recognition result in an acoustic dimension. The acoustic feature information is matched with multiple preset standard acoustic feature information, and the standard acoustic feature information with the highest matching degree is determined as the successfully matched target standard acoustic feature information. Furthermore, a preset acoustic score corresponding to the target standard acoustic feature information is determined as the acoustic score for each candidate recognition result.

[0038] In some possible embodiments, such as Figure 2 As shown, candidate recognition results can be input into a pre-trained acoustic model, which can then calculate the acoustic score of the candidate recognition results.

[0039] Similarly, the method for determining the language score of each candidate recognition result differs in different application scenarios:

[0040] In some possible embodiments, at least one candidate word corresponding to each candidate recognition result is identified. For example, the candidate recognition result can be matched with a preset word segmentation vocabulary, and at least one corresponding candidate word can be determined based on the matching result. The word segmentation language score of each candidate word can be determined. For example, each word in the preset word segmentation vocabulary is set with a corresponding word segmentation language score, and the word segmentation language score of each candidate word can be determined based on the preset word segmentation language score.

[0041] For example, a pre-defined word segmentation node graph can be set up, containing multiple word segmentation nodes connected by node edges. Each connected word segmentation node has a corresponding word segmentation language score, which is the confidence score of the corresponding word segmentation. The word segmentation language score can be obtained based on experimental data, for example... Figure 3 As shown, when the word segmentation node graph contains the words a, b, c, d, e, f, and g, a and b are connected by node edge L1, b and c by node edge L2, b and d by node edge L3, d and e by node edge L4, a and f by node edge L5, and f and g by node edge L6. The word segments a, b, c, d, e, f, and g have pre-set corresponding word segmentation language scores q1, q2, q3, q4, q5, q6, and q7. In this embodiment, the word segmentation language score of each candidate word is obtained by querying based on the word segmentation language scores contained in the word segmentation node graph.

[0042] Furthermore, after determining the segmentation language score of each candidate segmentation word, the sum of the segmentation language scores corresponding to at least one candidate segmentation word is calculated to determine the language score of each candidate recognition result.

[0043] In some possible embodiments, reference continues to be made to Figure 2 Further, the candidate recognition results can be input into a pre-trained language model, and the language model can calculate the language scores of the candidate recognition results.

[0044] Step 103: Identify the target hot words included in each candidate recognition result according to the preset hot word library of the client, and obtain the target hot word scores of the predetermined target hot words. The target hot word scores are determined according to the total number of times the target hot words are added to the preset hot word library.

[0045] The preset hot word library of the client contains multiple candidate hot words, and each candidate hot word can be understood as a relatively common navigation location of the client, etc. Therefore, even if there are two candidate recognition results with similar pronunciations, denoising can be performed based on the preset hot word library to avoid errors caused by similar sounds. For example, when the candidate voice navigation information contains "yingdudasha", at least one candidate recognition result may include "Yindu Building" and "Yingdu Building". However, "Yindu Building" belongs to the preset hot word library, and the target hot word score corresponding to "Yindu Building" will enhance its subsequent reference score for easy distinction from "Yingdu Building".

[0046] It should be noted that in different application scenarios, the method for obtaining the preset hot word library of the client is different, and examples are as follows:

[0047] In some possible embodiments, the historical navigation locations of the client can be obtained. The historical navigation locations can be all historical navigation locations in the navigation platform within a preset recent period, etc. The historical search parameters of the historical navigation locations are obtained, and the historical navigation locations corresponding to the historical search parameters that meet the preset hot word search conditions are determined as candidate hot words, and the preset hot word library of the client is generated according to all the candidate hot words.

[0048] Among them, in different application scenarios, the above historical search parameters are different, and the corresponding preset hot word search conditions are different. In some possible embodiments, the historical search parameter includes the historical navigation times. When the historical navigation times are greater than the preset times threshold, it is determined that the corresponding historical search parameter meets the preset hot word search conditions, and the relatively common navigation locations of the client are used as candidate hot words;

[0049] In some possible embodiments, the historical search parameter includes the input method of the historical navigation location. When the input method is the preset input method, it is determined that the historical search parameter meets the preset hot word search conditions. It can be understood that the preset input method can be an input method that can accurately reflect the search intention of the client, and the historical navigation locations input by the preset input method have a higher accuracy rate than the historical navigation locations input by voice.

[0050] For example, when the historical navigation location input based on the voice navigation command input by the client is inaccurate, the client usually inputs the accurate historical navigation location through other input methods, such as by calling the keyboard or by selecting the accurate historical navigation location on the map. Therefore, in this embodiment, the preset input method can be keyboard input or map selection input, etc.

[0051] In some possible embodiments, with client authorization, the client's address book location and other information can also be read and added to the hot word library as candidate hot words.

[0052] In the actual implementation process, the hot word score of each candidate hot word in the hot word library is also determined in advance.

[0053] In one embodiment of this disclosure, a preset language model can be pre-trained. The training sample data of the language model can be obtained from navigation locations input by other clients belonging to the same region as the client. In this embodiment, for candidate hot words added to the hot word library for the first time, they are input into the corresponding preset language model to obtain the hot word score of each candidate hot word. For candidate hot words that are not added to the hot word library for the first time, in order to enhance the recognition of the hot word, the hot word weight value of the candidate hot word that is not added for the first time is determined, and the hot word score of the candidate hot word corresponding to the addition operation is updated according to the hot word weight value.

[0054] In some possible implementations, the aforementioned hot word weight value is used, and the hot word weight of each candidate hot word is determined based on the total number of times each candidate hot word is added to the hot word database. The higher the total number of times, the higher the corresponding hot word weight, thereby further enhancing the weight of more commonly used candidate hot words. The reference hot word score of the corresponding candidate hot word is updated based on the hot word weight. Specifically, the third product value of the hot word weight and the current hot word score of the corresponding candidate hot word can be calculated, and the hot word score of the corresponding candidate hot word is updated based on the third product value.

[0055] In actual implementation, a preset unit weight increment value can be used. Each time a candidate hot word is detected before being added to the preset hot word library, this preset unit weight increment value is used as the hot word weight, and the hot word score of the corresponding candidate hot word is updated in real time based on this value. Specifically, the third product of the preset unit weight increment value and the current hot word score is calculated, and the hot word score of the corresponding candidate hot word is updated based on this third product value. The preset unit weight increment value can be set according to the needs of the scenario and is not limited here.

[0056] In one embodiment of this disclosure, after obtaining the above-mentioned hot word library, the target hot words contained in each candidate recognition result are obtained according to the preset hot word library, and the hot word score of the predetermined target hot words is obtained.

[0057] Step 104: Determine the reference score for each candidate recognition result based on the acoustic score, language score, and target hot word score.

[0058] In one embodiment of this disclosure, a reference score for each candidate recognition result is determined based on the acoustic score, language score, and target hot word score. This comprehensively considers the acoustic characteristics, language characteristics, and personalized preferences of the candidate recognition results, thus ensuring the accuracy of the final recognition result.

[0059] It should be noted that the method for determining the reference score for each candidate recognition result based on acoustic score, language score, and target hot word score differs in different application scenarios, as shown in the following example:

[0060] In some possible examples, the acoustic score is normalized to obtain a first normalized value. The language score and a preset first weight value are then calculated as a first product. The preset first weight value can be calibrated based on experimental data. The first product value is then normalized to obtain a second normalized value. The preset second weight value and the target hot word score are then calculated as a second product. The second product value is then normalized to obtain a third normalized value. In this example, the normalization method can be any normalization method in the prior art, which will not be listed here. Then, the first normalized value, the second normalized value, and the third normalized value are summed to obtain a reference score for each candidate recognition result.

[0061] In this embodiment, the first reference score can be obtained using the following formula (1), where Score(y) is the reference score, logP(y|x) is the first normalized value, and x is the voice navigation information, y is the candidate recognition result, and logP L (y) represents the language score, where L is the language model, λ is the first weight value, γ is the second weight value, and logP c (y) represents the target hot word score, γlogP c (y) is the third normalized value:

[0062] Score(y) = logP(y|x) + λlogP L (y)+γlogP c (y)Formula (1)

[0063] In some possible examples, the acoustic score, the language score, and the target hot word score can be normalized separately to obtain three normalized values, and the three normalized values ​​can be summed to obtain a reference score.

[0064] Step 105: Determine the target recognition result from at least one candidate recognition result based on the reference score.

[0065] In one embodiment of this disclosure, the candidate recognition result corresponding to the maximum value among all reference scores is determined as the target recognition result, and thus, the process continues to refer to... Figure 2 In the embodiments disclosed herein, based on the recognition of speech recognition results by combining language and acoustics, the target recognition results are further determined by combining a preset hot word library that reflects the personalized search habits of the client, thereby removing the influence of similar sounds and improving the recognition accuracy of the target recognition results.

[0066] In summary, the speech recognition method provided in this disclosure responds to received voice navigation information from a client, performs speech recognition on the voice navigation information to obtain at least one candidate recognition result, determines the acoustic score and the language score of each candidate recognition result, and then identifies the target hot words contained in each candidate recognition result according to the client's preset hot word library, and obtains the target hot word score of the predetermined target hot words. The target hot word score is determined based on the total number of times the target hot word is added to the client's preset hot word library. A reference score is determined for each candidate recognition result based on the acoustic score, language score, and target hot word score. The target recognition result is then determined from at least one candidate recognition result based on the reference score. In this technical solution, based on the speech recognition results obtained by combining language and acoustics, the client's hot words are introduced to determine the final recognition result. Since the client's hot words reflect personalized search habits, the influence of similar sounds can be removed, further improving the accuracy of the recognition results.

[0067] To implement the above embodiments, this disclosure also proposes a speech recognition device.

[0068] Figure 4 This is a schematic diagram of a speech recognition device provided in an embodiment of the present disclosure. The device can be implemented by software and / or hardware, and is generally integrated into an electronic device for speech recognition. Figure 4 As shown, the device includes: an identification module 410, a first determination module 420, a second determination module 430, an acquisition module 440, a third determination module 450, and a fourth determination module 460, wherein,

[0069] The recognition module 410 is used to respond to the received voice navigation information from the client and perform voice recognition on the voice navigation information to obtain at least one candidate recognition result;

[0070] The first determining module 420 is used to determine the acoustic score of each candidate recognition result;

[0071] The second determining module 430 is used to determine the language score of each candidate recognition result;

[0072] The acquisition module 440 is used to identify the target hot words contained in each candidate identification result according to the client's preset hot word library, and to obtain the target hot word score of the predetermined target hot word, wherein the target hot word score is determined according to the total number of times the target hot word is added to the client's preset hot word library;

[0073] The third determination module 450 is used to determine a reference score for each candidate recognition result based on the acoustic score, language score and target hot word score;

[0074] The fourth determining module 460 is used to determine the target recognition result from at least one candidate recognition result based on the reference score.

[0075] The speech recognition device provided in this disclosure can execute the speech recognition method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0076] In some possible embodiments, the first determining module 420 is configured to:

[0077] Extract the candidate acoustic feature information for each candidate recognition result;

[0078] The candidate acoustic feature information is matched with the preset standard acoustic feature information to determine the target standard acoustic feature information that has been successfully matched;

[0079] The preset acoustic score corresponding to the target standard acoustic feature information is determined as the acoustic score of each candidate recognition result.

[0080] In some possible embodiments, the second determining module 430 is configured to:

[0081] Identify at least one candidate word segmentation corresponding to each candidate identification result;

[0082] Determine the segmentation language score for each candidate word, and calculate the sum of the segmentation language scores corresponding to at least one candidate word to determine the language score of each candidate recognition result.

[0083] In some possible embodiments, the third determining module 450 is configured to:

[0084] The acoustic score is normalized to obtain a first normalized value;

[0085] Calculate the first product of the language score and the preset first weight value, and normalize the first product value to obtain a second normalized value;

[0086] Calculate the second product value of the preset second weight value and the target hot word score, and normalize the second product value to obtain the third normalized value;

[0087] The first normalized value, the second normalized value, and the third normalized value are summed to obtain a reference score for each candidate identification result.

[0088] In some possible embodiments, the fourth determining module 460 is configured to:

[0089] The candidate recognition result corresponding to the maximum value among all reference scores is determined as the target recognition result.

[0090] In some possible embodiments, it further includes: a hot word determination module, used for:

[0091] Determine the client's historical navigation locations and obtain the historical search parameters for those locations;

[0092] The historical navigation locations corresponding to historical search parameters that meet the preset hot word search conditions are identified as candidate hot words;

[0093] The client's preset hot word library is generated based on all candidate hot words.

[0094] In some possible embodiments, the historical search parameters include the number of historical navigation attempts, and when the number of historical navigation attempts exceeds a preset threshold, the historical search parameters are determined to meet preset hot word search conditions; and / or,

[0095] Historical search parameters include the input method of historical navigation locations. When the input method is a preset input method, the historical search parameters are determined to meet the preset hot word search conditions.

[0096] In some possible embodiments, the hot word score acquisition module is used for:

[0097] In response to the addition operation of the hot words to the preset hot word library, determine whether the candidate hot words corresponding to the addition operation are added to the preset hot word library for the first time;

[0098] When adding a word to the preset hot word library for the first time, the hot word score of the candidate hot word corresponding to the addition operation is determined according to the preset language model; when adding a word to the preset hot word library for the first time, the hot word weight value of the hot word corresponding to the addition operation is determined, and the hot word score of the candidate hot word corresponding to the addition operation is updated according to the hot word weight value.

[0099] In some possible embodiments, the hot word score acquisition module is used for:

[0100] Determine the total number of times the corresponding hot words are added to the preset hot word library.

[0101] Determine the keyword weights corresponding to the total number of occurrences, where the keyword weights are directly proportional to the total number of occurrences; or,

[0102] The preset unit weight increase value is determined to be the hot word weight value.

[0103] In some possible embodiments, the hot word score acquisition module is used for:

[0104] Get the current hot word score of the candidate hot words corresponding to the add operation;

[0105] Calculate the third product of the current hot word score and the hot word weight;

[0106] The hot word score of the candidate hot word corresponding to the addition operation is updated based on the sum of the current hot word score and the third product value.

[0107] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program / instructions, which, when executed by a processor, implements the speech recognition method in the above embodiments.

[0108] To implement the above embodiments, this disclosure also proposes a vehicle that includes the voice recognition device in the above embodiments or the electronic device mentioned in the following embodiments.

[0109] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.

[0110] The following is a detailed reference. Figure 5 The diagram illustrates a structural schematic suitable for implementing the electronic device 500 in the embodiments of this disclosure. The electronic device 500 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0111] like Figure 5 As shown, the electronic device 500 may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a memory 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0112] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0113] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a memory 508, or installed from a ROM 502. When the computer program is executed by the processor 501, it performs the functions defined in the speech recognition method of embodiments of this disclosure.

[0114] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0115] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0116] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0117] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:

[0118] The system receives voice navigation information from a client, performs speech recognition on the information to obtain at least one candidate recognition result, determines the acoustic score and language score of each candidate result, and then identifies target hot words contained in each candidate result based on the client's preset hot word library, obtaining the target hot word score for each pre-determined target hot word. The target hot word score is determined based on the total number of times the target hot word has been added to the client's preset hot word library. A reference score is determined for each candidate result based on the acoustic score, language score, and target hot word score. The target recognition result is then determined from at least one candidate result based on the reference score. In this technical solution, based on the speech recognition results combining language and acoustics, the client's hot words are introduced to determine the final recognition result. Since the client's hot words reflect personalized search habits, the influence of similar sounds can be removed, further improving the accuracy of the recognition results.

[0119] Electronic devices can be programmed with computer program code in one or more programming languages ​​or combinations thereof to perform the operations of this disclosure. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0121] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0122] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0123] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0124] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0125] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0126] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A speech recognition method, characterized in that, include: Upon receiving voice navigation information from the client, the system performs voice recognition on the voice navigation information to obtain at least one candidate recognition result. The acoustic score of each candidate recognition result is determined, and at least one candidate word corresponding to each candidate recognition result is identified. A preset word segmentation node graph is queried to obtain the word segmentation language score of each candidate word. The sum of the word segmentation language scores corresponding to the at least one candidate word is calculated to determine the language score of each candidate recognition result. The word segmentation node graph contains multiple word segmentation nodes connected by node edges, and each connected word segmentation node has a corresponding word segmentation language score. The client identifies the target hot words contained in each candidate identification result according to the client's preset hot word library, and obtains the target hot word score of the predetermined target hot word, wherein the target hot word score is determined according to the total number of times the target hot word is added to the preset hot word library; A reference score is determined for each candidate recognition result based on the acoustic score, the language score, and the target hot word score; The target recognition result is determined from the at least one candidate recognition result based on the reference score.

2. The method as described in claim 1, characterized in that, Determining the acoustic score for each candidate recognition result includes: Extract the candidate acoustic feature information for each of the candidate recognition results; The candidate acoustic feature information is matched with the preset standard acoustic feature information to determine the target standard acoustic feature information that has been successfully matched; The preset acoustic score corresponding to the target standard acoustic feature information is determined as the acoustic score of each candidate recognition result.

3. The method as described in claim 1, characterized in that, The step of determining a reference score for each candidate recognition result based on the acoustic score, the language score, and the target hot word score includes: The acoustic score is normalized to obtain a first normalized value; Calculate the first product of the language score and the preset first weight value, and normalize the first product value to obtain a second normalized value; Calculate the second product value of the preset second weight value and the target hot word score, and normalize the second product value to obtain a third normalized value; The first normalized value, the second normalized value, and the third normalized value are summed to obtain a reference score for each candidate identification result.

4. The method according to any one of claims 1-3, characterized in that, Determining the target recognition result from the at least one candidate recognition result based on the reference score includes: The candidate recognition result corresponding to the maximum value among all the reference scores is determined as the target recognition result.

5. The method as described in any one of claims 1-3, characterized in that, Before identifying the target hot words contained in each candidate identification result according to the client's preset hot word library, the method further includes: Determine the client's historical navigation locations and obtain the historical search parameters for those locations; The historical navigation locations corresponding to historical search parameters that meet the preset hot word search conditions are identified as candidate hot words; The client's preset hot word library is generated based on all the candidate hot words.

6. The method as described in claim 5, characterized in that, The historical search parameters include the number of historical navigation attempts. When the number of historical navigation attempts exceeds a preset threshold, the historical search parameters are determined to meet the preset hot word search conditions; and / or, The historical search parameters include the input method of historical navigation locations. When the input method is a preset input method, the historical search parameters are determined to meet the preset hot word search conditions.

7. The method as described in claim 5, characterized in that, Before obtaining the target hot word score of the predetermined target hot word, the method further includes: In response to the addition operation of acquiring hot words and adding them to the preset hot word library, determine whether the candidate hot words corresponding to the addition operation are being added to the preset hot word library for the first time; When adding a word to the preset hot word library for the first time, the hot word score of the candidate hot word corresponding to the addition operation is determined according to the preset language model; when adding a word to the preset hot word library for the second time, the hot word weight value corresponding to the addition operation is determined, and the hot word score of the candidate hot word corresponding to the addition operation is updated according to the hot word weight value.

8. The method as described in claim 7, characterized in that, Determining the hot word weight value corresponding to the addition operation includes: Determine the total number of times the hot words corresponding to the addition operation are added to the preset hot word library. Determine the hot word weights corresponding to the total number of occurrences, wherein the hot word weights are directly proportional to the total number of occurrences; or, The preset unit weight increase value is determined to be the weight value of the hot word.

9. The method as described in claim 7 or 8, characterized in that, The step of updating the hot word score of the candidate hot word corresponding to the addition operation based on the hot word weight value includes: Obtain the current hot word score of the candidate hot word corresponding to the addition operation; Calculate the third product of the current hot word score and the hot word weight; The hot word score of the candidate hot word corresponding to the addition operation is updated based on the sum of the current hot word score and the third product value.

10. A voice recognition device, characterized in that, include: The recognition module is used to respond to the received voice navigation information from the client and perform voice recognition on the voice navigation information to obtain at least one candidate recognition result; The first determining module is used to determine the acoustic score of each of the candidate recognition results; The second determining module is used to identify at least one candidate word corresponding to each candidate recognition result, query a preset word segmentation node graph to obtain the word segmentation language score of each candidate word, and calculate the sum of the word segmentation language scores corresponding to the at least one candidate word to determine the language score of each candidate recognition result. The word segmentation node graph contains multiple word segmentation nodes connected by node edges, and each connected word segmentation node has a corresponding word segmentation language score. The acquisition module is used to identify the target hot words contained in each candidate identification result according to the preset hot word library of the client, and to obtain the target hot word score of the predetermined target hot word, wherein the target hot word score is determined according to the total number of times the target hot word is added to the preset hot word library; The third determining module is used to determine a reference score for each candidate recognition result based on the acoustic score, the language score, and the target hot word score; The fourth determining module is used to determine the target recognition result from the at least one candidate recognition result based on the reference score.

11. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the speech recognition method according to any one of claims 1-6.

12. A vehicle, characterized in that, The vehicles include: The voice recognition device as claimed in claim 10, or the electronic device as claimed in claim 11.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program for performing the speech recognition method according to any one of claims 1-6.