Identification result determination method and device, equipment, storage medium and vehicle

By sorting and text adjustment of the initial text results of speech recognition, the candidate text results are extended, the problem of low accuracy of speech recognition is solved, and the accuracy of recognition is improved.

CN120030140APending Publication Date: 2025-05-23BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311568904.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The current speech recognition technology has low accuracy, especially when dealing with homophones, similar sounds, and similar initial vowels, resulting in the speech segment not being correctly recognized.

Method used

By obtaining the initial text results of the voice clip, sorting processes are performed to determine the pending text results, and text adjustments are made to obtain extended text results. Then, the candidate text results including the initial text results and the extended text results are sorted to determine the target text results.

Benefits of technology

Through text adjustment and sorting processing, the number of potential text results for speech clips is widened, and the possibility of including correct results in text results is improved, thereby improving the accuracy of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030140A_ABST
    Figure CN120030140A_ABST
Patent Text Reader

Abstract

The invention relates to an identification result determination method and device, equipment, a storage medium and a vehicle. The method comprises the steps of obtaining initial text results generated by performing voice recognition on a voice segment, performing sorting processing on the initial text results, and determining a to-be-processed text result in the initial text results; performing text adjustment on the to-be-processed text result to obtain an extended text result; sorting the candidate text results, and determining a target text result in the candidate text results; wherein the candidate text result comprises an initial text result and an extended text result. According to the embodiment of the invention, the text result determined through voice recognition is expanded in the form of text adjustment, the number of potential text results of the voice segment is increased, and the possibility that the text results include correct text results is improved, so that the possibility that the target text result is a correct result is improved, and the user experience is improved. And the voice recognition accuracy of the voice segment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of speech recognition technology, and in particular to a method, device, equipment, storage medium and vehicle for determining a recognition result. Background Art

[0002] With the development of speech recognition technology, the application scenarios of speech recognition are becoming more and more extensive. Through speech recognition, speech fragments can be recognized as text results, and the tasks corresponding to the text results can be executed subsequently.

[0003] In the related art, the pronunciation probability of each frame in the speech segment can be obtained through the acoustic model, and then the text information corresponding to the entire speech segment can be determined according to each pronunciation probability. However, the accuracy of current speech recognition is low. Summary of the invention

[0004] In order to solve the above technical problems, the present disclosure provides a method, device, equipment, storage medium and vehicle for determining a recognition result.

[0005] In a first aspect, the present disclosure provides a method for determining a recognition result, the method comprising:

[0006] Acquire initial text results generated by performing speech recognition on the speech segment, sort the initial text results, and determine text results to be processed in the initial text results;

[0007] Performing text adjustment on the text result to be processed to obtain an extended text result;

[0008] The candidate text results are sorted to determine the target text results among the candidate text results; wherein the candidate text results include the initial text results and the extended text results.

[0009] In a second aspect, the present disclosure provides a recognition result determination device, the device comprising:

[0010] A first determination module is used to obtain initial text results generated by performing speech recognition on the speech segment, sort the initial text results, and determine the text results to be processed in the initial text results;

[0011] An adjustment module, used for performing text adjustment on the text result to be processed to obtain an extended text result;

[0012] The second determination module is used to sort the candidate text results and determine the target text results in the candidate text results; wherein the candidate text results include the initial text results and the extended text results.

[0013] In a third aspect, an embodiment of the present disclosure further provides an electronic device, including:

[0014] processor;

[0015] A memory for storing executable instructions;

[0016] The processor is used to read executable instructions from the memory and execute the executable instructions to implement the recognition result determination method of the first aspect mentioned above.

[0017] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, wherein the storage medium stores a computer program. When the computer program is executed by a processor, the processor implements the recognition result determination method of the first aspect.

[0018] In a fifth aspect, an embodiment of the present disclosure further provides a vehicle, comprising at least one of the following: the above-mentioned recognition result determination device; the above-mentioned electronic device; and the above-mentioned computer-readable storage medium.

[0019] Compared with the prior art, the technical solution provided by the embodiment of the present disclosure has the following advantages: a recognition result determination method, device, equipment and storage medium of the embodiment of the present disclosure obtains the initial text result generated by voice recognition of the voice segment, and sorts the initial text result to determine the to-be-processed text result in the initial text result; performs text adjustment on the to-be-processed text result to obtain the extended text result; sorts the candidate text result to determine the target text result in the candidate text result; wherein the candidate text result includes the initial text result and the extended text result. Using the above technical solution, the voice segment is voice recognized to determine the initial text result, and the to-be-processed text result in the initial text result is expanded in the form of text adjustment to obtain the extended text result, and the target text result in the initial text result and the extended text result is determined. The text result determined by voice recognition is expanded in the form of text adjustment, which broadens the number of potential text results of the voice segment, increases the possibility of including the correct text result in the text result, thereby increasing the possibility of the target text result being the correct result, that is, the accuracy of voice recognition of the voice segment is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0022] Figure 1 A flowchart of a method for determining a recognition result provided by an embodiment of the present disclosure;

[0023] Figure 2 A flowchart of another identification result determination method provided by an embodiment of the present disclosure;

[0024] Figure 3 A schematic diagram of determining an extended text result provided by an embodiment of the present disclosure;

[0025] Figure 4 An entity graph for determining an example of an extended text result provided by an embodiment of the present disclosure;

[0026] Figure 5 A schematic diagram of another method for determining an extended text result provided by an embodiment of the present disclosure;

[0027] Figure 6 Another example of determining an extended text result entity graph provided by an embodiment of the present disclosure;

[0028] Figure 7 A schematic diagram of a model structure of a recognition result determination method provided for an example of the present disclosure;

[0029] Figure 8 A schematic diagram of the structure of a recognition result determination device provided in an embodiment of the present disclosure;

[0030] Fig. 9 A schematic diagram of the hardware circuit structure of a recognition result determination device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0032] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0033] With the development of speech recognition technology, the application scenarios of speech recognition are becoming more and more extensive. Through speech recognition, speech fragments can be recognized as text results, and the tasks corresponding to the text results can be executed subsequently.

[0034] In the related technology, the pronunciation probability of each frame in the speech segment can be obtained through the acoustic model, and then the maximum path is determined according to the maximum value of each pronunciation probability, and the text result corresponding to the entire speech segment is determined according to the maximum path. At the same time, the text results corresponding to other paths with larger values ​​can also be provided to the user as a reference for speech recognition.

[0035] However, even if the text results corresponding to multiple paths with larger values ​​are provided to the user, the similarity between the text result and the text result corresponding to the largest path is higher. Since the possibility of the speech fragment not being correctly recognized is greater due to homophones, similar sounds, and similar initials and finals, the current speech recognition accuracy is low.

[0036] In order to solve the above problems, the embodiments of the present disclosure provide a method, device, equipment, storage medium and vehicle for determining a recognition result.

[0037] Figure 1 The present invention provides a flowchart of a recognition result determination method provided by an embodiment of the present invention. The method can be executed by a recognition result determination device, wherein the device can be implemented by software and / or hardware, and can generally be integrated in an electronic device. The recognition result determination device and / or the electronic device can be configured in a first vehicle. Figure 1 As shown, the method includes:

[0038] Step 101, obtaining initial text results generated by performing speech recognition on a speech segment, sorting the initial text results, and determining text results to be processed in the initial text results.

[0039] The voice segment may be a voice instruction segment in which the user instructs the device to perform a task, and this embodiment does not limit the voice segment. For example, the voice segment may be a voice segment instructing multimedia playback, such as a voice segment instructing the playback of a singer's song; or the voice segment may be a voice segment instructing the control of a vehicle component, such as a voice segment instructing the opening of a car window.

[0040] The initial text result may be a text result obtained by recognizing a voice segment using an artificial intelligence method. This embodiment does not limit the number of the initial text results. For example, the number of the initial text results may be 3. The initial text result may include wake-up words, intent words, entity text, etc. The text result to be processed may be a text result determined by screening the evaluation value in the initial text result. For example, the text result to be processed may be the initial text result with the largest evaluation value. This embodiment does not limit the number of the text results to be processed. For example, the number of the text results to be processed may be one or more.

[0041] In this embodiment, the user can have a voice conversation with the recognition result determination device, which can convert the user's voice into a voice segment, and perform voice recognition on the voice segment through artificial intelligence to obtain multiple initial text results, and determine the initial text result with a higher evaluation value among the multiple initial text results as the text result to be processed.

[0042] In some embodiments of the present disclosure, obtaining an initial text result generated by performing speech recognition on a speech segment includes:

[0043] Input a speech segment into a preset acoustic model to obtain a model text result output by the acoustic model and an evaluation value corresponding to the model text result; sort the model text results according to the evaluation value, and determine the model text results in the first number as the initial text result.

[0044] The acoustic model may be a neural network model with speech recognition function, and this embodiment does not limit the neural network type, neural network structure, etc. of the acoustic model. The model text result may be all text results determined by the acoustic model performing speech recognition on the speech segment. The evaluation value may be the probability of characterizing the text result as a correct text result or a standard text result. The first number may be a pre-set number, and this embodiment does not limit the first number. For example, the first number may be 3.

[0045] In this embodiment, after the recognition result determination device generates a voice segment according to the user's voice, the voice segment is input into a pre-trained acoustic model. The acoustic model performs voice recognition on the voice segment and outputs a plurality of model text results, each of which has a corresponding evaluation value. The recognition result determination device sorts the model text results according to the evaluation values ​​from high to low, and determines the model text results in the first number as the initial text results.

[0046] In the above scheme, the first number of model text results are selected as the initial text results for subsequent processing, and the final text results are determined based on the initial text results with higher evaluation values, thereby improving the accuracy of speech recognition of speech fragments.

[0047] In some embodiments of the present disclosure, determining the text results to be processed in the initial text results includes: sorting the initial text results according to the evaluation values, and determining the first second number of initial text results as the text results to be processed.

[0048] The second number may be a preset number, and the second number may be a number not greater than the first number. This embodiment does not limit the first number. For example, if the first number is 3, the second number may be 1.

[0049] In this embodiment, after the recognition result determination device determines the initial text results, it sorts the initial text results in descending order of evaluation value, and determines the first second number of initial text results as the text results to be processed.

[0050] In the above scheme, the first second number of initial text results are selected as the text results to be processed for subsequent text adjustment, and the text results to be processed with higher evaluation values ​​are used as the basis for subsequent text adjustment, thereby improving the credibility of the text results obtained after text adjustment, and thus improving the accuracy of speech recognition of speech fragments.

[0051] Step 102: perform text adjustment on the text result to be processed to obtain an extended text result.

[0052] The extended text result may be a text result obtained by extending the text dimension of part or all of the text in the to-be-processed text result. That is, the extended text result may be a text result that is expanded and recalled based on the existing to-be-processed text result.

[0053] In an embodiment of the present disclosure, the recognition result determination device can extract text from the text result to be processed, and perform one or more text adjustments including text replacement, text deletion, and text addition on part of or all of the text to obtain an extended text result based on the extension of the text result to be processed.

[0054] Step 103, sorting the candidate text results to determine the target text results in the candidate text results; wherein the candidate text results include initial text results and extended text results.

[0055] The candidate text result may be a union of the initial text result and the extended text result.

[0056] In the disclosed embodiment, the initial text result is determined based on the acoustic model, and the to-be-processed text result with a higher evaluation value in the initial text result is expanded to obtain the extended text result. Furthermore, the recognition result determination device can score and sort the initial text result and the extended text result, and determine the text result with the highest score value as the target text result.

[0057] In the embodiments of the present disclosure, there are many methods for determining the target text result according to the candidate text result, which are not limited in this embodiment. Examples are as follows:

[0058] In an optional implementation, determining a target text result among candidate text results includes: inputting the candidate text results into a preset language model and decoder for scoring, obtaining a score value for each candidate text result, and determining the candidate text result with the largest score value as the target text result.

[0059] In another optional implementation, the candidate text results are sorted to determine the target text results in the candidate text results, including:

[0060] The candidate text results are scored according to their scoring information to obtain a score value of the candidate text results; wherein the scoring information includes at least one of historical playback information, priority information, and user feedback information; and the candidate text result corresponding to the maximum score value is used as the target text result.

[0061] The scoring information may be information used to score the text result. When the text result is a multimedia name, the historical playback information may be information recording the historical playback of the text result. This embodiment does not limit the historical playback information. For example, the historical playback information may include the historical playback duration, the historical playback times, etc. The priority information may be a preset priority level of the text result. The user feedback information may be information provided by the user regarding the text result. The user feedback information may include correct recognition, incorrect recognition, etc.

[0062] In this embodiment, the recognition result determination device can determine the scoring information of each candidate text result, and the scoring information may include multiple dimensions such as historical playback, priority, user feedback, etc. The scoring information of each dimension may have a corresponding scoring weight. Determine the number of scoring information of each candidate text result in each dimension, and then determine the score value of the candidate text result according to the score of the scoring information of each dimension and its corresponding scoring weight. Further, the candidate text result with the largest score value is determined as the final target text result of the voice segment.

[0063] In the above scheme, the final text result can be determined more comprehensively and accurately based on the scoring information.

[0064] The recognition result determination method provided by the embodiment of the present disclosure obtains the initial text result generated by voice recognition of the speech segment, and sorts the initial text result to determine the to-be-processed text result in the initial text result; performs text adjustment on the to-be-processed text result to obtain an extended text result; sorts the candidate text results to determine the target text result in the candidate text result; wherein the candidate text result includes the initial text result and the extended text result. In the embodiment of the present disclosure, voice recognition is performed on the speech segment to determine the initial text result, and the to-be-processed text result in the initial text result is expanded in the form of text adjustment to obtain an extended text result, and the target text result in the initial text result and the extended text result is determined. The text result determined by speech recognition is expanded in the form of text adjustment, which broadens the number of potential text results of the speech segment, increases the possibility of including the correct text result in the text result, thereby increasing the possibility of the target text result being the correct result, that is, the accuracy of speech recognition of the speech segment is improved.

[0065] Figure 2 A flowchart of another method for determining a recognition result provided by an embodiment of the present disclosure is shown in FIG. Figure 2 As shown, in some embodiments of the present disclosure, text adjustment is performed on the text result to be processed to obtain an extended text result, including:

[0066] Step 201, extracting the entity text to be processed whose text sequence is located after the start text in the text result to be processed; wherein the start text includes the wake-up word.

[0067] Among them, the text order can be the order in which the text is generated. The wake-up word can be a word that starts voice recognition. This embodiment does not limit the wake-up word. For example, the wake-up word can be "classmate". The start text is also called the start word. The start text can be part of the text in the text result to be processed that indicates the execution of the task. The start text can include a wake-up word and an indicative verb. The indicative verb is also called an intention word. For example, in the case of indicating music playing, the wake-up word can be classmates, the indicative verb can be play, and the start text can be classmates play. Entity text is also called subject or keyword. The entity text can be the text that records the object in the task in the text result to be processed. This embodiment does not limit the entity text. For example, the entity text can be a text that records the name of a song, etc. The entity text to be processed can be an entity text to be processed for text adjustment.

[0068] In this embodiment, the recognition result determination device performs word segmentation processing on the text result to be processed, intercepts the demonstrative verb and the text before the demonstrative verb in the text result to be processed as the start text, and takes the text after the demonstrative verb as the entity text to be processed.

[0069] Step 202: Adjust the entity text to be processed to obtain an extended entity text.

[0070] Among them, the extended entity text can be an entity text obtained by expanding the entity text to be processed in terms of text dimensions, that is, the extended entity text can be an entity text determined by expanding based on the existing entity text.

[0071] In the embodiments of the present disclosure, the recognition result determination device can perform one or more text adjustments on the entity text to be processed, such as text replacement, text deletion, and text addition, to obtain an extended entity text extended based on the entity text to be processed. Alternatively, a recall index library can be preset, and the corresponding relationship between entity texts is recorded in the form of key-value pairs in the recall index library. The entity text to be processed is queried in the recall index library to determine the extended entity text corresponding to the entity text to be processed.

[0072] In some embodiments of the present disclosure, adjusting the entity text to be processed to obtain an extended entity text includes the following steps:

[0073] Step a1: Divide the entity text to be processed into multiple entity words.

[0074] Among them, the entity word can be a word that records a single entity, and the entity word can be understood as the basic unit word that composes the entity text to be processed. For example, if the entity text to be processed records the singer name and the song name, the entity words can include the entity words that record the singer name and the entity words that record the song name.

[0075] In this embodiment, the recognition result determination device further performs word segmentation processing on the entity text to be processed, and divides the entity text to be processed into multiple basic entity words.

[0076] Step a2: Delete one or more characters in at least one entity word respectively to obtain multiple adjusted entity texts corresponding to the entity text to be processed.

[0077] Among them, the adjusted entity text can be an entity text determined by adjusting the entity text to be processed with the entity word as the basic unit.

[0078] In this embodiment, for each entity word, the recognition result determination device can delete one or more characters in the entity word to obtain one or more adjusted entity words corresponding to each entity word. Further, according to the order of the entity words in the entity text to be processed, different adjusted entity words corresponding to each entity word are respectively selected for splicing to obtain multiple adjusted entity texts.

[0079] Step a3: Determine the extended entity text based on the adjusted entity text.

[0080] In this embodiment, there are multiple methods for determining the extended entity text according to the adjusted entity text, which are not limited in this embodiment. Examples are described as follows:

[0081] In an optional implementation, determining the extended entity text based on the adjusted entity text includes: using the adjusted entity text as the extended entity text. That is, in this implementation, the adjusted entity text and the extended entity text are the same.

[0082] Figure 3 A schematic diagram of determining an extended text result provided by an embodiment of the present disclosure, such as Figure 3 As shown, after dividing the text results to be processed into the start text and the entity text to be processed, the entity words in the recognized text to be processed are deleted to obtain the corresponding adjusted entity text, and the adjusted entity text is used as the extended entity text corresponding to the entity text to be processed for subsequent splicing of the extended text results.

[0083] Figure 4 An example entity graph for determining an extended text result provided by an embodiment of the present disclosure, such as Figure 4 As shown in the figure, if the correct name of the singer is "Zhang San", the correct name of the song is "Ocean Song", if the text result to be processed is "Classmates play Zhang Sansi's Ocean Song", then the starting text is "Classmates play" and the entity text to be processed is "Zhang Sansi's Ocean Song". The entity words include "Zhang Sansi" and "Ocean Song". The entity word "Zhang Sansi" is deleted and "Zhang San" is obtained. Then the entity text is adjusted to "Zhang San's Ocean Song". In this example, by deleting the characters in the entity words, more candidates are provided, which increases the probability of the final text result being accurate.

[0084] In another optional implementation, determining the extended entity text based on the adjusted entity text includes:

[0085] The adjusted entity text is processed by phonetic encoding to obtain the text phonetic encoding; a target entity text that successfully matches the text phonetic encoding is determined among multiple candidate entity texts, and the target entity text is determined as the extended entity text.

[0086] Wherein, the text pinyin encoding can be a record of the entity text in the form of pinyin encoding. The candidate entity text can be a pre-set text correct entity text, and the relationship between the entity texts included in the same candidate entity text is also correct. Taking the singer name and song name as an example, the singer name and song name recorded in the candidate entity text are correct, and the corresponding relationship between the singer name and song name recorded in the same candidate entity text is correct. Each candidate entity text can have one or more corresponding candidate pinyin encodings, and the candidate pinyin encoding can be the correct standard pinyin encoding corresponding to the candidate entity text, and the candidate pinyin encoding can be an adjusted pinyin encoding determined by encoding adjustment based on the standard pinyin encoding, and the encoding adjustment can include front and back nasal adjustments and / or flat tongue and retroflex adjustments, etc. The target entity text can be a candidate entity text in which the candidate pinyin encoding successfully matches the text pinyin encoding.

[0087] Figure 5 A schematic diagram of another method for determining an extended text result provided by an embodiment of the present disclosure, such as Figure 5 As shown, after the text result to be processed is divided into the starting text and the entity text to be processed, the entity words in the recognized text to be processed are deleted to obtain the corresponding adjusted entity text, and the text pinyin encoding of the adjusted entity text is determined, and the text pinyin encoding is matched with the candidate pinyin encoding of the candidate entity text, and the candidate entity text corresponding to the successfully matched candidate pinyin encoding is determined as the target entity text, and the target entity text is used as the extended entity text corresponding to the entity text to be processed, and the subsequent splicing of the extended text results is performed.

[0088] Figure 6 Another example of determining an extended text result entity graph provided by the embodiment of the present disclosure is as follows: Figure 6 As shown, if the correct name of the singer is "Zhang San", the correct name of the song is "Ocean Song", if the result of the text to be processed is "Classmates play Zhang San's Haiyangge", the starting text is "Classmates play", and the entity text to be processed is "Zhang San's Haiyangge", delete "Yang", and adjust the entity text to "Zhang San's Haiyangge". The pinyin encoding of the text is "zhangsandehaige", and the target entity text that successfully matches the pinyin encoding of the text is "Zhang San's Ocean Song". And, the target entity text is used as the extended entity text. In this example, the text dimension of the entity text to be processed is first adjusted, and then the pinyin dimension is matched. The entity text to be processed can be corrected from both the text and pinyin dimensions, which expands the scope of entity text adjustment and increases the probability of matching the correct entity text.

[0089] Step 203, concatenate the extended entity text after the start text to obtain an extended text result.

[0090] In this embodiment, after the extended entity text corresponding to the entity text to be processed is determined, the extended entity text is concatenated after the start text of the text result to be processed to obtain an extended text result.

[0091] For example, Figure 4 As shown, the starting text is "Classmates play", the extended entity text is "Zhang San's Ocean Song", and the extended text result is "Classmates play Zhang San's Ocean Song".

[0092] Next, the recognition result determination method in the embodiment of the present disclosure is further explained through a specific example.

[0093] Figure 7 A schematic diagram of a model structure of a recognition result determination method provided by the example of the present disclosure, such as Figure 7 As shown in FIG. 1 , a recall module is added based on the acoustic model, language model and decoder to perform recall processing on the entity text to be processed.

[0094] Specifically, the model text results with the first number of evaluation values ​​output by the acoustic model are determined as the initial text results, and the model text results with the second number of evaluation values ​​are determined as the text results to be processed, and the entity text to be processed in the text results to be processed (i.e., the keywords or words in the text results to be processed) are extended and recalled. The entity text to be processed is determined in a pre-set recall index library to determine whether there is an extended entity text. If so, an extended text result is generated according to the extended entity text, and N extended text results and M initial text results are input into the language model and decoder for scoring and sorting, wherein N and M are positive integers.

[0095] Among them, the extended recall of the entity text to be processed includes: text matching recall and pinyin matching recall. In the text matching recall, the entity text to be processed is recalled based on text matching. In the pinyin matching recall, the entity text to be processed is recalled based on pinyin matching.

[0096] In the process of scoring the text results, the scoring can be based on user feedback, the number and frequency of occurrences of the entity text, the playback time of the content corresponding to the entity text, the number of on-demand views corresponding to the entity text, etc.

[0097] In the above scheme, the extended text results are generated based on text matching recall and pinyin matching recall, and homophone replacement, initials and finals adjustment and other processing are realized, thereby improving the recall rate and accuracy of the text results.

[0098] Figure 8 A schematic diagram of the structure of a recognition result determination device provided by an embodiment of the present disclosure is shown.

[0099] In some embodiments of the present disclosure, Figure 8The identification result determination device shown can be executed by an electronic device or a server. The electronic device can include but is not limited to a mobile terminal such as a vehicle-mounted terminal, and a fixed terminal such as a vehicle domain controller. The server can be a server cluster or a cloud server.

[0100] like Figure 8 As shown, the recognition result determination device 800 may include: a first determination module 801 , an adjustment module 802 , and a second determination module 803 .

[0101] The first determination module 801 is used to obtain initial text results generated by performing speech recognition on the speech segment, sort the initial text results, and determine the text results to be processed in the initial text results;

[0102] An adjustment module 802 is used to perform text adjustment on the text result to be processed to obtain an extended text result;

[0103] The second determination module 803 is used to sort the candidate text results and determine the target text results in the candidate text results; wherein the candidate text results include the initial text results and the extended text results.

[0104] Optionally, the first determining module 801 is configured to:

[0105] Inputting the speech segment into a preset acoustic model to obtain a model text result output by the acoustic model and an evaluation value corresponding to the model text result;

[0106] The model text results are sorted according to the evaluation values, and the model text results that are in the first first number are determined as the initial text results.

[0107] Optionally, the first determining module 801 is configured to:

[0108] The initial text results are sorted according to the evaluation values, and the first second number of initial text results are determined as the text results to be processed.

[0109] Optionally, the adjustment module 802 includes:

[0110] An extraction submodule, used to extract the entity text to be processed whose text sequence is located after the start text in the text result to be processed; wherein the start text includes a wake-up word;

[0111] An adjustment submodule, used for performing text adjustment on the entity text to be processed to obtain an extended entity text;

[0112] The splicing submodule is used to splice the extended entity text after the start text to obtain the extended text result.

[0113] Optionally, the adjustment submodule includes:

[0114] A division unit, used for dividing the entity text to be processed into a plurality of entity words;

[0115] A deleting unit, used to delete one or more characters in at least one of the entity words, respectively, to obtain a plurality of adjusted entity texts corresponding to the entity text to be processed;

[0116] A determination unit is used to determine the extended entity text based on the adjusted entity text.

[0117] Optionally, the determining unit is used to:

[0118] Performing pinyin encoding processing on the adjustment entity text to obtain a text pinyin encoding;

[0119] A target entity text that successfully matches the text phonetic encoding is determined among multiple candidate entity texts, and the target entity text is determined as the extended entity text.

[0120] Optionally, the second determining module 803 is configured to:

[0121] Scoring the candidate text result according to the scoring information of the candidate text result to obtain a score value of the candidate text result; wherein the scoring information includes: at least one of historical playback information, priority information, and user feedback information;

[0122] The candidate text result corresponding to the maximum score value is used as the target text result.

[0123] The recognition result determination device provided in the embodiments of the present disclosure can execute the recognition result determination method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0124] Fig. 9 A schematic diagram of the hardware circuit structure of a recognition result determination device provided by an embodiment of the present disclosure is shown.

[0125] like Fig. 9 As shown, the recognition result determination device 900 may include a controller 901 and a memory 902 storing computer program instructions.

[0126] Specifically, the controller 901 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0127] The memory 902 may include a large capacity memory for information or instructions. By way of example and not limitation, the memory 902 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a tape, or a universal serial bus (USB) drive or a combination of two or more of these. Where appropriate, the memory 902 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 902 may be inside or outside the integrated gateway device. In a particular embodiment, the memory 902 is a non-volatile solid-state memory. In a particular embodiment, the memory 902 includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (Electrically Erasable Programmable ROM, EEPROM), an electrically rewritable ROM (EAROM) or a flash memory, or a combination of two or more of these.

[0128] The controller 901 reads and executes the computer program instructions stored in the memory 902 to perform the steps of the recognition result determination method provided in the embodiment of the present disclosure.

[0129] In one example, the recognition result determination device 900 may further include a transceiver 903 and a bus 904. Fig. 9 As shown, the controller 901, the memory 902 and the transceiver 903 are connected via a bus 904 and communicate with each other.

[0130] The bus 904 includes hardware, software or both. For example, but not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a Memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, the bus 904 may include one or more buses. Although embodiments of the present application describe and illustrate a particular bus, the present application contemplates any suitable bus or interconnect.

[0131] The following is an embodiment of a computer-readable storage medium provided in an embodiment of the present disclosure. The computer-readable storage medium and the recognition result determination method of the above-mentioned embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the computer-readable storage medium, reference can be made to the embodiment of the above-mentioned recognition result determination method.

[0132] This embodiment provides a storage medium including computer executable instructions, which are used to perform a recognition result determination method when executed by a computer processor.

[0133] Of course, the storage medium containing computer executable instructions provided by the embodiment of the present disclosure is not limited to the above method operations, and the computer executable instructions can also execute related operations in the recognition result determination method provided by any embodiment of the present disclosure.

[0134] Through the above description of the implementation method, the technical personnel in the relevant field can clearly understand that the present disclosure can be implemented by means of software and necessary general hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions to enable a computer cloud platform (which can be a personal computer, a server, or a network cloud platform, etc.) to execute the recognition result determination method provided by each embodiment of the present disclosure.

[0135] The embodiment of the present disclosure also provides a vehicle, comprising at least one of the following: the above-mentioned recognition result determination device; the above-mentioned electronic device; and the above-mentioned computer-readable storage medium.

[0136] Note that the above are only preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art will understand that the present disclosure is not limited to the specific embodiments herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present disclosure. Therefore, although the present disclosure is described in more detail through the above embodiments, the present disclosure is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present disclosure, and the scope of the present disclosure is determined by the scope of the appended claims.

Claims

1. A method for determining a recognition result, It is characterized in that include: Acquire initial text results generated by performing speech recognition on the speech segment, sort the initial text results, and determine text results to be processed in the initial text results; Performing text adjustment on the text result to be processed to obtain an extended text result; The candidate text results are sorted to determine the target text results among the candidate text results; wherein the candidate text results include the initial text results and the extended text results.

2. The method according to claim 1, It is characterized in that The step of obtaining an initial text result generated by performing speech recognition on the speech segment includes: Inputting the speech segment into a preset acoustic model to obtain a model text result output by the acoustic model and an evaluation value corresponding to the model text result; The model text results are sorted according to the evaluation values, and the model text results that are in the first first number are determined as the initial text results.

3. The method according to claim 1, It is characterized in that The determining of the to-be-processed text results in the initial text results includes: The initial text results are sorted according to the evaluation values, and the first second number of initial text results are determined as the text results to be processed.

4. The method according to claim 1, It is characterized in that The step of performing text adjustment on the text result to be processed to obtain an extended text result includes: Extracting the entity text to be processed whose text sequence is located after the start text in the text result to be processed; wherein the start text includes a wake-up word; Performing text adjustment on the entity text to be processed to obtain extended entity text; The extended entity text is concatenated after the start text to obtain the extended text result.

5. The method according to claim 4, It is characterized in that The step of adjusting the entity text to be processed to obtain the extended entity text includes: Dividing the entity text to be processed into multiple entity words; Deleting one or more characters in at least one of the entity words respectively to obtain a plurality of adjusted entity texts corresponding to the entity text to be processed; The extended entity text is determined based on the adjusted entity text.

6. The method according to claim 5, It is characterized in that The determining the extended entity text based on the adjusted entity text includes: Performing pinyin encoding processing on the adjustment entity text to obtain a text pinyin encoding; A target entity text that successfully matches the text phonetic encoding is determined among multiple candidate entity texts, and the target entity text is determined as the extended entity text.

7. The method according to claim 1, It is characterized in that The step of sorting the candidate text results to determine the target text results in the candidate text results includes: Scoring the candidate text result according to the scoring information of the candidate text result to obtain a score value of the candidate text result; wherein the scoring information includes: at least one of historical playback information, priority information, and user feedback information; The candidate text result corresponding to the maximum score value is used as the target text result.

8. A recognition result determination device, It is characterized in that include: A first determination module is used to obtain initial text results generated by performing speech recognition on the speech segment, sort the initial text results, and determine the text results to be processed in the initial text results; An adjustment module, used for performing text adjustment on the text result to be processed to obtain an extended text result; The second determination module is used to sort the candidate text results and determine the target text results in the candidate text results; wherein the candidate text results include the initial text results and the extended text results.

9. An electronic device, It is characterized in that The electronic device comprises: Processor and memory; The processor is used to execute the steps of the method according to any one of claims 1 to 7 by calling the program or instruction stored in the memory.

10. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a program or instruction, and the program or instruction enables a computer to execute the steps of the method according to any one of claims 1 to 7.

11. A vehicle, It is characterized in that Include at least one of the following: The recognition result determination device as described in claim 8 above; The electronic device as claimed in claim 9; The computer readable storage medium of claim 10.