Voice Recognition Modification Method, Device, Electronic Device and Storage Medium
Through user correction and self-learning optimization methods, the problem of recognition errors in speech recognition technology when dealing with similar or unclear speech is solved, improving user experience and recognition accuracy.
Patent Information
- Application Number
- CN202310459939.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-25
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-04-25
AI Technical Summary
When existing speech recognition technology processes speech information with similar or unclear pronunciations, it is difficult to effectively distinguish and improve, resulting in a reduction in recognition errors and user experience.
By collecting the user's voice commands, voice recognition results are generated, and whether the user's expected results are met. If it is not satisfied, the user's correction operation is received and the recognition results are corrected to satisfy the user's expectations. At the same time, establish a speech recognition optimization model and mapping relationship, self-learning optimization process, and improve the accuracy of automatic recognition.
It effectively solves the problem of recognition of similar pronunciations or unclear voice information, improves the user experience, and ensures that the recognition results meet user expectations.
Smart Images

Figure CN116343792B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speech recognition, and particularly relates to a speech recognition modification method, device, electronic device and storage medium. Background Art
[0002] ASR (Automatic Speech Recognition) is a technology that converts speech into text. For terminal devices, it generally includes cloud recognition and local recognition. Among them, cloud recognition has a large computing power, a large background training data set, better recognition effects, and wide corpus coverage, but it requires an Internet connection; local recognition terminal devices are restricted by storage, power consumption, and computing power, and the corpus is often fixed, and the recognition effect for extended corpus is relatively poor. Moreover, there are many homophones in Chinese, which may lead to the recognition content not being the result that the user wants.
[0003] Since there are many generalized extended expressions for the same instruction, it is impossible to completely cover and accurately recognize the instruction during the early-stage recognition optimization. First, users are often prone to trigger expressions outside the established recognition optimization, which may lead to the recognized instruction results not meeting the user's expectations; second, when recognizing the voice commands of users with accents, it is difficult to trigger effective commands, resulting in incorrect recognition results all the time. Therefore, both of the above situations will lead to semantic recognition errors or unpredictability.
[0004] In related technologies, when a user's voice command is recognized, most often use methods of word or speech correction to automatically correct and change the collected user voice command, so as to achieve the result expected by the user.
[0005] However, when correcting errors through words, there are often limited choices for the words to be corrected, so the result that the user wants cannot be given, and for texts with similar pronunciations, effective discrimination and improvement cannot be carried out, and there are often situations where multiple recognitions are incorrect; while for speech correction, the pronunciation requirements for secondary recognition are relatively high. When encountering users with unclear pronunciation or accents, there are often problems with incorrect recognition in secondary correction, so the misrecognized corpus cannot be optimized, resulting in the same problem being triggered when subsequent users trigger the same or similar corpus, thus reducing the user's driving experience, which urgently needs to be solved. Summary of the Invention
[0006] This application provides a speech recognition modification method, device, electronic device and storage medium to solve the problems that effective discrimination and improvement cannot be carried out for speech information with similar or unclear pronunciations, thereby reducing the user's experience.
[0007] An embodiment of the first aspect of the present application provides a method for modifying speech recognition, including the following steps: collecting a first speech command of a user; generating a first speech recognition result according to the first speech command, and determining whether the first speech recognition result meets the first expected recognition result of the user; and if the first speech recognition result does not meet the first expected recognition result, receiving a first correction operation generated by the user based on the first speech recognition result, and correcting the first speech recognition result according to the first correction operation, so that the corrected first speech recognition result meets the first expected recognition result.
[0008] According to the above technical means, by allowing the user to modify the recognized speech recognition result by themselves, the result can meet the user's expected requirements, thereby improving the user's experience.
[0009] Optionally, after correcting the first speech recognition result according to the first correction operation, it further includes: establishing a speech recognition optimization model based on the first speech command and the first expected recognition result, and / or establishing a mapping relationship between the first speech recognition result and the first expected recognition result; performing speech recognition according to the speech recognition optimization model and / or the mapping relationship.
[0010] According to the above technical means, a mapping relationship between the speech command and the correction result is established through the self-learning function, so that when the user triggers the same speech information, the expected result of the user can be automatically recognized, thereby improving the user's experience.
[0011] Optionally, the above method for modifying speech recognition further includes: collecting a second speech command of the user, and determining whether the second speech command is the same as the first speech command; if the second speech command is the same as the first speech command, recognizing the second speech command based on the speech recognition optimization model to obtain a second speech recognition result, and determining whether the second speech recognition result meets the first expected recognition result of the user; if the second speech recognition result does not meet the first expected recognition result of the user, obtaining the correction result corresponding to the first speech recognition result from the mapping relationship, and correcting the second speech recognition result according to the correction result corresponding to the first speech recognition result.
[0012] According to the above technical means, by comparing the collected second speech command with the first speech command, when the result does not meet the user's expectations, it can be automatically corrected according to the mapping relationship, so as to meet the user's usage requirements.
[0013] Optionally, after determining whether the second voice recognition result meets the user's first expected recognition result, the method further includes: if the second voice recognition result meets the user's first expected recognition result, generating a first control instruction for the vehicle according to the second recognition result; and controlling the vehicle to perform corresponding control actions according to the first control instruction.
[0014] According to the above technical means, by matching the optimized correction result and performing the interaction actions required by the user according to the correction result, the driving experience of the user is improved.
[0015] Optionally, after determining whether the second voice command is the same as the first voice command, the method further includes: if the second voice command is different from the first voice command, generating a third voice recognition result according to the second voice command, and determining whether the third voice recognition result meets the user's second expected recognition result; if the third voice recognition result does not meet the second expected recognition result, receiving a second correction operation generated by the user based on the third voice recognition result, and correcting the third voice recognition result according to the second correction operation, so that the corrected third voice recognition result meets the second expected recognition result.
[0016] According to the above technical means, when it is recognized that the voice command is inconsistent with the previous command, the new voice command can be recognized and optimized based on the voice recognition optimization model, so as to meet the user's usage requirements.
[0017] Optionally, after correcting the first voice recognition result according to the first correction operation, the method further includes: generating a second control instruction for the vehicle according to the corrected first voice recognition result; and controlling the vehicle to perform corresponding actions according to the second control instruction.
[0018] According to the above technical means, the vehicle is controlled to perform corresponding actions through the corrected voice command, thereby enhancing the user's usage experience.
[0019] Optionally, before collecting the user's first voice command, the method further includes: determining whether the voice recognition operation of the vehicle is triggered; if the voice recognition operation of the vehicle is triggered, collecting the user's first voice command.
[0020] According to the above technical means, by triggering the voice recognition operation of the vehicle to collect the user's voice command, the collection accuracy of the voice command is improved.
[0021] A second aspect embodiment of the present application provides a voice recognition modification device, including: a collection module for collecting a first voice command of a user; a judgment module for generating a first voice recognition result according to the first voice command and judging whether the first voice recognition result meets the user's first expected recognition result; and a correction module for, if the first voice recognition result does not meet the first expected recognition result, receiving a first correction operation generated by the user based on the first voice recognition result and correcting the first voice recognition result according to the first correction operation, so that the corrected first voice recognition result meets the first expected recognition result.
[0022] Optionally, after correcting the first voice recognition result according to the first correction operation, the correction module further includes: a construction unit for establishing a voice recognition optimization model according to the first voice command and the first expected recognition result, and / or establishing a mapping relationship between the first voice recognition result and the first expected recognition result; a first recognition unit for performing voice recognition according to the voice recognition optimization model and / or the mapping relationship.
[0023] Optionally, the above voice recognition modification device further includes: a first judgment unit for collecting a second voice command of the user and judging whether the second voice command is the same as the first voice command; a second recognition unit for, if the second voice command is the same as the first voice command, recognizing the second voice command based on the voice recognition optimization model to obtain a second voice recognition result and judging whether the second voice recognition result meets the user's first expected recognition result; a correction unit for, if the second voice recognition result does not meet the user's first expected recognition result, obtaining a correction result corresponding to the first voice recognition result from the mapping relationship and correcting the second voice recognition result according to the correction result corresponding to the first voice recognition result.
[0024] Optionally, after judging whether the second voice recognition result meets the user's first expected recognition result, the second recognition unit is further configured to: if the second voice recognition result meets the user's first expected recognition result, generate a first control command for the vehicle according to the second recognition result; control the vehicle to perform corresponding control actions according to the first control command.
[0025] Optionally, after determining whether the second voice command is the same as the first voice command, the determining unit is further configured to: if the second voice command is different from the first voice command, generate a third voice recognition result according to the second voice command, and determine whether the third voice recognition result meets the user's second expected recognition result; if the third voice recognition result does not meet the second expected recognition result, receive a second correction operation generated by the user based on the third voice recognition result, and correct the third voice recognition result according to the second correction operation, so that the corrected third voice recognition result meets the second expected recognition result.
[0026] Optionally, after correcting the first voice recognition result according to the first correction operation, the correction module further includes: a generating unit, configured to generate a second control command for the vehicle according to the corrected first voice recognition result; a control unit, configured to control the vehicle to perform corresponding actions according to the second control command.
[0027] Optionally, before collecting the user's first voice command, the collecting module further includes: a second determining unit, configured to determine whether the voice recognition operation of the vehicle is triggered; a collecting unit, configured to collect the user's first voice command if the voice recognition operation of the vehicle is triggered.
[0028] An embodiment of the third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the voice recognition modification method as described in the above embodiment.
[0029] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the voice recognition modification method as described in the above embodiment.
[0030] In the embodiment of the present application, by collecting the user's first voice command, generating a first voice recognition result, and determining whether the first voice recognition result meets the user's first expected recognition result, when the first expected recognition result is not met, receiving the first correction operation generated by the user based on the first voice recognition result, and correcting the first voice recognition result according to the first correction operation, so that the corrected first voice recognition result meets the first expected recognition result. Thereby, problems such as the inability to effectively distinguish and improve voice information with similar pronunciations or unclear pronunciations, thereby reducing the user experience, are solved.
[0031] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Brief Description of the Drawings
[0032] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, where:
[0033] Figure 1 It is a flowchart of a method for modifying speech recognition according to an embodiment of the present application;
[0034] Figure 2 It is a schematic diagram of the process for modifying the original recognition result according to an embodiment of the present application;
[0035] Figure 3 It is a schematic diagram of the automatic correction process for secondary trigger recognition according to an embodiment of the present application;
[0036] Figure 4 It is a block schematic diagram of a speech recognition modification device according to an embodiment of the present application;
[0037] Figure 5 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present application.
[0038] Description of the reference numerals: 10 - speech recognition modification device; 100 - acquisition module, 200 - judgment module, 300 - correction module Detailed Description of the Embodiments
[0039] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.
[0040] The voice recognition modification method, device, electronic device, and storage medium according to the embodiments of the present application will be described below with reference to the accompanying drawings. In view of the problem in the above-mentioned background technology that it is impossible to effectively distinguish and improve voice information with similar pronunciations or unclear pronunciations, thereby reducing the user experience, the present application provides a voice recognition modification method. In this method, by collecting the user's first voice command, a first voice recognition result is generated, and it is determined whether the first voice recognition result meets the user's first expected recognition result. When the first expected recognition result is not met, the user's first correction operation generated based on the first voice recognition result is received, and the first voice recognition result is corrected according to the first correction operation, so that the corrected first voice recognition result meets the first expected recognition result. Thereby, the problems such as the inability to effectively distinguish and improve voice information with similar pronunciations or unclear pronunciations, thereby reducing the user experience, are solved. When it is recognized that the user's voice command does not meet the result expectation, the recognized voice command is corrected, and a mapping relationship between the voice command and the correction result is established through the self-learning function, so that when the user triggers the same voice information, the user's expected result can be automatically recognized, thereby improving the user experience.
[0041] Specifically, Figure 1 is a flowchart of a voice recognition modification method provided by an embodiment of the present application.
[0042] As Figure 1 shown, the voice recognition modification method includes the following steps:
[0043] In step S101, the user's first voice command is collected.
[0044] Further, in an embodiment of the present application, before collecting the user's first voice command, it further includes: determining whether the voice recognition operation of the vehicle is triggered; if the voice recognition operation of the vehicle is triggered, the user's first voice command is collected.
[0045] Specifically, as Figure 2 shown, during the driving process of the user, if the user wants to implement the interaction function of the vehicle, the user first needs to trigger the voice recognition operation of the vehicle and send a voice command to the vehicle. At this time, the vehicle collects the first voice command issued by the user.
[0046] In step S102, a first voice recognition result is generated according to the first voice command, and it is determined whether the first voice recognition result meets the user's first expected recognition result.
[0047] Specifically, in the embodiments of the present application, after the vehicle collects the first voice command issued by the user, it recognizes the voice information issued by the user according to the content of the first voice command, and displays the recognized voice information in the vehicle typewriter to generate a first voice recognition result, and determines whether the recognition result meets the recognition result expected by the user, that is, the first expected recognition result.
[0048] In step S103, if the first voice recognition result does not meet the first expected recognition result, then the vehicle receives the first correction operation generated by the user based on the first voice recognition result, and corrects the first voice recognition result according to the first correction operation, so that the corrected first voice recognition result meets the first expected recognition result.
[0049] Further, in an embodiment of the present application, after correcting the first voice recognition result according to the first correction operation, the method further includes: generating a second control command for the vehicle according to the corrected first voice recognition result; controlling the vehicle to perform corresponding actions according to the second control command.
[0050] Specifically, in the embodiments of the present application, if the first voice recognition result displayed on the vehicle typewriter does not meet the first expected recognition result of the user, at this time, the user can double-click the vehicle typewriter to make the first voice recognition result in the typewriter in a selected state, and after the correction button is displayed on the typewriter, the user clicks the correction button to edit and correct the first voice recognition result through the system keyboard, so as to correct the first voice recognition result to the first expected recognition result required by the user, and click the confirmation button to complete the correction operation after the correction is completed, forming the first expected recognition result. At the same time, after the user completes the correction, the vehicle receives the corrected first voice recognition result of the user, so that the corrected first voice recognition result meets the first expected recognition result of the user, and generates the control command required by the user according to the corrected first voice recognition result, so as to control the vehicle to perform corresponding interaction actions through the control command.
[0051] Further, in an embodiment of the present application, after correcting the first voice recognition result according to the first correction operation, the method further includes: establishing a voice recognition optimization model based on the first voice command and the first expected recognition result, and / or establishing a mapping relationship between the first voice recognition result and the first expected recognition result; performing voice recognition according to the voice recognition optimization model and / or the mapping relationship.
[0052] Specifically, after the first voice recognition result is corrected according to the first correction operation in the embodiments of the present application, the corrected first voice recognition result can be recorded and uploaded to the recognition module as a hot word for recognition optimization. The optimization method mainly includes two aspects. First, a mapping relationship between the first voice recognition result and the first expected recognition result is established and stored by mapping. Second, a voice recognition optimization model is established according to the first voice command and the first expected recognition result, so as to automatically match the corresponding first expected recognition result when the same voice command of the user is recognized, thereby controlling the vehicle to perform corresponding interaction actions.
[0053] For example, in the embodiments of the present application, if the user wants to turn on the electronic parking function by voice, first, the user sends a voice message of "da kai dian zi zhu che" to the vehicle. If the recognition result of the vehicle typewriter is "turn on electronic registration", since this recognition result does not meet the user's first expected recognition result, at this time, the user can double-click the typewriter to edit and correct this recognition result, correct the originally recognized "turn on electronic registration" to "turn on electronic parking", and then click the confirmation button to complete the correction operation. At this time, after the vehicle voice module queries the first expected recognition result, it issues a "turn on electronic parking" command and uses "turn on electronic parking" as the recognition optimization statement for modeling optimization. At the same time, a mapping relationship between "turn on electronic registration" and "turn on electronic parking" is established, so that when the user triggers the same voice message again, the vehicle will automatically recognize the user's first expected recognition result, thereby controlling the vehicle to perform corresponding interaction actions.
[0054] Further, in an embodiment of the present application, the above voice recognition modification method further includes: collecting a second voice command of the user and determining whether the second voice command is the same as the first voice command; if the second voice command is the same as the first voice command, then recognizing the second voice command based on the voice recognition optimization model to obtain a second voice recognition result, and determining whether the second voice recognition result meets the user's first expected recognition result; if the second voice recognition result does not meet the user's first expected recognition result, then obtaining the correction result corresponding to the first voice recognition result from the mapping relationship and correcting the second voice recognition result according to the correction result corresponding to the first voice recognition result.
[0055] Specifically, as Figure 3As shown, when the user needs to use voice control to make the vehicle perform relevant interaction actions again during driving, after the user issues a voice command to the vehicle, the vehicle collects the user's second voice command and determines whether the currently collected second voice command is the same as the previously collected first voice command. If the second voice command is the same as the first voice command, the second voice command is recognized based on the voice recognition optimization model to obtain a second voice recognition result, and it is determined whether the second voice recognition result meets the user's first expected recognition result. When the second voice recognition result does not meet the user's first expected recognition result, at this time, the correction result corresponding to the first voice recognition result is obtained from the established mapping relationship, and the second voice recognition result is corrected according to the correction result corresponding to the first voice recognition result, so that the second voice recognition result meets the user's first expected recognition result. The corrected second voice recognition result is displayed on the vehicle typewriter, and the voice command information of the user is obtained through the result displayed on the typewriter, so as to issue an interaction command according to the voice command information to perform corresponding interaction actions.
[0056] For example, in the embodiment of the present application, when the user triggers the same command or the same pronunciation again, such as issuing the voice information "da kai dian zi zhu che", since the correction result has been optimized for recognition, the vehicle recognition module will first match the optimized first voice recognition result. If the recognition result is still incorrect after recognition optimization and is recognized as the wrong voice information "open electronic registration", at this time, the corrected mapping text can be found in the mapping storage through the first voice recognition result, that is, the wrong information "open electronic registration" is found from the mapping relationship to obtain the correction result "open electronic parking brake", and "open electronic parking brake" is used to update the recognition result display of the typewriter, so as to issue a command to control the vehicle to perform the electronic parking brake action.
[0057] Further, in an embodiment of the present application, after determining whether the second voice recognition result meets the user's first expected recognition result, it further includes: if the second voice recognition result meets the user's first expected recognition result, a first control command for the vehicle is generated according to the second recognition result; the vehicle is controlled to perform corresponding control actions according to the first control command.
[0058] Specifically, in the embodiment of the present application, when the second voice command is recognized based on the voice recognition optimization model to obtain a second voice recognition result, and it is determined that the second voice recognition result meets the user's first expected recognition result, that is, when the "dakai dian zi zhu che" issued by the user is displayed as "open electronic parking brake", a control command for the vehicle to open the electronic parking brake is generated according to this recognition result, and the vehicle is controlled to perform corresponding parking actions according to the control command.
[0059] Further, in an embodiment of the present application, after determining whether the second voice command is the same as the first voice command, the following steps are further included: If the second voice command is different from the first voice command, generate a third voice recognition result according to the second voice command, and determine whether the third voice recognition result meets the user's second expected recognition result; If the third voice recognition result does not meet the second expected recognition result, receive the second correction operation generated by the user based on the third voice recognition result, and correct the third voice recognition result according to the second correction operation, so that the corrected third voice recognition result meets the second expected recognition result.
[0060] Specifically, in the embodiment of the present application, after the vehicle collects the second voice command issued by the user, if it is recognized that the second voice command is different from the first voice command, generate a third voice recognition result according to the second voice command and display it in the vehicle typewriter. At the same time, determine whether the third voice recognition result meets the user's second expected recognition result. If the third voice recognition result does not meet the second expected recognition result issued by the user, at this time, the user double-clicks the vehicle typewriter to edit and correct the third voice recognition result to correct the third voice recognition result to the second expected recognition result required by the user. After the user finishes the correction, receive the corrected third voice recognition result of the user, so that the corrected third voice recognition result meets the user's second expected recognition result, and record, optimize and store the user's third voice recognition result, and generate the control command required by the user according to the corrected third voice recognition result, so as to control the vehicle to perform corresponding interaction actions through this control command. It should be noted that the optimization method here is the same as the above optimization method and will not be specifically described here.
[0061] According to the voice recognition modification method of the embodiment of the present application, by collecting the user's first voice command, generating a first voice recognition result, and determining whether the first voice recognition result meets the user's first expected recognition result, when it does not meet the first expected recognition result, receive the first correction operation generated by the user based on the first voice recognition result, and correct the first voice recognition result according to the first correction operation, so that the corrected first voice recognition result meets the first expected recognition result. Thus, the problems that it is impossible to effectively distinguish and improve the voice information with similar pronunciation or unclear pronunciation, thereby reducing the user experience, etc. are solved. When it is recognized that the user's voice command does not meet the result expectation, the recognized voice command is corrected, and the mapping relationship between the voice command and the correction result is established through the self-learning function, so that when the user triggers the same voice information, the expected result of the user can be automatically recognized, thereby enhancing the user experience.
[0062] Next, a voice recognition modification device according to an embodiment of the present application will be described with reference to the accompanying drawings.
[0063] Figure 5It is a block diagram of the voice recognition modification device according to an embodiment of the present application.
[0064] As Figure 5 shown, the voice recognition modification device 10 includes: an acquisition module 100, a judgment module 200, and a correction module 300.
[0065] Among them, the acquisition module 100 is used to acquire the first voice command of the user;
[0066] The judgment module 200 is used to generate a first voice recognition result according to the first voice command, and judge whether the first voice recognition result meets the user's first expected recognition result; and
[0067] The correction module 300 is used to, if the first voice recognition result does not meet the first expected recognition result, receive the first correction operation generated by the user based on the first voice recognition result, and correct the first voice recognition result according to the first correction operation, so that the corrected first voice recognition result meets the first expected recognition result.
[0068] Further, in an embodiment of the present application, after correcting the first voice recognition result according to the first correction operation, the correction module 300 further includes: a construction unit and a first recognition unit.
[0069] Among them, the construction unit is used to establish a voice recognition optimization model according to the first voice command and the first expected recognition result, and / or establish a mapping relationship between the first voice recognition result and the first expected recognition result;
[0070] The first recognition unit is used to perform voice recognition according to the voice recognition optimization model and / or the mapping relationship.
[0071] Further, in an embodiment of the present application, the above-mentioned voice recognition modification device 10 further includes: a first judgment unit, a second recognition unit, and a correction unit.
[0072] Among them, the first judgment unit is used to acquire the second voice command of the user and judge whether the second voice command is the same as the first voice command;
[0073] The second recognition unit is used to, if the second voice command is the same as the first voice command, recognize the second voice command based on the voice recognition optimization model to obtain a second voice recognition result, and judge whether the second voice recognition result meets the user's first expected recognition result;
[0074] The correction unit is used to, if the second voice recognition result does not meet the user's first expected recognition result, obtain the correction result corresponding to the first voice recognition result from the mapping relationship, and correct the second voice recognition result according to the correction result corresponding to the first voice recognition result.
[0075] Further, in an embodiment of the present application, after determining whether the second voice recognition result meets the user's first expected recognition result, the second recognition unit is further configured to:
[0076] If the second voice recognition result meets the user's first expected recognition result, generate a first control instruction for the vehicle according to the second recognition result;
[0077] Control the vehicle to perform corresponding control actions according to the first control instruction.
[0078] Further, in an embodiment of the present application, after determining whether the second voice command is the same as the first voice command, the determination unit is further configured to:
[0079] If the second voice command is different from the first voice command, generate a third voice recognition result according to the second voice command, and determine whether the third voice recognition result meets the user's second expected recognition result;
[0080] If the third voice recognition result does not meet the second expected recognition result, receive the second correction operation generated by the user based on the third voice recognition result, and correct the third voice recognition result according to the second correction operation so that the corrected third voice recognition result meets the second expected recognition result.
[0081] Further, in an embodiment of the present application, after correcting the first voice recognition result according to the first correction operation, the correction module 300 further includes: a generation unit and a control unit.
[0082] The generation unit is configured to generate a second control instruction for the vehicle according to the corrected first voice recognition result;
[0083] The control unit is configured to control the vehicle to perform corresponding actions according to the second control instruction.
[0084] Further, in an embodiment of the present application, before collecting the user's first voice command, the collection module 100 further includes: a second determination unit and a collection unit.
[0085] The second determination unit is configured to determine whether the voice recognition operation of the vehicle is triggered;
[0086] The collection unit is configured to collect the user's first voice command if the voice recognition operation of the vehicle is triggered.
[0087] The voice recognition modification device according to the embodiment of the present application collects the user's first voice command, generates a first voice recognition result, and determines whether the first voice recognition result meets the user's first expected recognition result. When it does not meet the first expected recognition result, it receives the first correction operation generated by the user based on the first voice recognition result, and corrects the first voice recognition result according to the first correction operation, so that the corrected first voice recognition result meets the first expected recognition result. Thus, it solves the problems that it is impossible to effectively distinguish and improve the voice information with similar pronunciations or unclear pronunciations, thereby reducing the user experience. When it is recognized that the user's voice command does not meet the result expectation, the recognized voice command is corrected, and the mapping relationship between the voice command and the correction result is established through the self-learning function, so that when the user triggers the same voice information, the user's expected result can be automatically recognized, thereby improving the user experience.
[0088] Figure 5 The structural schematic diagram of the electronic device provided by the embodiment of the present application. The electronic device may include:
[0089] A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.
[0090] When the processor 502 executes the program, it implements the voice recognition modification method provided in the above embodiment.
[0091] Further, the electronic device further includes:
[0092] A communication interface 503 for communication between the memory 501 and the processor 502.
[0093] The memory 501 is used to store a computer program executable on the processor 502.
[0094] The memory 501 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0095] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 can be interconnected via a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 only a thick line is used in Figure 5 , but it does not mean that there is only one bus or one type of bus.
[0096] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other via an internal interface.
[0097] The processor 502 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.
[0098] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-described method for modifying speech recognition is implemented.
[0099] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0100] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0101] Any process or method description shown in a flowchart or described otherwise herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0102] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, etc.
[0103] Those of ordinary skill in the art of the present technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and when executed, includes one or a combination of the steps of the method embodiments.
[0104] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for modifying speech recognition, characterized in that, Including the following steps: Collect the user's first voice command; Generate a first speech recognition result according to the first voice command, and determine whether the first speech recognition result meets the user's first expected recognition result; And If the first speech recognition result does not meet the first expected recognition result, receive the first correction operation generated by the user based on the first speech recognition result, and correct the first speech recognition result according to the first correction operation, so that the corrected first speech recognition result meets the first expected recognition result; After correcting the first speech recognition result according to the first correction operation, it further includes: Establish a speech recognition optimization model according to the first voice command and the first expected recognition result, and establish a mapping relationship between the first speech recognition result and the first expected recognition result; perform speech recognition according to the speech recognition optimization model and the mapping relationship; It further includes: Collect the user's second voice command, and determine whether the second voice command is the same as the first voice command; if the second voice command is the same as the first voice command, recognize the second voice command based on the speech recognition optimization model to obtain a second speech recognition result, and determine whether the second speech recognition result meets the user's first expected recognition result; if the second speech recognition result does not meet the user's first expected recognition result, obtain the correction result corresponding to the first speech recognition result from the mapping relationship, and correct the second speech recognition result according to the correction result corresponding to the first speech recognition result.
2. The method according to claim 1, characterized in that, After determining whether the second speech recognition result meets the user's first expected recognition result, it further includes: If the second speech recognition result meets the user's first expected recognition result, generate a first control command for the vehicle according to the second recognition result; Control the vehicle to perform corresponding control actions according to the first control command.
3. The method according to claim 1, characterized in that, After determining whether the second voice command is the same as the first voice command, it further includes: If the second voice command is not the same as the first voice command, generate a third speech recognition result according to the second voice command, and determine whether the third speech recognition result meets the user's second expected recognition result; If the third speech recognition result does not meet the second expected recognition result, receive the second correction operation generated by the user based on the third speech recognition result, and correct the third speech recognition result according to the second correction operation, so that the corrected third speech recognition result meets the second expected recognition result.
4. The method according to claim 1, characterized in that, After correcting the first speech recognition result according to the first correction operation, it further includes: Generate a second control command for the vehicle according to the corrected first speech recognition result; Control the vehicle to perform corresponding actions according to the second control command.
5. The method according to claim 1, characterized in that, Before collecting the user's first voice command, it further includes: Determine whether the speech recognition operation of the vehicle is triggered; If the speech recognition operation of the vehicle is triggered, collect the user's first voice command.
6. A device for modifying speech recognition, characterized in that, Including: A collection module, configured to collect a first voice command of a user; A judgment module, configured to generate a first voice recognition result according to the first voice command, and judge whether the first voice recognition result meets a first expected recognition result of the user; And A correction module, configured to, if the first voice recognition result does not meet the first expected recognition result, receive a first correction operation generated by the user based on the first voice recognition result, and correct the first voice recognition result according to the first correction operation, so that the corrected first voice recognition result meets the first expected recognition result.
7. An electronic device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the voice recognition modification method according to any one of claims 1-6.
8. A computer-readable storage medium, on which a computer program is stored, characterized in that, The program is executed by the processor to be used for implementing the voice recognition modification method according to any one of claims 1-4.
Citation Information
Patent Citations
Information processing method and electronic device
CN105808197A
Information inputting method, information inputting device and calculation equipment
CN106601254A