Voice Information Processing Method, Device, Vehicle, and Readable Storage Medium
By receiving and identifying user's voice information in smart cars and performing scene analysis, and directly processing the function page of preset scenes on the graphical user interface of the on-board intelligent system, the existing voice control system has solved the problems of high error rate and slow response speed, achieving higher accuracy and response efficiency.
Patent Information
- Application Number
- CN202410525382.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-04-28
AI Technical Summary
In the existing intelligent car voice control system, the voice control process link is long, resulting in a high error rate and reducing the accuracy and response speed of voice control.
By receiving voice information from users in the vehicle, recognition and scene analysis are performed. If the scene analysis result is a preset scene, the corresponding function page will be directly processed in the graphical user interface of the on-board intelligent system to reduce logical processing steps.
It improves the accuracy and response efficiency of the voice control process that conforms to the preset scenarios, reduces the generation of errors, and improves the user experience.
Smart Images

Figure CN118314899B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech processing, and particularly to a method and device for processing speech information, a vehicle, and a readable storage medium. Background Art
[0002] With the rapid development of intelligent vehicles, the functions and complexities of in-vehicle intelligent systems (such as screens) equipped in vehicles are also continuously increasing. Existing intelligent vehicles are generally equipped with a voice control system. Through voice commands, users can control the in-vehicle intelligent system accordingly, thereby controlling the corresponding functions of the vehicle.
[0003] In related technologies, in the process of implementing voice control, the voice system mainly relies on natural language recognition ability to perform semantic recognition on the user's voice commands, and then uses the backend service to perform logical processing on the voice recognition result to form a control command. Finally, the in-vehicle service executes the command according to the control command. The above voice control process has a long link, a high probability of causing errors, resulting in easy errors in the voice control process, thereby reducing the accuracy of the voice control process and affecting the user's driving experience. Moreover, the above voice control process completely relies on the backend service to perform logical processing on the voice recognition result to obtain the control command, increasing the system operation load of the backend service and affecting the response speed of the voice control process. Summary of the Invention
[0004] To solve or partially solve the problems existing in related technologies, the present application provides a method and device for processing speech information, a vehicle, and a readable storage medium, which can quickly and accurately recognize and respond to speech information that meets a preset scenario, and improve the user experience.
[0005] The first aspect of the present application provides a method for processing speech information, including:
[0006] Receiving speech information issued by a user in the vehicle cockpit;
[0007] Recognizing the speech information to obtain a corresponding speech recognition text;
[0008] Performing scene parsing on the speech recognition text to obtain a corresponding scene parsing result;
[0009] When the scene parsing result is a preset scene, processing a function page corresponding to the speech recognition text on the graphical user interface of the in-vehicle intelligent system; wherein, the speech recognition text includes function elements mapped to the function page.
[0010] In some embodiments, the performing scene parsing on the speech recognition text to obtain a corresponding scene parsing result includes:
[0011] Determine the sentence type of the speech recognition text according to the keywords in the speech recognition text;
[0012] When the sentence type belongs to a preset sentence type, determine that the scene analysis result of the speech recognition text is a preset scene.
[0013] In some embodiments, the determining the sentence type of the speech recognition text according to the keywords in the speech recognition text includes:
[0014] Obtain the action element and function element in the speech recognition text according to the keywords in the speech recognition text;
[0015] When the action element belongs to a preset operation instruction and the function element belongs to a preset function page, determine the sentence type of the speech recognition text.
[0016] In some embodiments, the processing the function page corresponding to the speech recognition text on the graphical user interface of the vehicle-mounted intelligent system when the scene analysis result is a preset scene includes:
[0017] When the scene analysis result is a preset scene, obtain the action element and function element of the speech recognition text;
[0018] Generate a corresponding function request according to the action element and the function element;
[0019] In response to the function request, process the function page corresponding to the speech recognition text on the graphical user interface of the vehicle-mounted intelligent system.
[0020] In some embodiments, the generating a corresponding function request according to the action element and the function element includes:
[0021] Query the mark corresponding to the function page in the pre-stored function mapping information according to the function element;
[0022] Generate a corresponding function request according to the action element and the mark.
[0023] In some embodiments, the method further includes:
[0024] Pre-generate function mapping information according to each function page and the corresponding mark and store it.
[0025] In some embodiments, the method further includes:
[0026] When the scene analysis result is not a preset scene, process the speech recognition text through a preset speech processing system to generate a corresponding speech instruction for the vehicle-mounted intelligent system to respond to.
[0027] The second aspect of the present application provides a voice information processing device, including:
[0028] An information receiving module, configured to receive voice information sent by a user in a vehicle cockpit;
[0029] A voice recognition module, configured to recognize the voice information to obtain a corresponding voice recognition text;
[0030] A scene analysis module, configured to perform scene analysis on the voice recognition text to obtain a corresponding scene analysis result;
[0031] A display processing module, configured to process a function page corresponding to the voice recognition text on a graphical user interface of an in-vehicle intelligent system when the scene analysis result is a preset scene; wherein, the voice recognition text includes function elements mapped to the function page.
[0032] The third aspect of the present application provides a vehicle, including:
[0033] A processor; and
[0034] A memory, on which executable code is stored, and when the executable code is executed by the processor, the processor is caused to execute the method as described above.
[0035] The fourth aspect of the present application provides a computer-readable storage medium, on which executable code is stored, and when the executable code is executed by a processor of a vehicle, the processor is caused to execute the method as described above.
[0036] The technical solution provided by the present application may include the following beneficial effects:
[0037] In the technical solution of the present application, the voice information sent by the user is recognized to obtain a corresponding voice recognition text, and then the voice recognition text is subjected to scene analysis; when the scene analysis result is a preset scene, a function page corresponding to the voice recognition text is processed on a graphical user interface of an in-vehicle intelligent system. By the above method, in the voice control process that meets the preset scene, the process of obtaining an operation instruction according to logical processing can be omitted, the generation of errors can be reduced, the correct rate of instruction execution can be improved, and the voice information of the preset scene can be "directly" controlled in the in-vehicle intelligent system, thereby improving the response efficiency of voice control.
[0038] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings
[0039] The above and other objects, features, and advantages of the present application will become more apparent by describing the exemplary embodiments of the present application in more detail with reference to the accompanying drawings. In the exemplary embodiments of the present application, the same reference numerals generally represent the same components.
[0040] Figure 1 is a schematic flowchart of the voice information processing method shown in the embodiments of the present application;
[0041] Figure 2 is another schematic flowchart of the voice information processing method shown in the embodiments of the present application;
[0042] Figure 3 is another schematic flowchart of the voice information processing method shown in the embodiments of the present application;
[0043] Figure 4 is a schematic structural diagram of the voice information processing device shown in the embodiments of the present application;
[0044] Figure 5 is another schematic structural diagram of the voice information processing device shown in the embodiments of the present application;
[0045] Figure 6 is another schematic structural diagram of the voice information processing device shown in the embodiments of the present application;
[0046] Figure 7 is a schematic structural diagram of the vehicle shown in the embodiments of the present application. Detailed Embodiments
[0047] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0048] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0049] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, "a plurality of" means two or more, unless otherwise specifically defined.
[0050] In the related art, in the process of implementing voice control in a voice system, it generally includes: S10, obtaining a voice command of a user; S20, performing semantic recognition on the voice command of the user by using natural language recognition ability; S30, performing logical processing on the voice recognition result by using a back-end service to form a control command; S40, executing the command by an on-terminal service according to the control command. In the above voice control process, the link from step S10 to S40 is relatively long, and the probability of causing errors is relatively large. In particular, there are likely to be errors in the logical processing of the semantic recognition result in step S30, resulting in the voice control process being prone to errors, thereby reducing the accuracy of the voice control process and affecting the driving experience of the user. Moreover, in step S30 of the above voice control process, mainly relying on the back-end service to perform logical processing on the voice recognition result to obtain the control command increases the system operation load of the back-end service and affects the response speed of the voice control process.
[0051] In view of the above problems, an embodiment of this application provides a voice information processing method, which can eliminate the process of obtaining an operation command according to logical processing during the voice control process that meets a preset scenario, reduce the generation of errors, and can achieve "direct" control of the vehicle-mounted intelligent system for the voice information of the preset scenario, improving the response efficiency of the voice control process.
[0052] The technical solutions of the embodiments of this application are described in detail below with reference to the accompanying drawings. Figure 1 It is a schematic flowchart of the voice information processing method shown in an embodiment of this application.
[0053] See Figure 1 , the voice information processing method of this application includes:
[0054] S110, receiving voice information sent by a user in a vehicle cockpit.
[0055] In this step, the voice information sent by the user can be collected in real time through a microphone installed in the vehicle cockpit and then processed by the vehicle-mounted intelligent system in the vehicle. Among them, the vehicle-mounted intelligent system can be an intelligent terminal with a UI display interface.
[0056] S120, perform speech recognition on the speech information to obtain the corresponding speech recognition text.
[0057] In this step, perform speech recognition on the acquired speech information according to the related technology, and convert the speech information into the corresponding speech recognition text.
[0058] It should be understood that the speech recognition text can be in the text data format. For example, the in-vehicle intelligent system collects the user's speech information as "XXX, help me open the air conditioner adjustment page, XXX". After semantic analysis of the above speech information, the above speech information can be converted into a speech recognition text with the corresponding content of "XXX, help me open the air conditioner adjustment page, XXX".
[0059] In some embodiments, during the process of performing speech recognition on the speech information, the obtained speech recognition text can be corrected through semantic analysis. That is to say, when the content of the speech information is expressed incorrectly, it is corrected through semantic analysis and the corrected speech recognition text is obtained. In this way, the accuracy of subsequent parsing of the speech recognition text can be improved.
[0060] S130, perform scene parsing on the speech recognition text to obtain the corresponding scene parsing result.
[0061] In this step, perform scene parsing according to the obtained speech recognition text, so as to judge the scene type corresponding to the speech information, that is, obtain the corresponding scene parsing result.
[0062] In some embodiments, the scene parsing result can include: a preset scene or a non-preset scene. Among them, the preset scene can be a situation where the speech information sent by the user is for processing a preset function page. The non-preset scene can be a situation where the speech information sent by the user is for processing a non-preset function page.
[0063] S140, when the scene parsing result is a preset scene, process the function page corresponding to the speech recognition text on the graphical user interface of the in-vehicle intelligent system; wherein, the speech recognition text contains function elements mapped to the function page.
[0064] In this step, when it is determined that the scene parsing result is a preset scene according to the scene parsing result, according to the function elements mapped to the function page in the speech recognition text, process the corresponding function page on the graphical user interface of the in-vehicle intelligent system.
[0065] Among them, the function page can be a specific page displayed on the graphical user interface of the in-vehicle intelligent system, such as function pages like the air conditioner adjustment page, the vehicle body inspection page, the map navigation page, etc.; the function element can be a noun with a specific function located in the speech recognition text, such as "air conditioner adjustment page".
[0066] In some implementations, the processing of the function page may include but is not limited to: performing an opening operation or a closing operation. When performing an opening operation on the corresponding function page in the graphical user interface of the vehicle-mounted intelligent system, the function page that has not been activated may be activated and displayed on top, or the function page that has been activated and located in the background may be displayed on top, so that the user can interact with the function page through the graphical user interface of the vehicle-mounted intelligent system. Accordingly, when performing a closing operation on the corresponding function page in the graphical user interface of the vehicle-mounted intelligent system, the function page that has been displayed on top or that has been activated and located in the background may be closed.
[0067] In some embodiments, a preset scene may correspond to multiple function pages. It should be understood that the function pages in the in-vehicle intelligent system can be divided into two types: local pages and third-party pages. Local pages may refer to the function pages that come with the system in the in-vehicle intelligent system, such as the air conditioning adjustment page, volume adjustment page, address book function page, etc. in the in-vehicle intelligent system. Among them, multiple function pages corresponding to the preset scene may refer to local pages. That is to say, in the technical solution of the present application, when the scene analysis result is a preset scene, the function page of the local page may be processed in the graphical user interface of the in-vehicle intelligent system. Of course, the third-party page refers to the page of the third-party application software installed in the in-vehicle intelligent system.
[0068] It should be understood that the user's voice information is diverse and cannot be exhaustively listed. By setting the preset scenarios to cover the local function pages, and using the situations in which the user directly controls the commonly used function pages as preset scenarios, the probability of users making errors during voice control can be effectively reduced, thereby improving the stability of the voice control process, while improving the efficiency of using local functions and improving user experience.
[0069] In this embodiment, the technical solution of the present application obtains a voice recognition text by recognizing the voice information sent by the user, and performs scene analysis on the voice recognition text. When the scene analysis result is a preset scene, the function page corresponding to the functional element of the voice recognition text is displayed or closed in the graphical user interface of the vehicle-mounted intelligent system. Through the above method, in the voice control process that meets the preset scene, the process of obtaining the operation instruction according to the logical processing can be omitted, the error generation can be reduced, and the voice information of the preset scene can be "directly" controlled in the vehicle-mounted intelligent system, which improves the response efficiency of the voice control process and improves the user experience.
[0070] Figure 2 FIG. 1 is another flow chart of the voice information processing method shown in the embodiment of the present application. Figure 1 The embodiments are further described on the basis of the embodiments shown.
[0071] See also Figure 2, the voice information processing method of the present application includes:
[0072] S210, receiving the voice information issued by the user in the vehicle cockpit.
[0073] In this step, the voice information issued by the user in the vehicle cockpit is received.
[0074] S220, identifying the voice information to obtain the corresponding voice recognition text.
[0075] In this step, the obtained voice information is identified according to the related technology, and the voice information is converted into the corresponding voice recognition text.
[0076] S230, performing scene parsing on the voice recognition text to obtain the corresponding scene parsing result.
[0077] In this step, scene parsing is performed according to the obtained voice recognition text, so as to judge the scene type corresponding to the voice information and obtain the corresponding scene parsing result. The type of the scene parsing result can be a preset scene that can be directly responded to and quickly reached; or a non-preset scene that cannot be directly recognized and responded to and requires further logical judgment.
[0078] S240, when the scene parsing result is a preset scene, opening or closing the function page corresponding to the voice recognition text on the graphical user interface of the in-vehicle intelligent system; wherein, the voice recognition text includes function elements mapped to the function page.
[0079] In this step, when it is determined that the scene parsing result is a preset scene according to the scene parsing result, the corresponding function page is processed on the graphical user interface of the in-vehicle intelligent system according to the function elements mapped to the function page in the voice recognition text.
[0080] S250, when the scene parsing result is a non-preset scene, processing the voice recognition text through a preset voice processing system to generate a corresponding voice command for the in-vehicle intelligent system to respond.
[0081] In this step, when it is determined that the scene parsing result is a non-preset scene according to the scene parsing result, the voice recognition text is logically processed through a preset voice processing system to generate a voice command corresponding to the voice recognition text, so that the in-vehicle intelligent system responds according to the voice command.
[0082] It should be understood that the scene parsing result can be a preset scene or a non-preset scene, that is to say, the voice information initiated by the user can obtain the scene parsing result of the preset scene or the scene parsing result of the non-preset scene. According to the different scene parsing results, the technical solution of the present application can realize different response controls for the in-vehicle intelligent system.
[0083] In some embodiments, the preset voice processing system processes the speech recognition text. It can be that the voice processing system in the related art processes semantic parsing, logical processing, and instruction output in sequence according to the speech recognition text, so as to obtain the voice instruction to be displayed or closed on the graphical user interface of the vehicle-mounted intelligent system. Among them, the preset voice processing system can be obtained by using the existing technology, which will not be elaborated here.
[0084] In this embodiment, for the technical solution of the present application, when the scene parsing result is a preset scene, the function page corresponding to the speech recognition text is processed on the graphical user interface of the vehicle-mounted intelligent system, so as to realize the "direct" control of the voice information of the preset scene in the vehicle-mounted intelligent system; when the scene parsing result is a non-preset scene, the preset voice processing system processes the speech recognition text to generate a corresponding voice instruction for the vehicle-mounted intelligent system to respond. Through the above method, according to the matching result of the voice information and the preset scene, different response controls can be realized for the vehicle-mounted intelligent system, making the voice control process more flexible in response, realizing the classification processing of the voice control process of the preset scene and the voice control process of the non-preset scene, reducing the probability of the voice processing system generating incorrect voice instructions, improving the accuracy of the voice control process, and effectively improving the response efficiency of the voice control process.
[0085] Figure 3 It is another flow schematic diagram of the voice information processing method shown in the embodiment of the present application. This embodiment is further elaborated on the basis of Figure 1 the embodiment shown.
[0086] See Figure 3 , the voice information processing method of the present application includes:
[0087] S310, receiving the voice information sent by the user in the vehicle cockpit.
[0088] S320, recognizing the voice information to obtain the corresponding speech recognition text.
[0089] S330, determining the sentence type of the speech recognition text according to the keywords in the speech recognition text.
[0090] In this step, the keywords in the obtained speech recognition text are recognized, and the recognized keywords are used to determine whether the sentence type of the speech recognition text is a preset sentence type.
[0091] Among them, the keywords can be set in advance. Specifically, the keywords can include verbs and functional nouns. The verbs can be, for example, verbs such as "open", "close", "start", "exit", etc., and the functional nouns can be, for example, page words with specific functions such as "air conditioning adjustment page", "volume adjustment page", etc.
[0092] For example, if the content of the speech recognition text is "XXX, help me open the air conditioner adjustment page, XXX", or "XXX, help me open the air conditioner adjustment page, XXX", then the keywords "open" and "air conditioner adjustment page" can be recognized according to the content of the above speech recognition text.
[0093] In some embodiments, the sentence type can be determined by combining the recognized keywords. For example, if the recognized keywords are "open" and "air conditioner adjustment page", then the sentence type is determined as "open" + "air conditioner adjustment page".
[0094] In some specific embodiments, determining the sentence type of the speech recognition text according to the keywords of the speech recognition text may include the following steps:
[0095] S331, obtain the action element and the function element in the speech recognition text according to the keywords of the speech recognition text.
[0096] In this step, the action element and the function element are obtained from the keywords extracted from the speech recognition text.
[0097] Among them, the action element may refer to a specific verb in the speech recognition text, such as "open" or "close", which is only an example here and is not limited; the function element may refer to a specific functional noun in the speech recognition text, such as "air conditioner adjustment page".
[0098] S332, when the action element belongs to the preset operation instruction and the function element belongs to the preset function page, determine the sentence type of the speech recognition text.
[0099] In this step, it is judged whether the action element belongs to the preset operation instruction and whether the function element belongs to the preset function page. When the action element belongs to the preset operation instruction and the function element belongs to the preset function page, then the sentence type of the speech recognition text is determined as the preset sentence type, otherwise it is a non-preset sentence type.
[0100] Among them, the preset operation instruction may be a specific operation instruction defined in advance, such as open or close; the preset function page may be the name of a specific function page that can be displayed on the screen defined in advance, such as the air conditioner adjustment page.
[0101] In some embodiments, the preset function page may be a function page stored in the function mapping information in advance. Among them, the function mapping information stored in advance may be a mapping table stored in advance, and in this mapping table, the preset function page information can be recorded. Specifically, the function page information may include, but is not limited to: the name of the function page.
[0102] S340. When the sentence pattern type belongs to the preset sentence pattern, determine that the scene analysis result of the speech recognition text is the preset scene.
[0103] In this step, when the sentence pattern type is the preset sentence pattern, then determine that the scene analysis result corresponding to the speech recognition text is the preset scene.
[0104] Among them, the preset sentence pattern can be: preset action element + preset function element. Specifically, the preset sentence pattern can be composed of a preset verb + the name of the preset function page.
[0105] S350. When the scene analysis result is the preset scene, obtain the action element and function element of the speech recognition text.
[0106] In this step, when the scene analysis result is determined to be the preset scene, then obtain the action element and function element of the speech recognition text.
[0107] In some embodiments, the action element and function element can be obtained from the keywords of the speech recognition text, so that the extraction efficiency of the action element and function element can be improved.
[0108] S360. Generate a corresponding function request according to the action element and function element.
[0109] In this step, generate a function request for controlling the corresponding function page to execute the preset operation instruction corresponding to the action element according to the obtained action element and function element.
[0110] In some specific embodiments, generating a corresponding function request according to the action element and function element can be implemented according to the following steps:
[0111] S361. Query the mark corresponding to the function page in the pre-stored function mapping information according to the function element.
[0112] In this step, query the corresponding function page in the pre-stored function mapping information according to the function element, and then obtain the mark corresponding to the matched function page.
[0113] It should be understood that the pre-stored function mapping information can be a pre-stored mapping table, and each function page in the mapping table can have a corresponding mark. It should be noted that the mark can refer to the ID number of each function page stored in the mapping table.
[0114] In some embodiments, the technical solution of the present application can also pre-generate function mapping information according to each function page and the corresponding mark and store it. Among them, the function pages in the function mapping information can be local function pages that are opened or closed by the vehicle-mounted intelligent system in the graphical user interface.
[0115] In some embodiments, the names of associated function pages may have the same label (ID number). In this way, it is convenient to combine different but associated functions into the same function page for control. For example, the air conditioner adjustment page and the air volume adjustment page can adopt the same label. In this way, the air conditioner adjustment and the air volume adjustment can be combined into the same function page, which is convenient for users to perform interactive operations.
[0116] S362. Generate a corresponding function request according to the action element and the label.
[0117] In this step, according to the action element and the label, obtain the corresponding preset operation instruction and the ID number of the corresponding function page as input parameters, and generate a function request for controlling the in-vehicle intelligent system. It should be understood that the function request can be used to request to control the function page corresponding to the speech recognition text in the graphical user interface of the in-vehicle intelligent system, where the processing method corresponds to the preset operation instruction, that is, it can be to open or close the corresponding function page.
[0118] S370. Respond to the function request and process the function page corresponding to the speech recognition text in the graphical user interface of the in-vehicle intelligent system.
[0119] In this step, respond to the generated function request to open or close the function page corresponding to the speech recognition text in the graphical user interface of the in-vehicle intelligent system.
[0120] It should be understood that the graphical user interface of the in-vehicle intelligent system, according to the preset operation instruction corresponding to the action element, such as opening or closing the function page with the corresponding ID number.
[0121] In this embodiment, the technical solution of the present application determines the sentence type of the speech recognition text according to the keywords in the speech recognition text, and uses whether the sentence type belongs to the preset sentence type to determine whether the scene parsing result of the speech recognition text is the preset scene, so as to quickly judge whether the user's speech information corresponds to the preset scene, and generates a function request corresponding to the preset scene by using the action element and the function element in the speech recognition text, so as to process the function page corresponding to the speech recognition text in the graphical user interface of the in-vehicle intelligent system. Through the above method, the processing of the function page corresponding to the preset scene is abstracted, such as the start or close operation of the function page, to form a direct response effect, optimize the processing logic of the voice command, and improve the response speed and accuracy of the voice control process.
[0122] Corresponding to the foregoing method embodiment for realizing application functions, the present application also provides a voice information processing device, a vehicle, and corresponding embodiments.
[0123] Figure 4It is a schematic structural diagram of the voice information processing device shown in the embodiments of the present application.
[0124] See Figure 4 , the voice information processing device 400 of the present application includes an information receiving module 410, a voice recognition module 420, a scene parsing module 430, and a display processing module 440. Among them:
[0125] The information receiving module 410 is used to receive the voice information issued by the user in the vehicle cockpit.
[0126] The voice recognition module 420 is used to recognize the voice information to obtain the corresponding voice recognition text.
[0127] The scene parsing module 430 is used to perform scene parsing on the voice recognition text to obtain the corresponding scene parsing result.
[0128] The display processing module 440 is used to process the function page corresponding to the voice recognition text on the graphical user interface of the in-vehicle intelligent system when the scene parsing result is a preset scene; wherein, the voice recognition text includes function elements mapped to the function page.
[0129] In some embodiments, the display processing module 440 can also be used to process the voice recognition text through a preset voice processing system to generate a corresponding voice command for the in-vehicle intelligent system to respond when the scene parsing result is a non-preset scene.
[0130] Figure 5 It is another schematic structural diagram of the voice information processing device shown in the embodiments of the present application.
[0131] See Figure 5 , the voice information processing device 400 of the present application includes an information receiving module 410, a voice recognition module 420, a scene parsing module 430, a display processing module 440, and a mapping module 450.
[0132] Among them, the mapping module 450 is used to generate function mapping information and store it in advance according to each function page and the corresponding mark.
[0133] Figure 6 It is another schematic structural diagram of the voice information processing device shown in the embodiments of the present application.
[0134] See Figure 6 , in the voice information processing device 400 of the present application, the scene parsing module 430 may include: a sentence pattern determination module 431 and a scene determination module 432; the display processing module 440 may include: an element acquisition module 441, a request generation module 442, and a page processing module 443.
[0135] The sentence pattern type module 431 is used to determine the sentence pattern type of the speech recognition text according to the keywords in the speech recognition text.
[0136] In some specific embodiments, the sentence pattern determination module 431 is used to obtain the action element and function element in the speech recognition text according to the keywords in the speech recognition text; when the action element belongs to a preset operation instruction and the function element belongs to a preset function page, determine the sentence pattern type of the speech recognition text.
[0137] The scene determination module 432 is used to determine that the scene parsing result of the speech recognition text is a preset scene when the sentence pattern type belongs to a preset sentence pattern.
[0138] The element acquisition module 441 is used to acquire the action element and function element of the speech recognition text when the scene parsing result is a preset scene.
[0139] The request generation module 442 is used to generate a corresponding function request according to the action element and function element.
[0140] In some specific embodiments, the request generation module 442 is used to query the tag corresponding to the function page in the pre-stored function mapping information according to the function element; generate a corresponding function request according to the action element and the tag.
[0141] The page processing module 443 is used to respond to the function request and process the function page corresponding to the speech recognition text in the graphical user interface of the vehicle-mounted intelligent system.
[0142] In this embodiment, the technical solution of the present application performs scene parsing on the speech recognition text recognized from the voice information sent by the user. When the scene parsing result is a preset scene, the function page corresponding to the speech recognition text is opened or closed in the graphical user interface of the vehicle-mounted intelligent system. By the above method, the voice control that meets the preset scene can be directly responded, the process of obtaining the operation instruction according to the logical processing can be omitted, the generation of errors can be reduced, and the "direct access" control of the vehicle-mounted intelligent system for the voice information of the preset scene can be realized, the response efficiency of the voice control process can be improved, and the user experience can be improved.
[0143] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.
[0144] Figure 7 It is a schematic structural diagram of a vehicle shown in an embodiment of the present application.
[0145] See Figure 7 , the vehicle 1000 includes a memory 1010 and a processor 1020.
[0146] The processor 1020 can be a Central Processing Unit (CPU), or it can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0147] The memory 1010 can include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM can store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage device can be a read-write storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory 1010 can include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory 1010 can include a removable storage device that can be read and / or written, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, super density disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and instantaneous electronic signals transmitted wirelessly or wired.
[0148] An executable code is stored on the memory 1010. When the executable code is processed by the processor 1020, it can cause the processor 1020 to execute some or all of the methods described above.
[0149] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, which includes computer program code instructions for performing some or all of the steps in the above method of the present application.
[0150] Alternatively, the present application can also be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium), on which executable code (or a computer program or computer instruction code) is stored. When the executable code (or the computer program or computer instruction code) is executed by a processor of a vehicle (or a server, etc.), the processor is caused to execute some or all of the steps of the above method according to the present application.
[0151] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.
Claims
1. A method for processing speech information, characterized in that: include: Receive voice messages from users in the vehicle cabin; Recognize the voice information to obtain corresponding voice recognition text; The speech recognition text is subjected to scene analysis to obtain a corresponding scene analysis result; wherein, according to the keywords of the speech recognition text, the action elements and function elements in the speech recognition text are obtained; when the action element belongs to a preset operation instruction, and the function element belongs to a preset function page, the sentence type of the speech recognition text is determined; when the sentence type belongs to a preset sentence type, the scene analysis result of the speech recognition text is determined to be a preset scene; When the scene analysis result is a preset scene, a function page corresponding to the speech recognition text is processed in a graphical user interface of the in-vehicle intelligent system; wherein the speech recognition text contains function elements mapped to the function page.
2. The method according to claim 1, characterized in that The scene analysis result includes: a preset scene or a non-preset scene; The preset scenario is a situation where the voice message sent by the user is used to process a preset function page; the non-preset scenario is a situation where the voice message sent by the user is used to process a non-preset function page.
3. The method according to claim 1, characterized in that The keywords include verbs and functional nouns; The preset sentence pattern is: preset action element+preset function element.
4. The method according to claim 1, characterized in that When the scene analysis result is a preset scene, processing a function page corresponding to the speech recognition text in the graphical user interface of the vehicle-mounted intelligent system includes: When the scene analysis result is a preset scene, obtaining action elements and function elements of the speech recognition text; Generate a corresponding function request according to the action element and the function element; In response to the function request, a function page corresponding to the speech recognition text is processed on a graphical user interface of the in-vehicle intelligent system.
5. The method according to claim 4, characterized in that The generating a corresponding function request according to the action element and the function element includes: According to the functional element, searching for a tag corresponding to the functional page in pre-stored functional mapping information; A corresponding function request is generated according to the action element and the tag.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Function mapping information is generated and stored in advance according to each function page and the corresponding mark.
7. The method according to claim 1, characterized in that The method further comprises: When the scene analysis result is a non-preset scene, the voice recognition text is processed by a preset voice processing system to generate a corresponding voice command for the in-vehicle intelligent system to respond.
8. A voice information processing device, characterized in that: include: An information receiving module, used to receive voice information sent by a user in the vehicle cabin; A speech recognition module, used to recognize the speech information and obtain corresponding speech recognition text; A scene analysis module, used to perform scene analysis on the speech recognition text to obtain a corresponding scene analysis result; wherein, according to the keywords of the speech recognition text, the action elements and function elements in the speech recognition text are obtained; when the action element belongs to a preset operation instruction and the function element belongs to a preset function page, the sentence type of the speech recognition text is determined; when the sentence type belongs to a preset sentence type, the scene analysis result of the speech recognition text is determined to be a preset scene; A display processing module is used to process the function page corresponding to the speech recognition text in the graphical user interface of the vehicle-mounted intelligent system when the scene analysis result is a preset scene; wherein the speech recognition text contains function elements mapped to the function page.
9. A vehicle, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having executable code stored thereon, characterized in that: When the executable code is executed by a processor of a vehicle, the processor is caused to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Voice recognition method and device, electronic equipment and storage medium
CN110675870A