Voice control method and device, vehicle and storage medium

By recognizing and executing the target sequence of in-vehicle voice control commands, the problem of in-vehicle voice control technology being unable to satisfy multiple command intentions has been solved, achieving an extremely smooth voice control experience for users.

CN118098229BActive Publication Date: 2026-01-02CHERY AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410166038.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-05
Publication Date
2026-01-02
Estimated Expiration
2044-02-05

AI Technical Summary

Technical Problem

Existing in-vehicle voice control technology cannot simultaneously satisfy multiple voice command intentions of users, resulting in a poor user experience.

Method used

By recognizing one or more voice control commands in the user's voice content and executing multiple voice control commands sequentially according to the target execution order, including vehicle control, navigation, music, weather and telephone commands, the system identifies voice endpoint detection values ​​to determine the number of commands and adjusts ambiguous commands to be executed at the end when there are ambiguous commands, guiding the user to issue clear commands or complete ambiguous commands.

Benefits of technology

It enables users to execute multiple commands simultaneously by speaking, enhancing the user's voice control experience, providing an extremely smooth operating feel, and solving the problems of inaccurate command recognition and lack of feedback during execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118098229B_ABST
    Figure CN118098229B_ABST
Patent Text Reader

Abstract

The application discloses a voice control method and device, a vehicle and a storage medium. The method comprises the following steps: obtaining voice content of a user; identifying one or more voice control instructions in the voice content; and if the voice content comprises multiple voice control instructions, sequentially executing the multiple voice control instructions according to a target execution sequence. Thus, the problem that multiple voice instruction intentions of a user cannot be simultaneously met in the prior art and the user experience is poor is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of vehicle voice technology, in particular to a voice control method and device, a vehicle and a storage medium. BACKGROUND

[0002] The current vehicle voice control technology evolves from single-intent natural speech recognition to multi-intent natural speech recognition, supporting vehicle control operations of arbitrary voice instructions of users and vehicle control operations of various combinations of voice instructions.

[0003] At present, major car companies support at most three instructions of users and cross two function domains, but when multiple instructions are issued by the user at the same time, or the function domain span between the instructions is more, there are problems such as inaccurate instruction recognition and no feedback during execution, which cannot meet multiple instruction intentions of the user at the same time, further causing poor user experience. SUMMARY

[0004] The present application provides a voice control method, device, vehicle and storage medium to solve the problems of not being able to meet multiple voice instruction intentions of the user at the same time and poor user experience in related technologies.

[0005] The first aspect of the present application provides a voice control method, comprising the following steps: obtaining voice content of a user; identifying one or more voice control instructions in the voice content; if the voice content includes multiple voice control instructions, executing the multiple voice control instructions in turn according to a target execution order.

[0006] Optionally, in an embodiment of the present application, the voice control instruction includes one or more of a vehicle control instruction, a navigation instruction, a music instruction, a weather instruction and a telephone instruction.

[0007] Optionally, in an embodiment of the present application, identifying one or more voice control instructions in the voice content includes: obtaining a voice endpoint detection value of each voice control instruction in the voice content; if the voice endpoint detection value is less than a preset value, it is determined that the voice content includes multiple voice control instructions, otherwise it is determined that the voice content includes one voice control instruction.

[0008] Optionally, in an embodiment of the present application, executing the multiple voice control instructions includes: if the voice control instruction is a vehicle control instruction, processing each voice control instruction to obtain a semantic result of each voice control instruction, converting the semantic result into a vehicle control instruction, and delivering the vehicle control instruction to a regional control unit to execute the vehicle control instruction using the regional control unit; if the voice control instruction is a navigation instruction, a music instruction, a weather instruction or a telephone instruction, the voice control instruction is transmitted to the corresponding software to execute the voice control instruction.

[0009] Optionally, in an embodiment of the present application, the plurality of voice control instructions are executed in the target execution order, comprising: identifying whether there is an ambiguous voice control instruction in the plurality of voice control instructions; if there is an ambiguous voice control instruction, adjusting the ambiguous voice control instruction to be executed at the end, otherwise executing the plurality of voice control instructions in the order of the instructions.

[0010] Optionally, in an embodiment of the present application, after the plurality of voice control instructions are executed in the target execution order, comprising: obtaining an execution result of the voice control instructions; and broadcasting the execution result of the voice control instructions after the voice control instructions are executed, or broadcasting the execution result of all the voice control instructions after all the voice control instructions are executed.

[0011] Optionally, in an embodiment of the present application, after the ambiguous voice control instruction is adjusted to be executed at the end, comprising: guiding the user to issue a clear voice control instruction; and / or guiding the user to complete the ambiguous voice control instruction.

[0012] An embodiment of the second aspect of the present application provides a voice control device, comprising: an obtaining module configured to obtain voice content of a user; an identifying module configured to identify one or more voice control instructions in the voice content; and an executing module configured to execute the plurality of voice control instructions in the target execution order if the voice content includes a plurality of voice control instructions.

[0013] Optionally, in an embodiment of the present application, the voice control instruction comprises one or more of a vehicle control instruction, a navigation instruction, a music instruction, a weather instruction, and a phone instruction.

[0014] Optionally, in an embodiment of the present application, the identifying module is further configured to: obtain a voice endpoint detection value of each voice control instruction in the voice content; and determine that the voice content includes a plurality of voice control instructions if the voice endpoint detection value is less than a preset value, otherwise determine that the voice content includes one voice control instruction.

[0015] Optionally, in an embodiment of the present application, the executing module is further configured to: if the voice control instruction is a vehicle control instruction, process each voice control instruction to obtain a semantic result of each voice control instruction, convert the semantic result into a vehicle control instruction, and deliver the vehicle control instruction to a regional control unit to execute the vehicle control instruction by using the regional control unit; and if the voice control instruction is a navigation instruction, a music instruction, a weather instruction, or a phone instruction, transmit the voice control instruction to a corresponding software to execute the voice control instruction.

[0016] Optionally, in an embodiment of the present application, the execution module is further configured to: identify whether there is an ambiguous voice control instruction in the plurality of voice control instructions; if there is an ambiguous voice control instruction, adjust the ambiguous voice control instruction to be executed at the end, otherwise execute the plurality of voice control instructions according to the order of the voice control instructions.

[0017] Optionally, in an embodiment of the present application, further comprising: a broadcasting module configured to: after the plurality of voice control instructions are executed according to the target execution order, obtain an execution result of the voice control instructions; and after the voice control instructions are executed, broadcast the execution result of the voice control instructions, or after all the voice control instructions are executed, broadcast the execution result of all the voice control instructions.

[0018] Optionally, in an embodiment of the present application, after the ambiguous voice control instruction is adjusted to be executed at the end, the method further comprises: guiding the user to issue a clear voice control instruction; and / or guiding the user to complete the ambiguous voice control instruction.

[0019] A third aspect of embodiments of the present application provides a vehicle, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to perform the voice control method of the above embodiments.

[0020] A fourth aspect of embodiments of the present application provides a computer readable storage medium having a computer program stored thereon, the program being executed by a processor to perform the voice control method of the above embodiments.

[0021] Therefore, the present application has at least the following beneficial effects:

[0022] Embodiments of the present application can identify voice control instructions in the voice content of the user, and execute the plurality of voice control instructions according to the target execution order, so as to meet the multiple instruction intentions of the user, and bring the user an extremely smooth voice control instruction experience, and improve the experience of the user. Therefore, the technical problem that the multiple voice instruction intentions of the user cannot be met simultaneously in the related art, and the experience of the user is poor, is solved.

[0023] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0024] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, including the appended drawings.

[0025] Figure 1 A flowchart of the voice control method according to the embodiments of the present application is provided.

[0026] Figure 2 This is a flowchart of a voice control method provided according to an embodiment of this application;

[0027] Figure 3 This is a schematic diagram of a voice control device provided according to an embodiment of this application;

[0028] Figure 4 This is a structural schematic diagram of a vehicle provided according to an embodiment of this application. Detailed Implementation

[0029] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0030] The voice control method, apparatus, vehicle, and storage medium of this application are described below with reference to the accompanying drawings. Addressing the issues mentioned in the background section regarding the current inability to achieve a seamless user experience where commands can be given and executed simultaneously, and the inevitable interruption of streaming control by multiple rounds of dialogue during streaming conversations, this application provides a voice control method. In this method, when a user issues multiple voice control commands, the commands are executed according to the target execution order, enabling simultaneous speaking and execution of commands and satisfying multiple user command intentions. This solves the problems in related technologies where multiple user voice command intentions cannot be satisfied simultaneously, resulting in a poor user experience.

[0031] Specifically, Figure 1 This is a flowchart illustrating a voice control method provided in an embodiment of this application.

[0032] like Figure 1 As shown, the voice control method includes the following steps:

[0033] In step S101, the user's voice content is acquired.

[0034] The user's voice content can include one or more voice control commands, such as "open the car window, turn on the air conditioner, adjust the seat"; the user's voice content can be acquired through the vehicle's microphone.

[0035] In step S102, one or more voice control commands are identified in the voice content.

[0036] The voice control commands include one or more of the following: vehicle control commands, navigation commands, music commands, weather commands, and telephone commands. Vehicle control commands include the adjustment of windows, air conditioning, seats, etc.

[0037] It should be noted that the voice control instruction of the embodiment of the present application includes a vehicle control instruction, a navigation instruction, a music instruction, a weather instruction, and a telephone instruction, and each domain supports the first round of interaction instruction of the business vertical class, and the number of business vertical classes of the same domain instruction is not limited.

[0038] In the embodiment of the present application, identifying one or more voice control instructions in the voice content includes: obtaining a voice endpoint detection value of each voice control instruction in the voice content; if the voice endpoint detection value is less than a preset value, it is determined that the voice content includes multiple voice control instructions, otherwise it is determined that the voice content includes one voice control instruction.

[0039] The preset value can be set according to specific circumstances, such as 600ms or 700ms, which is not limited.

[0040] It can be understood that the embodiment of the present application can judge the VAD (Voice Activity Detection) value between each voice control instruction. When the VAD value is less than the preset value, such as VAD<600ms, it is determined that the voice content of the user includes multiple voice control instructions, and if VAD>600ms, it is determined that the voice content of the user includes only one voice control instruction.

[0041] In step S103, if the voice content includes multiple voice control instructions, the multiple voice control instructions are executed in turn according to the target execution order.

[0042] The multiple voice control instructions can be referred to as streaming instructions.

[0043] It can be understood that the embodiment of the present application executes multiple voice control instructions in turn according to the target execution order when there are multiple voice control instructions, so as to realize instruction saying and executing at the same time, and bring the user an extremely smooth voice instruction operation experience.

[0044] In the embodiment of the present application, the multiple voice control instructions are executed in turn according to the target execution order, including: identifying whether there is an ambiguous voice control instruction in the multiple voice control instructions; if there is an ambiguous voice control instruction, the ambiguous voice control instruction is adjusted to the end for execution, otherwise the multiple voice control instructions are executed according to the order of the instructions.

[0045] It can be understood that the application can adjust the ambiguous voice control instruction to the end of execution when there is an ambiguous voice control instruction in multiple voice control instructions, so as to prevent secondary interaction during instruction execution and affect the fluency of instruction issuance, otherwise, the voice control instruction is executed according to the order of instruction issuance, thereby solving the problem that multiple rounds of dialogues in a streaming dialogue process interrupt instruction execution.

[0046] For example, if the voice control instruction in the voice content includes "turn on the air conditioner to 24℃", "make a call", and "open the window", and "make a call" is an ambiguous user instruction, then "make a call" is adjusted to the last execution to prevent secondary dialogue. For example, if the voice control instruction in the voice content includes "turn on the air conditioner to 24℃", "open the window", and "query today's weather", and there is no ambiguous voice instruction in the voice control instruction in the voice content, then the voice control instruction is executed according to the order of instruction issuance, that is, the air conditioner is turned on to 24℃ first, then the window is opened, and finally today's weather is queried.

[0047] In the implementation of the application, after adjusting the ambiguous voice control instruction to the end of execution, the method includes guiding the user to issue a clear voice control instruction, and / or guiding the user to complete the ambiguous voice control instruction.

[0048] It can be understood that after adjusting the ambiguous voice control instruction to the end of execution, the application embodiment can guide the user to issue a clear voice control instruction, or guide the user to complete the ambiguous voice control instruction, so as to realize clear instruction issuance and execution.

[0049] For example, if the voice control instruction is navigation travel, the voice control instruction is an ambiguous user voice instruction, the user can be prompted to "please enter the correct navigation instruction, and the instruction must include the destination", or the user can be prompted to "please enter the navigation destination".

[0050] In the implementation of the application, the execution of multiple voice control instructions includes: if the voice control instruction is a vehicle control instruction, processing each voice control instruction to obtain a semantic result of each voice control instruction, converting the semantic result into a vehicle control instruction, and issuing the vehicle control instruction to a regional control unit to execute the vehicle control instruction by using the regional control unit; if the voice control instruction is a navigation instruction, a music instruction, a weather instruction, or a telephone instruction, the voice control instruction is transmitted to the corresponding software to execute the voice control instruction.

[0051] It can be understood that if the voice control instruction is a vehicle control instruction, the semantic result obtained by processing each voice control instruction is converted into a vehicle control Someip (Scalable service-Oriented Middleware over IP, IP-based scalable service-oriented middleware) instruction and is issued to the area control unit, and the vehicle control instruction is executed by using the area control unit; if the voice control instruction is a navigation instruction, a music instruction, a weather instruction, or a telephone instruction, the voice control instruction is transmitted to the corresponding software for execution.

[0052] In the implementation of the present application, after sequentially executing a plurality of voice control instructions in accordance with a target execution sequence, the following steps are included: obtaining an execution result of the voice control instruction; after the voice control instruction is executed, the execution result of the voice control instruction is broadcasted, or after all voice control instructions are executed, the execution result of all voice control instructions is broadcasted.

[0053] It can be understood that the execution result of the voice control instruction can be obtained, the execution result of the voice control instruction can be broadcasted after the voice control instruction is executed, or the execution result of the voice control instruction can be broadcasted after all voice control instructions are executed.

[0054] For example, if the plurality of voice control instructions include an instruction such as "query today's weather" that must be broadcasted to inform the user of the execution result, the execution result (such as "today's weather is sunny, northeast wind 1-2 level, temperature 23℃") is broadcasted; if there is no instruction that must be broadcasted to inform the user of the execution result among the plurality of voice control instructions, a unified broadcast such as "good" is performed after all voice control instructions are executed.

[0055] Specifically, the implementation of the present application distinguishes whether the original voice instruction or the streaming voice instruction by creating a streaming instruction (also referred to as a plurality of voice control instructions) recognition module, and enters the streaming voice instruction recognition module to perform a quick response. The streaming voice instruction module includes a demand skill recognition module, a streaming tail sentence recognition module, a TTS (Text-To-Speech, speech synthesis) broadcast module, and a post-processing module.

[0056] The demand skill recognition module is used to identify whether the skill is a streaming skill that needs to be executed; the streaming tail sentence recognition module is used to identify whether the current streaming voice is a tail sentence of a voice instruction; the TTS broadcast module is used to voice broadcast the feedback result after each instruction is issued and executed, including TTS broadcast during streaming execution, end and secondary interaction TTS broadcast; and the post-processing module is used to process secondary interaction or multi-round dialogue post-processing skills.

[0057] The vehicle machine system matched therewith includes a vehicle machine system processing module and an interface display module: the vehicle machine system processing module is used to process semantic protocols, to issue someip signals and other vehicle control type instructions after converting into vehicle control information, to provide feedback results of vehicle control instruction execution, and to provide TTS skills for which special processing is not required for part of the instructions to be broadcast; and the interface display module is used to display voice instructions issued by the user, and only occupies one row, with the execution instructions and execution results being refreshed from bottom to top, and the result types being presented as success, failure, and timeout, with the execution results being synchronized with TTS broadcasting.

[0058] Compared with the prior art, the voice processing is started after the user finishes speaking, and the corresponding vehicle control instructions are issued, the waiting time from when the user finishes speaking to when execution is started is determined according to the number of instructions issued by the user, and the time is not uniform, and the instructions cannot be executed while being spoken. The embodiments of the present application can improve the response speed of continuous multiple instructions of the user from the experience, execute the instructions while speaking, and further realize a more smooth and fast human-computer dialogue.

[0059] The voice control method of the present application will be described below through a specific embodiment, as shown in Figure 2 , which includes

[0060] I. Streamed voice instruction module

[0061] Determine whether the continuous voice instructions of the user are streamed voice instructions, by judging the VAD values between the voice instructions, if VAD < 600 ms, it is determined to be a streamed instruction; if VAD > 600 ms, it is determined to be an original voice instruction, and the normal voice instruction issuing is performed.

[0062] II. Demand instruction skill module

[0063] After determining that the voice instructions are streamed voice instructions, compare the pre-defined voice controllable domain skills including but not limited to vehicle control, navigation, music, weather, and telephone, if it is such a skill, proceed to the next step of streamed tail sentence determination; if it is not such a skill, skip execution and do not perform normal voice instruction issuing.

[0064] III. Streamed tail sentence instruction determination module

[0065] The last voice control instruction issued by the user in the streamed conversation or the instruction executed last in order on the terminal, this module includes three cases.

[0066] 1, is the end of the flow sentence instruction, with TTS broadcast. This case of TTS broadcast refers to the user's perception of the vehicle control instruction that needs to be broadcast by TTS, such as turning on the air conditioner, turning on the seat heating, etc. The user's voice instruction is issued, the voice assistant processes the semantics, sends the semantic result to the car system, the car system receives the semantic result and converts it into a vehicle control Someip instruction to issue to each ZCU (Zone Control Unit, regional control unit) of the vehicle for execution. After the operation is executed, the car system processing module returns the success or failure result of the operation to the voice assistant, and the voice assistant returns the result of this action to make a corresponding broadcast. The content of this broadcast is used as the end of this flow instruction, for example: Good, the air conditioner has been turned on.

[0067] 2, is the end of the flow sentence instruction, without TTS broadcast. This case of TTS broadcast refers to the user's visual perception of the vehicle control instruction that is very obvious, such as opening the window, adjusting the window, etc. The execution result of the instruction can not be presented through voice broadcast, and the execution process is as described in 2. The only difference is that the car system processing module returns the result of the feedback, and does not broadcast the reply of this skill, but the voice assistant makes a unified end TTS broadcast to prompt the user about the overall execution of the flow instruction, for example: Good.

[0068] 3, non-flow end sentence instruction. If this flow voice instruction is N-1 sentence instruction, it needs to enter the next step TTS broadcast module for processing.

[0069] Four, TTS broadcast module

[0070] 1, strong dependence on TTS broadcast skill. After entering the N-1 sentence flow instruction, the voice instruction is executed, and the execution process of the instruction is described in the first item in 3. The execution result of the instruction depends on the TTS broadcast to present, such as weather, vehicle condition query, etc. Intent skill, the user needs TTS broadcast to know the feedback result. If it is a non-end sentence (N-1) instruction, the strong dependence on TTS skill will produce process TTS broadcast.

[0071] 2, not dependent on TTS broadcast skill. Determine if the instruction is a multi-round dialogue. If it is not a multi-round dialogue, it is only an execution result that can be visually perceived, such as window control. The execution result of the instruction can not be presented through voice broadcast. If a multi-round dialogue is generated, the instruction enters the post-processing module, which will be processed first.

[0072] Five, post-processing module

[0073] In processing the multi-turn dialogue generated in the N-1 sentence stream instruction, post-processing needs to be performed, and the multi-turn dialogue will not generate a unified broadcast of TTS. The multi-turn dialogue will cause a secondary interaction between the voice assistant and the user, which will interrupt the current streaming processing and execution. For example, the voice instruction is navigation travel, the voice assistant judges that the voice information input by the user is an instruction with unclear intent, and then the voice assistant needs to gradually guide the user to make a clear instruction or complete the instruction content through a voice prompt.

[0074] According to the voice control method provided in the embodiments of the present application, the voice control instruction in the voice content of the user can be recognized, and the plurality of voice control instructions can be executed in the target execution order, so that the multiple instruction intents of the user can be met, and the user is brought an extremely smooth voice control instruction experience, and the experience of the user is improved.

[0075] Secondly, the voice control device provided in the embodiments of the present application is described with reference to the accompanying drawings.

[0076] Figure 3 is a block schematic diagram of the voice control device in the embodiments of the present application.

[0077] As shown in Figure 3 , the voice control device 10 includes an acquisition module 100, an identification module 200, and an execution module 300.

[0078] The acquisition module 100 is configured to acquire voice content of a user, the identification module 200 is configured to identify one or more voice control instructions in the voice content, and the execution module 300 is configured to sequentially execute the plurality of voice control instructions in a target execution order if the voice content includes the plurality of voice control instructions.

[0079] In the embodiments of the present application, the voice control instruction includes one or more of a vehicle control instruction, a navigation instruction, a music instruction, a weather instruction, and a telephone instruction.

[0080] In the embodiments of the present application, the identification module 200 is further configured to: acquire a voice endpoint detection value of each voice control instruction in the voice content; and determine that the voice content includes the plurality of voice control instructions if the voice endpoint detection value is less than a preset value, and otherwise determine that the voice content includes the one voice control instruction.

[0081] In the embodiments of the present application, the execution module 300 is further configured to: if the voice control instruction is the vehicle control instruction, process each voice control instruction to obtain a semantic result of each voice control instruction, convert the semantic result into the vehicle control instruction, and deliver the vehicle control instruction to a regional control unit to execute the vehicle control instruction by using the regional control unit; and if the voice control instruction is the navigation instruction, the music instruction, the weather instruction, or the telephone instruction, transmit the voice control instruction to a corresponding software to execute the voice control instruction.

[0082] In the embodiment of the present application, the execution module 300 is further configured to: identify whether there is an ambiguous voice control instruction in the plurality of voice control instructions; if there is an ambiguous voice control instruction, adjust the ambiguous voice control instruction to the end for execution, otherwise execute the plurality of voice control instructions according to the order of the instructions.

[0083] In the embodiment of the present application, the device 10 of the embodiment of the present application further comprises a broadcast module.

[0084] The broadcast module is configured to: after sequentially executing the plurality of voice control instructions according to the target execution order, obtain the execution result of the voice control instructions; after the voice control instructions are executed, broadcast the execution result of the voice control instructions, or after all the voice control instructions are executed, broadcast the execution result of all the voice control instructions.

[0085] In the embodiment of the present application, after adjusting the ambiguous voice control instruction to the end for execution, the method further comprises: guiding the user to issue a clear voice control instruction; and / or guiding the user to complete the ambiguous voice control instruction.

[0086] It should be noted that the foregoing explanation and description of the voice control method embodiments are also applicable to the voice control device of the embodiments, which will not be described here.

[0087] According to the voice control device provided in the embodiment of the present application, the voice control instructions in the voice content of the user can be identified, and the plurality of voice control instructions can be executed according to the target execution order, so that the multiple instruction intentions of the user can be met, the user can have an extremely smooth voice control instruction experience, and the experience of the user can be improved.

[0088] Figure 4 A structural schematic diagram of a vehicle is provided for the embodiment of the present application. The vehicle can include:

[0089] The memory 401, the processor 402, and the computer program stored in the memory 401 and executable on the processor 402.

[0090] The processor 402 executes the program to implement the voice control method provided in the above embodiments.

[0091] Further, the vehicle further includes:

[0092] The communication interface 403 is configured to communicate between the memory 401 and the processor 402.

[0093] The memory 401 is configured to store the computer program executable on the processor 402.

[0094] The memory 401 can include a high-speed RAM memory, and can also include a non-volatile memory, for example, at least one disk memory.

[0095] If the memory 401, the processor 402 and the communication interface 403 are implemented independently, the communication interface 403, the memory 401 and the processor 402 can be connected to each other through a bus and complete communication between each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 4 In the figure, only one thick line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.

[0096] Optionally, in a specific implementation, if the memory 401, the processor 402 and the communication interface 403 are integrated on a chip, the memory 401, the processor 402 and the communication interface 403 can complete communication between each other through an internal interface.

[0097] The processor 402 can be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0098] The embodiments of the present application also provide a computer readable storage medium, which has a computer program stored thereon, and the program is executed by a processor to implement the voice control method.

[0099] In the description of the application, reference to "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that a particular feature, structure, material or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the application. The illustrative appearances of the above-mentioned terms in various places in the specification are not necessarily referred to the same embodiment or example. Moreover, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples. Furthermore, in non-contradictory cases, those skilled in the art can combine and combine the features of different embodiments or examples described in the specification and the features of different embodiments or examples.

[0100] In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the application, the meaning of "N" is at least two, for example, two, three, etc., unless otherwise specifically limited.

[0101] Any process or method descriptions or descriptions of the flow diagrams in the specification can be understood as representing code modules, segments or portions of code which include one or more executable instructions for implementing the specified logic function(s) or process(es) and the preferred embodiments of the application include additional implementations in which the order of execution or the functions are not the same as those illustrated and described. It is therefore intended that the application embraces all such variations and modifications. The application also includes the combinations of the individual features recited in the claims.

[0102] It should be understood that parts of the application can be implemented in hardware, software, firmware or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. As in another embodiment, if implemented in hardware, any of the following technologies known in the art can be used in combination of one or more: discrete logic circuit with logic gate circuit for implementing logic functions on data signals, application specific integrated circuit with suitable combination logic gate circuit, programmable gate array (PGA), field programmable gate array (FPGA) and the like.

[0103] Those skilled in the art of the present technology can understand that all or part of the steps carried out by the above-mentioned embodiment method can be completed by program instructions to the relevant hardware, and the program can be stored in a computer readable storage medium. The program, when executed, includes one or a combination of steps of the method embodiments.

Claims

1. A voice control method, characterized by, The method comprises the following steps: obtaining voice content of a user; identifying one or more voice control instructions in the voice content; the identifying one or more voice control instructions in the voice content comprises: obtaining a voice endpoint detection value of each voice control instruction in the voice content; if the voice endpoint detection value is less than a preset value, it is determined that the voice content comprises multiple voice control instructions, otherwise it is determined that the voice content comprises one voice control instruction; if the voice content comprises multiple voice control instructions, the multiple voice control instructions are executed in a target execution order; the executing the multiple voice control instructions in the target execution order comprises: identifying whether there is an instruction-unclear voice control instruction in the multiple voice control instructions; if there is the instruction-unclear voice control instruction, the instruction-unclear voice control instruction is adjusted to be executed at the end, otherwise the multiple voice control instructions are executed in an instruction-issuing order; after the multiple voice control instructions are executed in the target execution order, the method comprises: obtaining an execution result of the voice control instruction; after the voice control instruction is executed, the execution result of the voice control instruction is broadcasted, or after all voice control instructions are executed, the execution results of all voice control instructions are broadcasted; after the instruction-unclear voice control instruction is adjusted to be executed at the end, the method comprises: guiding the user to issue an instruction-clear voice control instruction, and / or guiding the user to complete the instruction-unclear voice control instruction.

2. The voice control method of claim 1, wherein, The voice control instruction comprises one or more of a vehicle control instruction, a navigation instruction, a music instruction, a weather instruction, and a telephone instruction.

3. The voice control method of claim 2, wherein, The executing the multiple voice control instructions comprises: if the voice control instruction is the vehicle control instruction, processing each voice control instruction to obtain a semantic result of the voice control instruction, converting the semantic result into the vehicle control instruction, and issuing the vehicle control instruction to a regional control unit to execute the vehicle control instruction by using the regional control unit; if the voice control instruction is the navigation instruction, the music instruction, the weather instruction, or the telephone instruction, the voice control instruction is transmitted to corresponding software to execute the voice control instruction.

4. A voice control device, characterized by The method comprises: an obtaining module configured to obtain voice content of a user; an identifying module configured to identify one or more voice control instructions in the voice content; the identifying one or more voice control instructions in the voice content comprises: obtaining a voice endpoint detection value of each voice control instruction in the voice content; if the voice endpoint detection value is less than a preset value, it is determined that the voice content comprises multiple voice control instructions, otherwise it is determined that the voice content comprises one voice control instruction; The execution module is configured to sequentially execute the plurality of voice control instructions according to a target execution sequence if the voice content includes the plurality of voice control instructions; the sequentially executing the plurality of voice control instructions according to the target execution sequence includes: identifying whether there is an instruction-unclear voice control instruction in the plurality of voice control instructions; if there is the instruction-unclear voice control instruction, adjusting the instruction-unclear voice control instruction to be executed at the end, otherwise executing the plurality of voice control instructions according to an instruction-issuing sequence; after sequentially executing the plurality of voice control instructions according to the target execution sequence, including: obtaining an execution result of the voice control instruction; after the voice control instruction is executed, announcing the execution result of the voice control instruction, or after all voice control instructions are executed, announcing the execution result of all voice control instructions; after adjusting the instruction-unclear voice control instruction to be executed at the end, including: guiding a user to issue an instruction-clear voice control instruction, and / or guiding the user to complete the instruction-unclear voice control instruction.

5. A vehicle characterized by comprising: The computer program is stored in the memory and executable on the processor, and the processor executes the program to implement the voice control method according to any one of claims 1-3. The program is executed by the processor to implement the voice control method according to any one of claims 1-3.

6. A computer-readable storage medium having stored thereon a computer program, characterized in that, ​

Citation Information

Patent Citations

  • Information processing method, device, equipment, medium, product and vehicle

    CN117690428A

  • Text feature-based vehicle-mounted voice guiding method and device, equipment and medium

    CN118430535A