Voice interaction method and device of intelligent equipment and vehicle

CN120457481APending Publication Date: 2025-08-08BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380088669.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-23
Filing Date
2023-12-21
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the existing technology, the voice interaction responsiveness of vehicle-mounted terminals is poor, especially in a multi-user environment, and it is unable to effectively distinguish and respond to users who truly intend to engage in voice interaction.

Method used

By collecting user voice information, conducting semantic analysis, determining voice interaction scenarios and modes, filtering and generating voice interaction instructions in multi-person mode, and controlling smart devices to interact with multiple users, achieving full vehicle coverage.

Benefits of technology

It improves the responsiveness of voice interaction and ensures that users with real intentions can successfully execute voice commands. It is suitable for multi-person environments in vehicle terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120457481A_ABST
    Figure CN120457481A_ABST
Patent Text Reader

Abstract

The invention discloses a voice interaction method and device of intelligent equipment and a vehicle. The method comprises: acquiring voice information of a first user, performing semantic analysis on the voice information of the first user, and determining a voice interaction scene (S102); obtaining a voice interaction mode based on a corresponding relationship between the voice interaction scene and the voice interaction mode, and controlling the intelligent device to enter a multi-person mode in response to the voice interaction mode being a multi-person mode (S103); acquiring voice information of the plurality of users, and screening the voice information of the plurality of users based on the voice information of the plurality of users and the voice interaction scene to obtain voice information of at least one second user (S104); and performing semantic analysis on the voice information of the at least one second user, generating a first voice interaction instruction, and controlling the intelligent device to execute the first voice interaction instruction (S105), thereby improving the responsiveness of voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Voice interaction method and device for intelligent device and vehicle

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on the Chinese patent application with application number 202211663194.4 and application date of December 23, 2022, and claims the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby introduced into this application as a reference. Technical Field

[0003] The present disclosure relates to the field of vehicle technology, and in particular to a voice interaction method, apparatus, electronic device, storage medium, and vehicle for an intelligent device. Background Art

[0004] Currently, with the development of technologies such as artificial intelligence and natural language processing, voice interaction technology has been widely used in scenarios such as in-vehicle terminals, business processing, and mobile phones, making people's lives more convenient. For example, in the in-vehicle terminal scenario, a passenger can say "turn on the air conditioner", and the in-vehicle terminal can turn on the air conditioner through the vehicle control application, without the need for the passenger to turn it on manually. However, related technologies cannot achieve voice interaction with users who truly intend to interact with voice due to interference from other users, such as other users chatting or making phone calls. In other words, the voice interaction methods in related technologies have poor responsiveness.

[0005] Summary of the Invention

[0006] The present disclosure aims to solve one of the technical problems in the above-mentioned technologies at least to some extent.

[0007] To this end, the first objective of the present disclosure is to provide a voice interaction method for an intelligent device.

[0008] The second objective of the present disclosure is to provide a voice interaction device for an intelligent device.

[0009] A third objective of the present disclosure is to provide an electronic device.

[0010] A fourth object of the present disclosure is to provide a computer-readable storage medium.

[0011] A fifth object of the present disclosure is to provide a vehicle.

[0012] The first aspect of the present disclosure provides a voice interaction method for an intelligent device, comprising: collecting voice information of a first user, and performing semantic analysis on the voice information of the first user to determine a voice interaction scenario of the intelligent device; obtaining a voice interaction mode corresponding to the voice interaction scenario based on a correspondence between the voice interaction scenario and the voice interaction mode, and controlling the intelligent device to enter the multi-person mode in response to the voice interaction mode being a multi-person mode; collecting voice information of multiple users, and screening the voice information of the multiple users based on the voice information of the multiple users and the voice interaction scenario to obtain voice information of at least one second user; performing semantic analysis on the voice information of at least one second user to generate a first voice interaction instruction, and controlling the intelligent device to execute the first voice interaction instruction; wherein, the first user is a user who issues voice information for determining the voice interaction scenario of the intelligent device, and the second user is at least one user among the multiple users, and is a user who issues voice information for generating the first voice interaction instruction.

[0013] In addition, the voice interaction method of the smart device proposed in the above embodiment of the present disclosure may also have the following additional technical features:

[0014] In one embodiment of the present disclosure, the second user includes a third user and a fourth user, and the third user and the fourth user are different users in the same space. The semantic analysis of the voice information of at least one of the second users to generate a first voice interaction instruction includes: performing semantic analysis on the voice information of the third user to obtain that the voice interaction intention of the third user is to control the vehicle-mounted equipment of the target category, and generating a first interaction instruction based on the vehicle-mounted equipment of the target category and the identity of the third user, wherein the first interaction instruction carries the identity of the third user; performing semantic analysis on the voice information of the fourth user to obtain that the voice interaction intention of the fourth user is to control the vehicle-mounted equipment of the target category, replacing the identity of the third user in the first interaction instruction with the identity of the fourth user, and generating a second interaction instruction; wherein the first voice interaction instruction includes the first interaction instruction and the second interaction instruction.

[0015] In one embodiment of the present disclosure, the performing semantic analysis on the voice information of the fourth user to obtain the fourth user's voice interaction intention is to control the vehicle-mounted equipment of the target category, including: performing voice recognition on the fourth user's voice information to obtain the recognition text corresponding to the fourth user; judging whether the recognition text corresponding to the fourth user includes set keywords; if the recognition text corresponding to the fourth user includes the set keywords, determining that the fourth user's voice interaction intention is to control the vehicle-mounted equipment of the target category.

[0016] In one embodiment of the present disclosure, the semantic analysis of the voice information of at least one second user to generate the first voice interaction instruction includes: performing semantic analysis on the voice information of each second user to obtain the voice interaction intention of each second user; determining whether the voice interaction intentions of multiple second users are related; if the voice interaction intentions of multiple second users are related, fusing the voice interaction intentions of multiple second users to obtain a fused interaction intention; and generating the first voice interaction instruction based on the fused interaction intention.

[0017] In one embodiment of the present disclosure, based on the voice information of the plurality of users and the voice interaction scenario, the voice information of the plurality of users is screened to obtain the voice information of at least one second user, including: performing voice recognition on the voice information of the user to obtain the recognized text corresponding to the user; obtaining a template text library of the voice interaction scenario; obtaining the similarity between the recognized text corresponding to the user and the template text in the template text library; and determining the voice information of the user whose similarity is greater than or equal to a set threshold as the voice information of the second user.

[0018] In one embodiment of the present disclosure, the voice information of multiple users includes the voice information of the first user, and the method further includes: determining whether the similarity between the recognition text corresponding to the first user and the template text in the template text library is greater than or equal to the set threshold; if the similarity corresponding to the first user is less than the set threshold, controlling the smart device to switch from multi-player mode to single-player mode; performing semantic analysis on the voice information of the first user, generating a second voice interaction instruction, and controlling the smart device to execute the second voice interaction instruction.

[0019] In one embodiment of the present disclosure, controlling the smart device to enter the multiplayer mode includes: determining whether a page of an application associated with the voice interaction scenario is displayed on the display interface of the smart device, and / or whether the smart device is in full-duplex mode; if a page of an application associated with the voice interaction scenario is displayed on the display interface of the smart device, and / or the smart device is in full-duplex mode, identifying that the smart device meets the conditions for entering the multiplayer mode, and controlling the smart device to enter the multiplayer mode.

[0020] In one embodiment of the present disclosure, it also includes: if the page of the application associated with the voice interaction scenario is not displayed on the display interface of the smart device, and / or the smart device is in simplex mode, identifying that the smart device does not meet the conditions for entering multi-player mode, and controlling the smart device to enter single-player mode; performing semantic analysis on the voice information of the first user, generating a second voice interaction instruction, and controlling the smart device to execute the second voice interaction instruction.

[0021] In one embodiment of the present disclosure, it also includes: in response to the voice interaction mode being a single-player mode, controlling the smart device to enter the single-player mode; performing semantic analysis on the voice information of the first user, generating a second voice interaction instruction, and controlling the smart device to execute the second voice interaction instruction.

[0022] The second aspect of the present disclosure provides a voice interaction device for an intelligent device, comprising: a determination module for collecting voice information of a first user, and performing semantic analysis on the voice information of the first user to determine the voice interaction scenario of the intelligent device; a control module for obtaining a voice interaction mode corresponding to the voice interaction scenario based on the correspondence between the voice interaction scenario and the voice interaction mode, and controlling the intelligent device to enter the multi-person mode in response to the voice interaction mode being the multi-person mode; a screening module for collecting voice information of multiple users, and based on the voice information of the multiple users and the voice interaction scenario, screening the voice information of the multiple users to obtain voice information of at least one second user; an execution module for performing semantic analysis on the voice information of at least one second user, generating a first voice interaction instruction, and controlling the intelligent device to execute the first voice interaction instruction; wherein, the first user is a user who issues voice information for determining the voice interaction scenario of the intelligent device, and the second user is at least one user among the multiple users, and is a user who issues voice information for generating the first voice interaction instruction.

[0023] In one embodiment of the present disclosure, the second user includes a third user and a fourth user, and the third user and the fourth user are different users in the same space. The semantic analysis of the voice information of at least one of the second users to generate a first voice interaction instruction includes: performing semantic analysis on the voice information of the third user to obtain that the voice interaction intention of the third user is to control the vehicle-mounted equipment of the target category, and generating a first interaction instruction based on the vehicle-mounted equipment of the target category and the identity of the third user, wherein the first interaction instruction carries the identity of the third user; performing semantic analysis on the voice information of the fourth user to obtain that the voice interaction intention of the fourth user is to control the vehicle-mounted equipment of the target category, replacing the identity of the third user in the first interaction instruction with the identity of the fourth user, and generating a second interaction instruction; wherein the first voice interaction instruction includes the first interaction instruction and the second interaction instruction.

[0024] In one embodiment of the present disclosure, the execution module is also used to: perform voice recognition on the voice information of the fourth user to obtain the recognition text corresponding to the fourth user; determine whether the recognition text corresponding to the fourth user includes set keywords; if the recognition text corresponding to the fourth user includes the set keywords, determine that the fourth user's voice interaction intention is to control the vehicle-mounted equipment of the target category.

[0025] In one embodiment of the present disclosure, the execution module is also used to: perform semantic analysis on the voice information of each second user to obtain the voice interaction intention of each second user; determine whether the voice interaction intentions of multiple second users are related; if the voice interaction intentions of multiple second users are related, fuse the voice interaction intentions of multiple second users to obtain a fused interaction intention; and generate the first voice interaction instruction based on the fused interaction intention.

[0026] In one embodiment of the present disclosure, the screening module is also used to: perform voice recognition on the voice information of the user to obtain the recognition text corresponding to the user; obtain the template text library of the voice interaction scenario; obtain the similarity between the recognition text corresponding to the user and the template text in the template text library; and determine the voice information of the user whose similarity is greater than or equal to a set threshold as the voice information of the second user.

[0027] In one embodiment of the present disclosure, the voice information of multiple users includes the voice information of the first user, and the execution module is further used to: determine whether the similarity between the recognition text corresponding to the first user and the template text in the template text library is greater than or equal to the set threshold; if the similarity corresponding to the first user is less than the set threshold, control the smart device to switch from multi-player mode to single-player mode; perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

[0028] In one embodiment of the present disclosure, the control module is also used to: determine whether a page of an application associated with the voice interaction scenario is displayed on the display interface of the smart device, and / or whether the smart device is in full-duplex mode; if a page of an application associated with the voice interaction scenario is displayed on the display interface of the smart device, and / or the smart device is in full-duplex mode, identify that the smart device meets the conditions for entering multiplayer mode, and control the smart device to enter multiplayer mode.

[0029] In one embodiment of the present disclosure, the execution module is also used to: if the page of the application associated with the voice interaction scenario is not displayed on the display interface of the smart device, and / or the smart device is in simplex mode, identify that the smart device does not meet the conditions for entering multi-player mode, and control the smart device to enter single-player mode; perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

[0030] In one embodiment of the present disclosure, the execution module is also used to: in response to the voice interaction mode being a single-player mode, control the smart device to enter the single-player mode; perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

[0031] The third aspect embodiment of the present disclosure proposes an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the voice interaction method of the smart device as described in the first aspect embodiment of the present disclosure is implemented.

[0032] The fourth embodiment of the present application proposes a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the voice interaction method of the smart device as described in the first embodiment of the present disclosure is implemented.

[0033] The fifth aspect embodiment of the present application proposes a vehicle, comprising a voice interaction device of the smart device as described in the second aspect embodiment of the present disclosure; or the electronic device as described in the third aspect embodiment of the present disclosure; or the computer-readable storage medium as described in the fourth aspect embodiment of the present disclosure.

[0034] The correspondence between voice interaction scenarios and voice interaction modes can be taken into consideration to determine the voice interaction scenarios of smart devices, and when the voice interaction mode is multi-person mode, the smart device can be controlled to enter multi-person mode. The voice information and voice interaction scenarios of multiple users can be comprehensively considered, and the voice information of at least one second user can be screened out from the voice information of multiple users to generate a first voice interaction instruction, and the smart device can be controlled to execute the first voice interaction instruction. This helps to realize voice interaction with users who really have voice interaction intentions, improves the responsiveness of voice interaction, and can realize voice interaction between smart devices and at least one second user. It is particularly suitable for voice interaction scenarios of vehicle-mounted terminals, and can achieve full vehicle coverage of voice interaction of vehicle-mounted terminals.

[0035] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The above and / or additional aspects and advantages of the present disclosure will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0037] FIG1 is a flow chart of a voice interaction method for a smart device according to an embodiment of the present disclosure;

[0038] FIG2 is a schematic diagram of a vehicle according to one embodiment of the present disclosure;

[0039] FIG3 is a flow chart of a voice interaction method for a smart device according to another embodiment of the present disclosure;

[0040] FIG4 is a flow chart of a voice interaction method for a smart device according to another embodiment of the present disclosure;

[0041] FIG5 is a flow chart of a voice interaction method for a smart device according to another embodiment of the present disclosure;

[0042] FIG6 is a flow chart of a voice interaction method for a smart device according to another embodiment of the present disclosure;

[0043] FIG7 is a flow chart of a voice interaction method for a smart device according to another embodiment of the present disclosure;

[0044] FIG8 is a schematic diagram of the structure of a voice interaction device of an intelligent device according to an embodiment of the present disclosure;

[0045] FIG9 is a schematic structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0046] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.

[0047] The following describes the voice interaction method, apparatus, vehicle, electronic device, and storage medium of the smart device according to the embodiments of the present disclosure in conjunction with the accompanying drawings.

[0048] FIG1 is a flow chart of a voice interaction method for a smart device according to an embodiment of the present disclosure.

[0049] As shown in FIG1 , the voice interaction method of the smart device according to an embodiment of the present disclosure includes:

[0050] A wake-up instruction is received, the smart device is controlled to enter a wake-up state, and the identity of the first user is extracted from source information of the wake-up instruction.

[0051] It should be noted that the voice interaction method for a smart device according to the embodiments of the present disclosure is performed by a smart device, including mobile phones, laptops, desktop computers, vehicle-mounted terminals, smart home appliances, etc. The voice interaction method for a smart device according to the embodiments of the present disclosure can be performed by a voice interaction device according to the embodiments of the present disclosure. The voice interaction device according to the embodiments of the present disclosure can be configured in any smart device to perform the voice interaction method for a smart device according to the embodiments of the present disclosure.

[0052] It should be noted that there are not too many restrictions on the wake-up instructions. For example, they can include voice instructions, gesture instructions, touch instructions, etc.

[0053] It should be noted that the first user refers to the user who issues the wake-up command, that is, the first user is the wake-up person, and can be any user. Taking the execution subject as an in-vehicle terminal as an example, the first user can be any passenger in the vehicle, such as the driver's seat, the front passenger seat, or the rear seat.

[0054] It is understood that the source information of the wake-up command carries the identity of the first user, and the identity of the first user can be extracted from the source information of the wake-up command. It is also understood that different users may correspond to different identity identifiers. For example, taking the execution subject as an in-vehicle terminal, the corresponding identity identifiers for the main driver and front passenger are 001 and 002 respectively.

[0055] S102: Collect voice information of the first user, perform semantic analysis on the voice information of the first user, and determine a voice interaction scenario of the smart device.

[0056] It should be noted that there are not too many restrictions on voice interaction scenarios. For example, it can include song ordering scenarios, vehicle control scenarios, navigation scenarios, chat scenarios, etc.

[0057] In one embodiment, semantic analysis is performed on the voice information of the first user to determine the voice interaction scenario of the smart device, including performing voice recognition on the voice information of the first user to obtain the recognition text corresponding to the first user, and judging whether the recognition text corresponding to the first user includes keywords corresponding to the candidate scenario. If the recognition text corresponding to the first user includes keywords corresponding to the candidate scenario, the voice interaction scenario of the smart device is determined to be a candidate scenario.

[0058] It should be noted that different candidate scenarios may correspond to different keywords.

[0059] For example, the song-ordering scene can correspond to keywords such as "sing", "song", and "listen"; the vehicle control scene can correspond to keywords such as "seat", "air conditioning", and "window"; the navigation scene can correspond to keywords such as "navigation", "go", and "head to"; and the chat scene can correspond to keywords such as "weather", "good morning", and "week".

[0060] In some examples, as shown in Figure 2, taking the execution subject as an in-vehicle terminal, the vehicle includes three rows of seats, each row includes two seats, that is, the vehicle includes six seats in total. The first row of seats includes the driver's seat and the passenger seat, the second row of seats includes the second row of left seats and the second row of right seats, and the third row of seats includes the third row of left seats and the third row of right seats. The vehicle includes six MIC (Microphone) modules, located in the area surrounding the six seats. The MIC modules correspond one to one with the seats and are used to collect voice information from the passengers in the corresponding seats.

[0061] In some examples, a wake-up command can be received to control the vehicle-mounted terminal to enter the wake-up state, and the identity of the main driver passenger can be extracted from the source information of the wake-up command, that is, the main driver passenger is the wake-up person, the voice information of the main driver passenger is collected, and voice recognition is performed on the voice information of the main driver passenger to obtain the corresponding recognition text of the main driver passenger.

[0062] If the recognition text corresponding to the main driver passenger includes "I want to sing xx's song", it can be determined that the recognition text corresponding to the main driver passenger includes the "sing" keyword corresponding to the song ordering scene, and the voice interaction scene of the in-vehicle terminal can be determined to be a song ordering scene.

[0063] If the recognition text corresponding to the main driver and passenger includes "turn on seat heating", it can be determined that the recognition text corresponding to the main driver and passenger includes the "seat" keyword corresponding to the vehicle control scene, and the voice interaction scene of the vehicle terminal can be determined to be the vehicle control scene.

[0064] If the recognition text corresponding to the main driver passenger includes "Navigate to a nearby restaurant", it can be determined that the recognition text corresponding to the main driver passenger includes the "navigation" keyword corresponding to the navigation scene, and the voice interaction scene of the vehicle terminal can be determined to be a navigation scene.

[0065] If the recognition text corresponding to the main driver passenger includes "How is the weather today", it can be determined that the recognition text corresponding to the main driver passenger includes the "weather" keyword corresponding to the chat scene, and the voice interaction scene of the in-vehicle terminal can be determined to be a chat scene.

[0066] In some examples, as shown in Figure 2, a vehicle includes a HUD (Head Up Display), a central control screen, a secondary screen, and a rear screen. The central control screen is located in the area in front of the driver's seat, the secondary screen is located in the area in front of the front passenger seat, and the rear screen is located between the first and second rows of seats.

[0067] S103, based on the correspondence between the voice interaction scenario and the voice interaction mode, obtain the voice interaction mode corresponding to the voice interaction scenario, and in response to the voice interaction mode being the multi-person mode, control the smart device to enter the multi-person mode.

[0068] In the embodiments of the present disclosure, the voice interaction modes of the smart device include multiplayer mode and single-player mode. There is a correspondence between voice interaction scenarios and voice interaction modes. This correspondence can be pre-set and is not specifically defined herein. For example, song request scenarios, vehicle control scenarios, and navigation scenarios all correspond to multiplayer mode, while chat scenarios correspond to single-player mode.

[0069] For example, based on the correspondence between the song-ordering scene and the multi-person mode, it can be determined that the voice interaction mode corresponding to the song-ordering scene is the multi-person mode, and the smart device can be controlled to enter the multi-person mode.

[0070] S104: Collect voice information of multiple users, and filter the voice information of the multiple users based on the voice information of the multiple users and the voice interaction scenarios to obtain voice information of at least one second user.

[0071] It should be noted that the multiple users may include the first user (the wake-up person) and / or the remaining users (non-wake-up persons) except the first user. The second user is any user among the multiple users.

[0072] In one embodiment, based on the voice information and voice interaction scenarios of multiple users, the voice information of multiple users is screened to obtain the voice information of at least one second user, including performing semantic analysis on the user's voice information to obtain the user's voice interaction intention, judging whether the user's voice interaction intention matches the voice interaction scenario, and determining the matching user's voice information as the second user's voice information.

[0073] Continuing with Figure 2 as an example, if the voice interaction scenario of the in-vehicle terminal is a song-ordering scenario, the main driver passenger can say "I want to sing the songs of singing star 1", the front passenger passenger can say "I want to sing the songs of singing star 2", and the second row left passenger can say "Navigate to xx". Each of the above voice information can be semantically analyzed separately, and it can be found that the voice interaction intentions of the main driver passenger and the front passenger are both song-ordering intentions, and the voice interaction intention of the second row left passenger is navigation intention. It is judged that the voice interaction intentions of the main driver passenger and the front passenger match the song-ordering scenario, and it is judged that the voice interaction intention of the second row left passenger does not match the song-ordering scenario, and the voice information of the main driver passenger and the front passenger are both determined as the voice information of the second user.

[0074] S105, perform semantic analysis on voice information of at least one second user, generate a first voice interaction instruction, and control the smart device to execute the first voice interaction instruction; wherein, the first user is a user who issues voice information for determining the voice interaction scenario of the smart device, and the second user is at least one user among the multiple users, and is a user who issues voice information for generating the first voice interaction instruction.

[0075] In one embodiment, semantic analysis is performed on the voice information of at least one second user to generate a first voice interaction instruction, including performing semantic analysis on the voice information of the second user to obtain the voice interaction intention of the second user, and generating the first voice interaction instruction based on the voice interaction intention of the second user.

[0076] Continuing with Figure 2 as an example, if the voice interaction scenario of the in-vehicle terminal is a song-ordering scenario, the main driver passenger can say "I want to sing songs by singing star 1" and the front passenger passenger can say "I want to sing songs by singing star 2". The voice information of the main driver passenger can be semantically analyzed to determine that the voice interaction intention of the main driver passenger is to order songs by singing star 1. Based on the above voice interaction intention of the main driver passenger, the first voice interaction instruction of "order songs by singing star 1" can be generated.

[0077] The voice information of the front passenger can also be semantically analyzed to obtain the front passenger's voice interaction intention of requesting a song by singing star 2. Based on the above voice interaction intention of the front passenger, the first voice interaction instruction of "requesting a song by singing star 2" can be generated.

[0078] In one embodiment, controlling the smart device to execute the first voice interaction instruction includes sending the first voice interaction instruction to a target application deployed in the smart device, and controlling the target application to execute the first voice interaction instruction.

[0079] In some examples, the target application corresponding to the voice interaction scenario can be obtained based on the correspondence between the voice interaction scenario and the target application. For example, there is a correspondence between the song request scenario and the song request application, the vehicle control scenario and the vehicle control application, the navigation scenario and the navigation application, and the chat scenario and the chat application.

[0080] For example, if the voice interaction scenario of the smart device is a song-ordering scenario, the target application is a song-ordering application. The first voice interaction instruction includes "Order the song of singing star 1" and "Order the song of singing star 2". "Order the song of singing star 1" and "Order the song of singing star 2" can be sent to the song-ordering application to control the song-ordering application to execute "Order the song of singing star 1" and "Order the song of singing star 2".

[0081] In summary, according to the voice interaction method of the smart device in the embodiment of the present disclosure, the correspondence between the voice interaction scenario and the voice interaction mode can be taken into consideration to determine the voice interaction scenario of the smart device, and when the voice interaction mode is the multi-person mode, the smart device can be controlled to enter the multi-person mode. The voice information and voice interaction scenarios of multiple users can be comprehensively considered, and the voice information of at least one second user can be screened out from the voice information of multiple users to generate a first voice interaction instruction, and the smart device can be controlled to execute the first voice interaction instruction, which helps to realize voice interaction of users who really have voice interaction intentions, improves the responsiveness of voice interaction, and can realize voice interaction between the smart device and at least one second user. It is particularly suitable for voice interaction scenarios of vehicle-mounted terminals, and can achieve full vehicle coverage of voice interaction of vehicle-mounted terminals.

[0082] FIG3 is a flow chart of a voice interaction method for a smart device according to another embodiment of the present disclosure.

[0083] As shown in FIG3 , the voice interaction method of the smart device according to an embodiment of the present disclosure includes:

[0084] A wake-up instruction is received, the smart device is controlled to enter a wake-up state, and the identity of the first user is extracted from source information of the wake-up instruction.

[0085] S302: Collect voice information of the first user, perform semantic analysis on the voice information of the first user, and determine a voice interaction scenario of the smart device.

[0086] S303: Based on the correspondence between the voice interaction scenario and the voice interaction mode, obtain the voice interaction mode corresponding to the voice interaction scenario, and in response to the voice interaction mode being the multi-person mode, control the smart device to enter the multi-person mode.

[0087] S304: Collect voice information of multiple users, and filter the voice information of the multiple users based on the voice information of the multiple users and the voice interaction scenarios to obtain voice information of at least one second user.

[0088] For the relevant contents of steps S302-S304, please refer to the above embodiment and will not be repeated here.

[0089] S305, perform semantic analysis on the voice information of the third user, obtain the third user's voice interaction intention to control the target category of vehicle-mounted equipment, and generate a first interaction instruction based on the target category of vehicle-mounted equipment and the identity of the third user, wherein the first interaction instruction carries the identity of the third user.

[0090] S306, perform semantic analysis on the fourth user's voice information, obtain the fourth user's voice interaction intention to control the target category of vehicle-mounted equipment, replace the third user's identity in the first interaction instruction with the fourth user's identity, and generate a second interaction instruction.

[0091] In an embodiment of the present disclosure, the second user includes a third user and a fourth user, and the third user and the fourth user are different users in the same space. The first voice interaction instruction includes a first interaction instruction and a second interaction instruction.

[0092] It should be noted that the voice interaction scenario in the embodiments of the present disclosure is a vehicle control scenario. There are no excessive restrictions on the target category of vehicle-mounted equipment, for example, it may include seats, air conditioners, windows, etc.

[0093] In one embodiment, semantic analysis is performed on the fourth user's voice information to determine that the fourth user's voice interaction intention is to control an in-vehicle device of a target category. This includes performing voice recognition on the fourth user's voice information to obtain recognized text corresponding to the fourth user, determining whether the recognized text corresponding to the fourth user includes a setting keyword, and if the recognized text corresponding to the fourth user includes the setting keyword, determining that the fourth user's voice interaction intention is to control an in-vehicle device of the target category. Thus, in this method, when the recognized text corresponding to the fourth user includes the setting keyword, the fourth user's voice interaction intention can be determined to be to control an in-vehicle device of the target category.

[0094] It should be noted that there are not too many restrictions on setting keywords. For example, the set keywords include "me too", "also", "same", "all", etc.

[0095] Continuing with Figure 2, the vehicle terminal's voice interaction scenario is vehicle control. Secondary users include the driver, front passenger, and second-row left passenger. The driver can say, "Turn on the seat heating," the front passenger can say, "I want some too," and the second-row left passenger can say, "I want some too."

[0096] The driver's seat (the third user) can be semantically analyzed to determine that the driver's seat's interaction intent is to control the seat. Based on the seat and the driver's seat's identity, a first interaction command is generated. The first interaction command carries the driver's seat's identity. For example, the first interaction command could be "Turn on the heating function for passenger seat 001," where 001 is the driver's seat's identity.

[0097] The second interaction command is generated by replacing the first interaction command with the second interaction command. For example, the first interaction command "Turn on the heating function of passenger seat 001" can be replaced with "002" to generate the second interaction command "Turn on the heating function of passenger seat 002", where "002" is the second interaction command identifier.

[0098] The second-row left passenger's (fourth user) voice message can be semantically analyzed to determine that the second-row left passenger's voice interaction intent is to control the seat. The driver's seat ID in the first interaction command is replaced with the second-row left passenger's ID to generate a second interaction command. For example, the first interaction command "Turn on the heating function of passenger seat 001" can be replaced with 003 to generate a second interaction command "Turn on the heating function of passenger seat 003," where 003 is the second-row left passenger's ID.

[0099] S307: Control the smart device to execute the first voice interaction instruction.

[0100] For the relevant content of step S307, please refer to the above embodiment and will not be repeated here.

[0101] In summary, according to the voice interaction method of the smart device in the embodiment of the present disclosure, a first interaction instruction can be generated based on the voice information of the third user, and when the voice interaction intention of the fourth user is to control the vehicle-mounted equipment of the target category, the identity identifier of the third user in the first interaction instruction is replaced with the identity identifier of the fourth user to generate a second interaction instruction, which improves the efficiency of instruction generation and is particularly suitable for vehicle control scenarios of vehicle-mounted terminals.

[0102] FIG4 is a flow chart of a voice interaction method for a smart device according to another embodiment of the present disclosure.

[0103] As shown in FIG4 , the voice interaction method of the smart device according to an embodiment of the present disclosure includes:

[0104] A wake-up instruction is received, the smart device is controlled to enter a wake-up state, and the identity of the first user is extracted from source information of the wake-up instruction.

[0105] S402: Collect voice information of the first user, perform semantic analysis on the voice information of the first user, and determine a voice interaction scenario of the smart device.

[0106] S403: Based on the correspondence between the voice interaction scenario and the voice interaction mode, obtain the voice interaction mode corresponding to the voice interaction scenario, and in response to the voice interaction mode being the multi-person mode, control the smart device to enter the multi-person mode.

[0107] S404: Collect voice information of multiple users, and filter the voice information of the multiple users based on the voice information of the multiple users and the voice interaction scenarios to obtain voice information of at least one second user.

[0108] For the relevant contents of steps S402-S404, please refer to the above embodiment and will not be repeated here.

[0109] S405: Perform semantic analysis on the voice information of each second user to obtain the voice interaction intention of each second user.

[0110] S406: Determine whether the voice interaction intentions of multiple second users are related.

[0111] S407: If the voice interaction intentions of multiple second users are related, the voice interaction intentions of the multiple second users are fused to obtain a fused interaction intention.

[0112] S408: Generate a first voice interaction instruction based on the fusion interaction intention.

[0113] Continuing with Figure 2 as an example, the voice interaction scenario of the in-vehicle terminal is a navigation scenario. The second user's voice information includes the voice information of the main driver and the front passenger. The main driver can say "Navigate to a nearby restaurant" and the front passenger can say "Go to the first one."

[0114] Semantic analysis can be performed on the driver's voice information to determine that the driver's voice interaction intention is to navigate to a nearby restaurant. Semantic analysis can be performed on the front passenger's voice information to determine that the front passenger's voice interaction intention is to navigate to the first restaurant. Therefore, the voice interaction intentions of both the driver and front passenger are to navigate to the restaurant. The correlation between the voice interaction intentions of the driver and front passenger can be determined.

[0115] The voice interaction intentions of the driver and front passenger are fused, and the fused interaction intention is to navigate to the first nearby restaurant. Based on the fused interaction intention, the first voice interaction instruction of "navigate to the first nearby restaurant" is generated.

[0116] S409: Control the smart device to execute the first voice interaction instruction.

[0117] For the relevant content of step S409, please refer to the above embodiment and will not be repeated here.

[0118] In summary, according to the voice interaction method of the intelligent device of the embodiment of the present disclosure, the voice information of the second user is semantically analyzed to obtain the voice interaction intention of the second user, and it is determined whether the voice interaction intentions of multiple second users are related. If the voice interaction intentions of multiple second users are related, the voice interaction intentions of the multiple second users are fused to obtain the fused interaction intention. Based on the fused interaction intention, a first voice interaction instruction is generated, which is particularly suitable for navigation scenarios of vehicle-mounted terminals.

[0119] FIG5 is a flow chart of a voice interaction method for a smart device according to another embodiment of the present disclosure.

[0120] As shown in FIG5 , the voice interaction method of the smart device according to an embodiment of the present disclosure includes:

[0121] A wake-up instruction is received, the smart device is controlled to enter a wake-up state, and the identity of the first user is extracted from source information of the wake-up instruction.

[0122] S502: Collect voice information of the first user, perform semantic analysis on the voice information of the first user, and determine a voice interaction scenario of the smart device.

[0123] S503: Based on the correspondence between the voice interaction scenario and the voice interaction mode, obtain the voice interaction mode corresponding to the voice interaction scenario, and in response to the voice interaction mode being the multi-person mode, control the smart device to enter the multi-person mode.

[0124] For the relevant contents of steps S502-S503, please refer to the above embodiment and will not be repeated here.

[0125] S504: Collect voice information of multiple users, perform voice recognition on the user's voice information, and obtain recognition text corresponding to the user.

[0126] S505: Acquire a template text library for the voice interaction scenario.

[0127] S506: Obtain the similarity between the recognition text corresponding to the user and the template text in the template text library.

[0128] S507: Determine the voice information of the user whose similarity is greater than or equal to the set threshold as the voice information of the second user.

[0129] It should be noted that a template text library can be pre-set for a voice interaction scenario to store template texts for the voice interaction scenario. Different voice interaction scenarios may correspond to different template text libraries.

[0130] For example, the template text library for the song-ordering scenario may include template texts such as "sing", "song", and "listen"; the template text library for the vehicle control scenario may include template texts such as "seat", "air conditioning", "window", and "me too"; the template text library for the navigation scenario may include template texts such as "navigation", "go", "head to", and "nearest"; the template text library for the chat scenario may include template texts such as "weather", "good morning", and "week".

[0131] For example, continuing with Figure 2, if the voice interaction scenario on the in-vehicle terminal is vehicle control, the driver can say, "Turn on the seat heating," the front passenger can say, "Me too," and the second-row left passenger can say, "What's the weather like today?" A template text library for vehicle control scenarios can be obtained, including template texts such as "seat," "air conditioning," "window," and "Me too."

[0132] The similarity between the recognition texts "open", "seat", and "heating" corresponding to the main driver passenger and the above-mentioned template text can be obtained, where the similarity between the recognition text "seat" and the template text "seat" is 100%. In response to the similarity 100% being greater than the set threshold 80%, the voice information of the main driver passenger can be determined as the voice information of the second user.

[0133] The similarity between the recognition text "I also" and "want" corresponding to the front passenger and the above-mentioned template text can be obtained, where the similarity between the recognition text "I also" and the template text "I also" is 100%. In response to the similarity 100% being greater than the set threshold 80%, the voice information of the front passenger can be determined as the voice information of the second user.

[0134] The similarity between the recognition texts "today", "weather", and "how" corresponding to the second row left passenger and the above-mentioned template text can be obtained. If the similarity between each of the above-mentioned recognition texts and the template text is less than the set threshold of 80%, the voice information of the second row left passenger will not be determined as the voice information of the second user.

[0135] S508: Perform semantic analysis on the voice information of at least one second user, generate a first voice interaction instruction, and control the smart device to execute the first voice interaction instruction.

[0136] For the relevant content of step S508, please refer to the above embodiment and will not be repeated here.

[0137] In summary, according to the voice interaction method of the smart device in the embodiment of the present disclosure, voice recognition is performed on the user's voice information to obtain the recognition text corresponding to the user, a template text library of the voice interaction scenario is obtained, the similarity between the recognition text corresponding to the user and the template text in the template text library is obtained, and the voice information of the user whose similarity is greater than or equal to the set threshold is determined as the voice information of the second user.

[0138] FIG6 is a flow chart of a voice interaction method for a smart device according to another embodiment of the present disclosure.

[0139] As shown in FIG6 , the voice interaction method of the smart device according to an embodiment of the present disclosure includes:

[0140] A wake-up instruction is received, the smart device is controlled to enter a wake-up state, and the identity of the first user is extracted from source information of the wake-up instruction.

[0141] S602: Collect voice information of the first user, perform semantic analysis on the voice information of the first user, and determine a voice interaction scenario of the smart device.

[0142] S603: Based on the correspondence between the voice interaction scenario and the voice interaction mode, obtain the voice interaction mode corresponding to the voice interaction scenario, and in response to the voice interaction mode being the multi-person mode, control the smart device to enter the multi-person mode.

[0143] S604: Collect voice information of multiple users, perform voice recognition on the user's voice information, and obtain recognition text corresponding to the user.

[0144] S605: Acquire a template text library for the voice interaction scenario.

[0145] S606: Obtain the similarity between the recognition text corresponding to the user and the template text in the template text library.

[0146] For the relevant contents of steps S602-S606, please refer to the above embodiment and will not be repeated here.

[0147] S607: Determine whether the similarity between the recognized text corresponding to the first user and the template text in the template text library is greater than or equal to a set threshold.

[0148] S608: If the similarity corresponding to the first user is less than a set threshold, control the smart device to switch from multi-player mode to single-player mode.

[0149] S609: Perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

[0150] In an embodiment of the present disclosure, the voice information of the plurality of users includes the voice information of the first user (the wake-up person).

[0151] Continuing with Figure 2 as an example, if the voice interaction scenario of the vehicle terminal is a vehicle control scenario, and if the driver-passenger is the first user, the driver-passenger can say "What's the weather like today" and obtain the template text library of the vehicle control scenario. The template text library includes template texts such as "seat", "air conditioning", "window", and "me too".

[0152] The similarity between the recognition texts "today", "weather", and "how" corresponding to the main driver passenger and the above-mentioned template text can be obtained, and it can be determined whether the similarity between the recognition text corresponding to the main driver passenger and the above-mentioned template text is greater than or equal to the set threshold of 80%. If the similarity between each recognition text corresponding to the main driver passenger and the template text is less than the set threshold of 80%, the vehicle-mounted terminal can be controlled to switch from multi-person mode to single-player mode.

[0153] The voice information of the driver and passenger is semantically analyzed to generate a second voice interaction instruction of "play today's weather with voice". The "play today's weather with voice" can be sent to the chat application to control the chat application to execute "play today's weather with voice".

[0154] In summary, according to the voice interaction method of the smart device in the embodiment of the present disclosure, it is determined whether the similarity between the recognition text corresponding to the first user and the template text in the template text library is greater than or equal to the set threshold. When the similarity corresponding to the first user is less than the set threshold, the smart device is controlled to switch from multi-person mode to single-player mode, and the voice information of the first user is semantically analyzed to generate a second voice interaction instruction, and the smart device is controlled to execute the second voice interaction instruction to realize voice interaction between the smart device and the awakened person, which can meet the voice interaction needs of the awakened person.

[0155] FIG7 is a flow chart of a voice interaction method for a smart device according to another embodiment of the present disclosure.

[0156] As shown in FIG7 , the voice interaction method of the smart device according to an embodiment of the present disclosure includes:

[0157] A wake-up instruction is received, the smart device is controlled to enter a wake-up state, and the identity of the first user is extracted from source information of the wake-up instruction.

[0158] S702: Collect voice information of the first user, perform semantic analysis on the voice information of the first user, and determine a voice interaction scenario of the smart device.

[0159] S703: Based on the correspondence between the voice interaction scenario and the voice interaction mode, obtain the voice interaction mode corresponding to the voice interaction scenario.

[0160] For the relevant contents of steps S702-S703, please refer to the above embodiment and will not be repeated here.

[0161] S704: In response to the voice interaction mode being the multi-person mode, determine whether a page of an application associated with the voice interaction scenario is displayed on the display interface of the smart device, and / or whether the smart device is in full-duplex mode.

[0162] S705: If a page of an application associated with the voice interaction scenario is displayed on the display interface of the smart device, and / or the smart device is in full-duplex mode, identify that the smart device meets the conditions for entering multiplayer mode, and control the smart device to enter multiplayer mode.

[0163] S706: If the page of the application associated with the voice interaction scenario is not displayed on the display interface of the smart device, and / or the smart device is in simplex mode, identify that the smart device does not meet the conditions for entering multiplayer mode, and control the smart device to enter single-player mode.

[0164] It is understandable that different voice interaction scenarios can be associated with different applications. For example, the song-ordering scenario can be associated with the song-ordering application, the vehicle-control scenario can be associated with the vehicle-control application, and the navigation scenario can be associated with the navigation application.

[0165] It should be noted that the display interface of the smart device can be at least one. For example, continuing with Figure 2 as an example, the display interface may include a central control screen, a secondary screen, a rear screen, etc.

[0166] It should be noted that smart devices have simplex mode and full-duplex mode. Simplex mode refers to responding to the voice information of one user (for example, only responding to the voice information of the wake-up person), while full-duplex mode refers to responding to the voice information of multiple users.

[0167] In some examples, continuing with Figure 2 as an example, if the voice interaction scenario of the vehicle terminal is a song ordering scenario, it can be determined whether the search page of the song ordering application is displayed on at least one of the central control screen, the secondary screen, and the rear screen, and / or whether the vehicle terminal is in full-duplex mode.

[0168] If the search page of the song request application is displayed on the central control screen and the vehicle terminal is in full-duplex mode, it is identified that the vehicle terminal meets the conditions for entering the multi-person mode and the vehicle terminal is controlled to enter the multi-person mode.

[0169] If the search page of the song request application is not displayed on the central control screen, and / or the vehicle terminal is in simplex mode, it is identified that the vehicle terminal does not meet the conditions for entering the multi-player mode, and the vehicle terminal is controlled to enter the single-player mode.

[0170] In some examples, continuing with FIG. 2 as an example, if the voice interaction scenario of the vehicle terminal is a navigation scenario, it can be determined whether the search page of the navigation application is displayed on the central control screen, and / or whether the vehicle terminal is in full-duplex mode.

[0171] If the search page of the navigation application is displayed on the central control screen and the vehicle terminal is in full-duplex mode, it is identified that the vehicle terminal meets the conditions for entering the multi-person mode and the vehicle terminal is controlled to enter the multi-person mode.

[0172] If the search page of the navigation application is not displayed on the central control screen, and / or the vehicle terminal is in simplex mode, it is identified that the vehicle terminal does not meet the conditions for entering multi-player mode, and the vehicle terminal is controlled to enter single-player mode.

[0173] S707: Collect voice information of multiple users, and filter the voice information of the multiple users based on the voice information of the multiple users and the voice interaction scenarios to obtain voice information of at least one second user.

[0174] S708: Perform semantic analysis on the voice information of at least one second user, generate a first voice interaction instruction, and control the smart device to execute the first voice interaction instruction.

[0175] S709: Perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

[0176] For the relevant contents of steps S707-S709, please refer to the above embodiment and will not be repeated here.

[0177] In summary, according to the voice interaction method of the smart device in the embodiment of the present disclosure, when the voice interaction mode corresponding to the voice interaction scenario is the multi-player mode, it can be determined whether the page of the application associated with the voice interaction scenario is displayed on the display interface of the smart device, and / or whether the smart device is in full-duplex mode. Based on the judgment result, it can be identified whether the smart device meets the conditions for entering the multi-player mode to control the smart device to enter the multi-player mode or the single-player mode.

[0178] Based on any of the above embodiments, after obtaining the voice interaction mode corresponding to the voice interaction scenario in step S103, it is also possible to control the smart device to enter the single-player mode in response to the voice interaction mode being the single-player mode, perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

[0179] For example, based on the correspondence between the chat scene and the single-player mode, it can be determined that the voice interaction mode corresponding to the chat scene is the single-player mode, and the smart device can be controlled to enter the single-player mode.

[0180] Therefore, in this method, when the voice interaction mode corresponding to the voice interaction scenario is the single-player mode, the smart device is directly controlled to enter the single-player mode, and the voice information of the first user is semantically analyzed to generate a second voice interaction instruction, and the smart device is controlled to execute the first voice interaction instruction to realize the voice interaction between the smart device and the awakened person, which can meet the voice interaction needs of the awakened person.

[0181] In order to implement the above embodiments, the present disclosure also proposes a voice interaction device for a smart device.

[0182] FIG8 is a schematic structural diagram of a voice interaction apparatus of an intelligent device according to an embodiment of the present disclosure.

[0183] As shown in FIG8 , the voice interaction apparatus 100 of the smart device according to the embodiment of the present disclosure includes: a wake-up module, a determination module 120 , a control module 130 , a screening module 140 and an execution module 150 .

[0184] The wake-up module is used to receive a wake-up instruction, control the smart device to enter a wake-up state, and extract the identity of the first user from the source information of the wake-up instruction;

[0185] The determination module 120 is configured to collect the voice information of the first user, perform semantic analysis on the voice information of the first user, and determine the voice interaction scenario of the smart device;

[0186] The control module 130 is configured to obtain a voice interaction mode corresponding to the voice interaction scenario based on a correspondence between the voice interaction scenario and the voice interaction mode, and control the smart device to enter the multi-person mode in response to the voice interaction mode being the multi-person mode;

[0187] The screening module 140 is configured to collect voice information of multiple users and screen the voice information of the multiple users based on the voice information of the multiple users and the voice interaction scenario to obtain voice information of at least one second user;

[0188] The execution module 150 is used to perform semantic analysis on the voice information of at least one second user, generate a first voice interaction instruction, and control the smart device to execute the first voice interaction instruction; wherein, the first user is the user who issues the voice information for determining the voice interaction scenario of the smart device, and the second user is at least one user among the multiple users, and is the user who issues the voice information for generating the first voice interaction instruction.

[0189] In one embodiment of the present disclosure, the second user includes a third user and a fourth user, and the third user and the fourth user are different users in the same space. The execution module 150 is also used to: perform semantic analysis on the voice information of the third user, obtain the voice interaction intention of the third user to control the vehicle-mounted equipment of the target category, and generate a first interaction instruction based on the vehicle-mounted equipment of the target category and the identity of the third user, wherein the first interaction instruction carries the identity of the third user; perform semantic analysis on the voice information of the fourth user, obtain the voice interaction intention of the fourth user to control the vehicle-mounted equipment of the target category, replace the identity of the third user in the first interaction instruction with the identity of the fourth user, and generate a second interaction instruction; wherein, the first voice interaction instruction includes the first interaction instruction and the second interaction instruction.

[0190] In one embodiment of the present disclosure, the execution module 150 is also used to: perform voice recognition on the voice information of the fourth user to obtain the recognition text corresponding to the fourth user; determine whether the recognition text corresponding to the fourth user includes set keywords; if the recognition text corresponding to the fourth user includes the set keywords, determine that the fourth user's voice interaction intention is to control the vehicle-mounted equipment of the target category.

[0191] In one embodiment of the present disclosure, the execution module 150 is also used to: perform semantic analysis on the voice information of each second user to obtain the voice interaction intention of each second user; determine whether the voice interaction intentions of multiple second users are related; if the voice interaction intentions of multiple second users are related, fuse the voice interaction intentions of multiple second users to obtain a fused interaction intention; and generate the first voice interaction instruction based on the fused interaction intention.

[0192] In one embodiment of the present disclosure, the screening module 140 is also used to: perform voice recognition on the voice information of the user to obtain the recognition text corresponding to the user; obtain the template text library of the voice interaction scenario; obtain the similarity between the recognition text corresponding to the user and the template text in the template text library; and determine the voice information of the user whose similarity is greater than or equal to a set threshold as the voice information of the second user.

[0193] In one embodiment of the present disclosure, the voice information of multiple users includes the voice information of the first user, and the execution module 150 is further used to: determine whether the similarity between the recognition text corresponding to the first user and the template text in the template text library is greater than or equal to the set threshold; if the similarity corresponding to the first user is less than the set threshold, control the smart device to switch from multi-player mode to single-player mode; perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

[0194] In one embodiment of the present disclosure, the control module 130 is further used to: determine whether a page of an application associated with the voice interaction scenario is displayed on the display interface of the smart device, and / or whether the smart device is in full-duplex mode; if a page of an application associated with the voice interaction scenario is displayed on the display interface of the smart device, and / or the smart device is in full-duplex mode, identify that the smart device meets the conditions for entering multiplayer mode, and control the smart device to enter multiplayer mode.

[0195] In one embodiment of the present disclosure, the execution module 150 is also used to: if the page of the application associated with the voice interaction scenario is not displayed on the display interface of the smart device, and / or the smart device is in simplex mode, identify that the smart device does not meet the conditions for entering multi-player mode, and control the smart device to enter single-player mode; perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

[0196] In one embodiment of the present disclosure, the execution module 150 is also used to: in response to the voice interaction mode being a single-player mode, control the smart device to enter the single-player mode; perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

[0197] It should be noted that for details not disclosed in the voice interaction apparatus of the smart device of the embodiment of the present disclosure, please refer to the details disclosed in the voice interaction method of the smart device of the embodiment of the present disclosure, which will not be repeated here.

[0198] In summary, the voice interaction device of the smart device in the embodiment of the present invention can take into account the correspondence between the voice interaction scenario and the voice interaction mode to determine the voice interaction scenario of the smart device, and when the voice interaction mode is the multi-person mode, control the smart device to enter the multi-person mode. It can comprehensively consider the voice information and voice interaction scenarios of multiple users, filter out the voice information of at least one second user from the voice information of multiple users to generate a first voice interaction instruction, and control the smart device to execute the first voice interaction instruction, which helps to realize voice interaction of users who really have voice interaction intentions, improves the responsiveness of voice interaction, and can realize voice interaction between the smart device and at least one second user. It is particularly suitable for voice interaction scenarios of vehicle-mounted terminals, and can achieve full vehicle coverage of voice interaction of vehicle-mounted terminals.

[0199] In order to implement the above embodiment, as shown in Figure 9, the embodiment of the present disclosure proposes an electronic device 200, including: a memory 210, a processor 220, and a computer program stored in the memory 210 and executable on the processor 220. When the processor 220 executes the program, the voice interaction method of the above-mentioned smart device is implemented.

[0200] The electronic device of the embodiment of the present disclosure executes a computer program stored in the memory through the processor, and can determine the voice interaction scenario of the smart device by considering the correspondence between the voice interaction scenario and the voice interaction mode, and control the smart device to enter the multi-person mode when the voice interaction mode is the multi-person mode. It can comprehensively consider the voice information and voice interaction scenarios of multiple users, filter out the voice information of at least one second user from the voice information of multiple users to generate a first voice interaction instruction, and control the smart device to execute the first voice interaction instruction, which helps to realize voice interaction of users who really have the intention of voice interaction, improves the responsiveness of voice interaction, and can realize voice interaction between the smart device and at least one second user. It is particularly suitable for the voice interaction scenario of the vehicle-mounted terminal, and can realize the full vehicle coverage of the voice interaction of the vehicle-mounted terminal.

[0201] In order to implement the above embodiment, the embodiment of the present disclosure proposes a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the voice interaction method of the above smart device is implemented.

[0202] The computer-readable storage medium of the embodiment of the present disclosure, by storing a computer program and being executed by a processor, can determine the voice interaction scenario of the smart device by considering the correspondence between the voice interaction scenario and the voice interaction mode, and control the smart device to enter the multi-person mode when the voice interaction mode is the multi-person mode. It can comprehensively consider the voice information and voice interaction scenarios of multiple users, filter out the voice information of at least one second user from the voice information of multiple users to generate a first voice interaction instruction, and control the smart device to execute the first voice interaction instruction, which helps to realize voice interaction of users who really have the intention of voice interaction, improves the responsiveness of voice interaction, and can realize voice interaction between the smart device and at least one second user. It is particularly suitable for the voice interaction scenario of the vehicle-mounted terminal, and can realize the full vehicle coverage of the voice interaction of the vehicle-mounted terminal.

[0203] In order to implement the above embodiment, the present disclosure provides a vehicle, including the voice interaction device of the above smart device; or the above electronic device; or the above computer-readable storage medium.

[0204] The vehicle of the embodiment of the present disclosure can determine the voice interaction scenario of the smart device by considering the correspondence between the voice interaction scenario and the voice interaction mode, and control the smart device to enter the multi-person mode when the voice interaction mode is the multi-person mode. It can comprehensively consider the voice information and voice interaction scenarios of multiple users, filter out the voice information of at least one second user from the voice information of multiple users to generate a first voice interaction instruction, and control the smart device to execute the first voice interaction instruction, which helps to realize voice interaction of users who really have the intention of voice interaction, improves the responsiveness of voice interaction, and can realize voice interaction between the smart device and at least one second user. It is particularly suitable for voice interaction scenarios of vehicle-mounted terminals, and can realize full vehicle coverage of voice interaction of vehicle-mounted terminals.

[0205] In the description of the present disclosure, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation to the present disclosure.

[0206] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0207] In this disclosure, unless otherwise expressly specified or limited, terms such as "mounted," "connected," "connect," and "fixed" should be understood broadly. For example, they may refer to fixed connections, detachable connections, or integration; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components or interactions between two components. Those skilled in the art will understand the specific meanings of these terms in this disclosure based on specific circumstances.

[0208] In the present disclosure, unless otherwise expressly specified or limited, when a first feature is "above" or "below" a second feature, it may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Furthermore, when a first feature is "above," "above," or "above" a second feature, it may mean that the first feature is directly above or diagonally above the second feature, or simply means that the first feature is at a higher level than the second feature. When a first feature is "below," "below," or "below" a second feature, it may mean that the first feature is directly below or diagonally below the second feature, or simply means that the first feature is at a lower level than the second feature.

[0209] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and features of different embodiments or examples, unless they are mutually inconsistent.

[0210] Although the embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are illustrative and are not to be construed as limitations on the present disclosure. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present disclosure.

Claims

1. A voice interaction method for an intelligent device, comprising: Collecting voice information of a first user, performing semantic analysis on the voice information of the first user, and determining a voice interaction scenario of the smart device; Based on the correspondence between the voice interaction scenario and the voice interaction mode, obtaining the voice interaction mode corresponding to the voice interaction scenario, and in response to the voice interaction mode being a multi-person mode, controlling the smart device to enter the multi-person mode; Collecting voice information of multiple users, and filtering the voice information of the multiple users based on the voice information of the multiple users and the voice interaction scenario to obtain voice information of at least one second user; Performing semantic analysis on the voice information of at least one second user to generate a first voice interaction instruction, and controlling the smart device to execute the first voice interaction instruction; Among them, the first user is a user who sends voice information for determining the voice interaction scenario of the smart device, and the second user is at least one user among the multiple users, and is a user who sends voice information for generating the first voice interaction instruction.

2. The method according to claim 1, wherein The second user includes a third user and a fourth user, the third user and the fourth user are different users in the same space, and the semantic analysis of the voice information of at least one of the second users to generate the first voice interaction instruction includes: performing semantic analysis on the voice information of the third user to determine that the third user's voice interaction intention is to control an in-vehicle device of a target category, and generating a first interaction instruction based on the in-vehicle device of the target category and the identity of the third user, wherein the first interaction instruction carries the identity of the third user; performing semantic analysis on the voice information of the fourth user to determine that the fourth user's voice interaction intention is to control the in-vehicle device of the target category, replacing the identity identifier of the third user in the first interaction instruction with the identity identifier of the fourth user, and generating a second interaction instruction; The first voice interaction instruction includes the first interaction instruction and the second interaction instruction.

3. The method according to claim 2, wherein: The performing semantic analysis on the voice information of the fourth user to obtain that the fourth user's voice interaction intention is to control the in-vehicle device of the target category includes: Performing voice recognition on the voice information of the fourth user to obtain a recognition text corresponding to the fourth user; determining whether the recognition text corresponding to the fourth user includes a set keyword; If the recognition text corresponding to the fourth user includes the set keyword, it is determined that the voice interaction intention of the fourth user is to control the in-vehicle device of the target category.

4. The method according to claim 1, wherein The performing semantic analysis on the voice information of at least one second user to generate a first voice interaction instruction includes: Performing semantic analysis on the voice information of each second user to obtain the voice interaction intention of each second user; Determining whether the voice interaction intentions of the plurality of second users are related; If the voice interaction intentions of multiple second users are related, fusing the voice interaction intentions of the multiple second users to obtain a fused interaction intention; Based on the fusion interaction intention, the first voice interaction instruction is generated.

5. The method according to claim 1, wherein The filtering of the voice information of the plurality of users based on the voice information of the plurality of users and the voice interaction scenario to obtain the voice information of at least one second user includes: Performing voice recognition on the user's voice information to obtain a recognition text corresponding to the user; Obtaining a template text library for the voice interaction scenario; Obtaining the similarity between the recognition text corresponding to the user and the template text in the template text library; The voice information of the user whose similarity is greater than or equal to a set threshold is determined as the voice information of the second user.

6. The method according to claim 5, wherein: The voice information of the plurality of users includes the voice information of the first user, and the method further includes: Determining whether the similarity between the recognition text corresponding to the first user and the template text in the template text library is greater than or equal to the set threshold; If the similarity corresponding to the first user is less than the set threshold, controlling the smart device to switch from multi-player mode to single-player mode; Perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

7. The method according to any one of claims 1 to 6, wherein The controlling the smart device to enter the multi-person mode includes: Determining whether a page of an application associated with the voice interaction scenario is displayed on a display interface of the smart device, and / or whether the smart device is in full-duplex mode; If the page of the application associated with the voice interaction scenario is displayed on the display interface of the smart device, and / or the smart device is in full-duplex mode, it is identified that the smart device meets the conditions for entering multi-person mode, and the smart device is controlled to enter multi-person mode.

8. The method according to claim 7, further comprising: If a page of an application associated with the voice interaction scenario is not displayed on the display interface of the smart device, and / or the smart device is in simplex mode, identifying that the smart device does not meet the conditions for entering multiplayer mode, and controlling the smart device to enter single-player mode; Perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

9. The method according to any one of claims 1 to 6, further comprising: In response to the voice interaction mode being the single-player mode, controlling the smart device to enter the single-player mode; Perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

10. A voice interaction device for an intelligent device, comprising: a determination module, configured to collect voice information of a first user, perform semantic analysis on the voice information of the first user, and determine a voice interaction scenario of the smart device; A control module is used to obtain the voice based on the corresponding relationship between the voice interaction scene and the voice interaction mode. a voice interaction mode corresponding to the interaction scenario, and in response to the voice interaction mode being a multi-person mode, controlling the smart device to enter the multi-person mode; a screening module, configured to collect voice information of a plurality of users, and screen the voice information of the plurality of users based on the voice information of the plurality of users and the voice interaction scenario, to obtain voice information of at least one second user; an execution module, configured to perform semantic analysis on the voice information of at least one second user, generate a first voice interaction instruction, and control the smart device to execute the first voice interaction instruction; Among them, the first user is a user who sends voice information for determining the voice interaction scenario of the smart device, and the second user is at least one user among the multiple users, and is a user who sends voice information for generating the first voice interaction instruction.

11. The device according to claim 10, wherein The second user includes a third user and a fourth user, the third user and the fourth user are different users in the same space, and the semantic analysis of the voice information of at least one of the second users to generate the first voice interaction instruction includes: performing semantic analysis on the voice information of the third user to determine that the third user's voice interaction intention is to control an in-vehicle device of a target category, and generating a first interaction instruction based on the in-vehicle device of the target category and the identity of the third user, wherein the first interaction instruction carries the identity of the third user; performing semantic analysis on the voice information of the fourth user to determine that the fourth user's voice interaction intention is to control the in-vehicle device of the target category, replacing the identity identifier of the third user in the first interaction instruction with the identity identifier of the fourth user, and generating a second interaction instruction; The first voice interaction instruction includes the first interaction instruction and the second interaction instruction.

12. The device according to claim 11, wherein The execution module is further configured to: Performing voice recognition on the voice information of the fourth user to obtain a recognition text corresponding to the fourth user; determining whether the recognition text corresponding to the fourth user includes a set keyword; If the recognition text corresponding to the fourth user includes the set keyword, it is determined that the voice interaction intention of the fourth user is to control the in-vehicle device of the target category.

13. The device according to claim 10, wherein The execution module is further configured to: Performing semantic analysis on the voice information of each second user to obtain the voice interaction intention of each second user; Determining whether the voice interaction intentions of the plurality of second users are related; If the voice interaction intentions of multiple second users are related, fusing the voice interaction intentions of the multiple second users to obtain a fused interaction intention; Based on the fusion interaction intention, the first voice interaction instruction is generated.

14. The device according to claim 10, wherein The screening module is also used to: Performing voice recognition on the user's voice information to obtain a recognition text corresponding to the user; Obtaining a template text library for the voice interaction scenario; Obtaining the similarity between the recognition text corresponding to the user and the template text in the template text library; The voice information of the user whose similarity is greater than or equal to the set threshold is determined as the voice information of the second user. interest.

15. The device according to claim 14, wherein The voice information of the plurality of users includes the voice information of the first user, and the execution module is further configured to: Determining whether the similarity between the recognition text corresponding to the first user and the template text in the template text library is greater than or equal to the set threshold; If the similarity corresponding to the first user is less than the set threshold, controlling the smart device to switch from multi-player mode to single-player mode; Perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

16. The device according to any one of claims 10 to 15, wherein: The control module is further configured to: Determining whether a page of an application associated with the voice interaction scenario is displayed on a display interface of the smart device, and / or whether the smart device is in full-duplex mode; If the page of the application associated with the voice interaction scenario is displayed on the display interface of the smart device, and / or the smart device is in full-duplex mode, it is identified that the smart device meets the conditions for entering multi-person mode, and the smart device is controlled to enter multi-person mode.

17. The device according to claim 16, wherein The execution module is further configured to: If a page of an application associated with the voice interaction scenario is not displayed on the display interface of the smart device, and / or the smart device is in simplex mode, identifying that the smart device does not meet the conditions for entering multiplayer mode, and controlling the smart device to enter single-player mode; Perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

18. The device according to any one of claims 10 to 15, wherein: The execution module is further configured to: In response to the voice interaction mode being the single-player mode, controlling the smart device to enter the single-player mode; Perform semantic analysis on the voice information of the first user, generate a second voice interaction instruction, and control the smart device to execute the second voice interaction instruction.

19. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the voice interaction method for an intelligent device according to any one of claims 1 to 9 is implemented.

20. A computer-readable storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the voice interaction method of the smart device according to any one of claims 1 to 9 is implemented.

21. A vehicle comprising: The voice interaction device of the smart device according to claim 10; or the electronic device according to claim 19; Or the computer-readable storage medium of claim 20.

Citation Information

Cited By

  • Automobile intelligent cockpit voice interaction method and system based on swan monk system, and medium

    CN121191515A