Voice interaction method, device, electronic device, and storage medium

By determining the current environment information in the voice assistant, distinguishing different voice requests, and outputting corresponding response information, the problem of voice assistants being unable to meet the needs of multiple users at the same time is solved, thus improving user experience and driving safety.

CN114495932BActive Publication Date: 2025-10-28BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210169073.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-23
Publication Date
2025-10-28
Estimated Expiration
2042-02-23

AI Technical Summary

Technical Problem

The voice assistant cannot distinguish between different voice input sources, which means that it cannot meet the needs of multiple users at the same time during the broadcast, affecting driving safety and user experience.

Method used

By determining the current environment information, different voice requests are distinguished, and the first response information corresponding to the first voice request and the second response information corresponding to the second voice request are output, so as to meet the needs of multiple users at the same time.

Benefits of technology

It improves the intelligence and user-friendliness of the voice assistant, reduces the impact of children on driving, and enhances user experience and driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495932B_ABST
    Figure CN114495932B_ABST
Patent Text Reader

Abstract

This disclosure provides a voice interaction method, apparatus, electronic device, and storage medium, relating to the field of computer technology, and particularly to the fields of the Internet of Things, voice technology, and artificial intelligence. The specific implementation scheme is as follows: in response to receiving a second voice request during the output of a first response information corresponding to a first voice request, determining current environmental information; determining second response information corresponding to the second voice request based on the current environmental information; and outputting the first and second response information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to the fields of Internet of Things, voice technology, and artificial intelligence. Specifically, it relates to a voice interaction method, device, electronic device, and storage medium. Background Technology

[0002] Voice interaction is a new generation of interaction mode based on voice input. It can recognize voice information and provide feedback on the interaction results. Voice interaction is applied in various voice assistants. A voice assistant is an intelligent software application that can be installed on various terminals such as mobile phones, in-vehicle systems, computers, and other electronic devices. Voice assistants enable intelligent dialogue and instant question-and-answer interactions to help users solve a range of problems. Summary of the Invention

[0003] This disclosure provides a voice interaction method, apparatus, electronic device, and storage medium.

[0004] According to one aspect of this disclosure, a voice interaction method is provided, comprising: receiving a second voice request in response to outputting first response information corresponding to a first voice request, determining current environment information; determining second response information corresponding to the second voice request based on the current environment information; and outputting the first response information and the second response information.

[0005] According to another aspect of this disclosure, a voice interaction device is provided, comprising: a first determining module, configured to receive a second voice request and determine current environment information in response to outputting a first response information corresponding to a first voice request; a second determining module, configured to determine a second response information corresponding to the second voice request based on the current environment information; and an output module, configured to output the first response information and the second response information.

[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the voice interaction method of this disclosure.

[0007] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the voice interaction method of this disclosure.

[0008] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the voice interaction method of this disclosure.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0011] Figure 1 This illustration schematically shows an exemplary system architecture to which voice interaction methods and devices can be applied according to embodiments of the present disclosure;

[0012] Figure 2 A flowchart illustrating a voice interaction method according to an embodiment of the present disclosure is shown schematically.

[0013] Figure 3 A flowchart illustrating the application of the voice interaction method according to this disclosure in a navigation voice assistant is shown.

[0014] Figure 4 A block diagram of a voice interaction device according to an embodiment of the present disclosure is schematically shown; and

[0015] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0016] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0017] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.

[0018] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0019] Voice interaction involves the user inputting a wake-up word into a voice interaction device to activate the voice assistant within the device, followed by a dialogue between the voice assistant and the user. When the user activates the voice assistant, they can input voice information related to their needs. After recognizing the voice information, the voice assistant can perform feature extraction and analysis, and then output a recognition result relevant to the user's needs.

[0020] In real-world scenarios, voice assistants may receive multiple different voice messages simultaneously or within a short period. For example, when a user is driving using map navigation, they need to listen to voice navigation or interact with the map navigation through a voice assistant. If there are multiple users in the cockpit, including children, the driver needs to be distracted while driving and interacting with the children, which can easily affect driving safety. With the increasing prevalence of smart products such as smart speakers in homes, children are also more familiar with using smart products and intelligent dialogue. If children interact with the voice assistant while the map navigation is broadcasting navigation information, it will interfere with normal voice navigation and affect the driver's driving.

[0021] In realizing this disclosed concept, the inventors discovered that voice assistants cannot distinguish between different voice input sources. During background voice broadcasting, various other voice input sources can wake up the voice assistant. After determining the user's needs based on the recognized other voice input sources, the voice assistant will interrupt the ongoing broadcasting and provide new recognition results. Furthermore, voice assistants cannot simultaneously satisfy different input requests from the same user. For example, a navigation voice assistant cannot simultaneously fulfill two requests: one to play a song while broadcasting navigation instructions.

[0022] This disclosure provides a voice interaction method, apparatus, electronic device, and storage medium. The voice interaction method includes: receiving a second voice request in response to outputting first response information corresponding to a first voice request; determining current environment information; determining second response information corresponding to the second voice request based on the current environment information; and outputting the first response information and the second response information.

[0023] Figure 1 The illustration schematically depicts an exemplary system architecture to which voice interaction methods and devices can be applied according to embodiments of the present disclosure.

[0024] It is important to note that Figure 1The examples shown are merely examples of system architectures applicable to embodiments of this disclosure, intended to help those skilled in the art understand the technical content of this disclosure. They do not imply that embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For instance, in another embodiment, an exemplary system architecture to which the voice interaction methods and apparatus can be applied may include a terminal device. However, the terminal device may implement the voice interaction methods and apparatus provided in the embodiments of this disclosure without needing to interact with a server.

[0025] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0026] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).

[0027] Terminal devices 101, 102, and 103 can be various electronic devices with voice interaction capabilities, including but not limited to smartphones, tablets, laptops, desktop computers, and smart speakers, etc. Smart speakers can include in-vehicle smart speakers, etc.

[0028] Server 105 can be a server providing various services, such as a backend management server supporting the content browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated based on user requests) to the terminal devices. The server can be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service system, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") in terms of management difficulty and weak business scalability. The server can also be a server for a distributed system or a server integrated with blockchain technology.

[0029] It should be noted that the voice interaction method provided in this embodiment can generally be executed by terminal devices 101, 102, or 103. Correspondingly, the voice interaction device provided in this embodiment can also be disposed in terminal devices 101, 102, or 103.

[0030] Alternatively, the voice interaction method provided in this embodiment can generally be executed by server 105. Correspondingly, the voice interaction device provided in this embodiment can generally be located in server 105. The voice interaction method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the voice interaction device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0031] For example, when a user is reading an ebook online, terminal devices 101, 102, and 103 can acquire the target content in the ebook that the user is looking at, and then send the acquired target content to server 105. Server 105 analyzes the target content to determine its feature information; predicts content that the user is interested in based on the feature information; and extracts the content that the user is interested in. Alternatively, a server or server cluster capable of communicating with terminal devices 101, 102, and 103 and / or server 105 can analyze the target content and ultimately extract the content that the user is interested in.

[0032] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0033] Figure 2 A flowchart illustrating a voice interaction method according to an embodiment of the present disclosure is shown schematically.

[0034] like Figure 2 As shown, the method includes operations S210 to S230.

[0035] During operation S210, in response to receiving a second voice request while outputting first response information corresponding to the first voice request, the current environment information is determined.

[0036] In operation S220, based on the current environment information, the second response information corresponding to the second voice request is determined.

[0037] In operation S230, the first response information and the second response information are output.

[0038] According to embodiments of this disclosure, a first voice request may include a voice request output by a first user. A second voice request may include a voice request output by a second user. The first user and the second user may be the same user. The number of second users may be one or more. The first user may include adult users. The second user may include at least one of adult users and child users. The first response information may characterize response information related to the content requested by the first voice request. Current environmental information may include at least one of driving status information, driving route information, and driving road condition information, and may not be limited to these.

[0039] According to embodiments of this disclosure, in response to outputting first response information corresponding to the first voice request, the responder receiving the second voice request may include various types of voice assistants. For example, the voice assistant may include at least one of the following: a voice assistant in a vehicle infotainment system, a voice assistant in map software, and a voice assistant in a vehicle navigation system, and is not limited thereto.

[0040] According to embodiments of this disclosure, the second response information may include at least one of the following: target information related to the content requested by the second voice request, predefined information unrelated to the content requested by the second voice request, and empty information. If the second response information is empty, it indicates that the voice assistant has not responded to the second voice request.

[0041] According to embodiments of this disclosure, the method of outputting the first response information and the second response information may include: when the first response information is continuous or interrupted output information, upon receiving a second voice request and determining the second response information, the output of the first response information may be interrupted, and the second response information may be output; and the output of the first response information may continue after the second response information is completed. When the first response information is interrupted output information, the second response information may be output during the output interval of the first response information. For example, upon receiving a second voice request and determining the second response information, the second response information may be output during the interval between the completion of the segment being output in the first response information; and the output of the second response information may be interrupted and the first response information output when it is necessary to output the first response information. In practical applications, the limitation on the output method is not limited to this.

[0042] According to embodiments of this disclosure, for example, after receiving a first voice request from an adult user, the voice assistant can output first response information related to the content requested by the first voice request. During the output of the first response information, for example, if a child user wakes up the voice assistant and inputs a second voice request, the voice assistant can determine second response information for responding to the second voice request based on the environmental information corresponding to the time or time period in which the second voice request was received, and then output the second response information.

[0043] Through the above embodiments of this disclosure, by combining the analysis of environmental information and the information represented by each voice request, the voice assistant can respond appropriately to each voice request when multiple voice requests are received, thus meeting the needs of each user. Furthermore, since the output response information can still include both the first and second response information after receiving a second voice request, the output of the first response information is not interrupted by the output of the second response information, ensuring complete output of all response information and improving the user experience.

[0044] The following describes specific embodiments. Figure 2 The method shown will be further explained.

[0045] As users increasingly utilize voice assistants across various scenarios and with greater frequency, they have higher expectations for the functionality and intelligence of these assistants. For instance, with the widespread use of vehicles, family travel has become a significant mode of transportation, making the resolution of driving safety issues during family trips or trips with children a crucial user need.

[0046] According to embodiments of this disclosure, the second voice request includes a voice request related to a child. The aforementioned voice interaction method may further include: in response to receiving a second voice request during the output of first response information, adding child-related target identification information to audio information related to the second voice request; performing feature extraction on the audio information including the target identification information to obtain a feature extraction result, so as to determine target information related to the content requested by the second voice request based on the feature extraction result.

[0047] According to embodiments of this disclosure, when audio information with child-like voice source features and child-like grammatical features is identified, the audio information can be determined to be child-like audio information. During the voice assistant's output of response information, if child-like audio information is received, target identifier information representing a child user type can be added to the child-like audio information to mark it as child-like audio information. Adult audio information can be left unprocessed. When the voice assistant receives both child-like and adult audio information simultaneously, the child-like and adult audio information can be processed according to different requirements.

[0048] It should be noted that if multiple adult audio messages with different remaining characteristics are received at the same time, these multiple adult audio messages can be processed into the same or different requirements based on the differences in sound source characteristics and the differences in requested content.

[0049] According to embodiments of this disclosure, the target identification information includes child-related identification information. When child-related identification information is detected in audio information, features can be extracted based on the child's speech characteristics using a feature extraction method that matches those characteristics. For example, based on features such as disjointed speech and unclear pronunciation, feature extraction can be achieved by extracting word information from the child-related audio information, identifying other similar information with a correlation greater than a first threshold, and identifying other possible word information with a word similarity and speech similarity greater than a second threshold.

[0050] It should be noted that for adult audio information, feature extraction can be achieved by extracting complete sentences, semantic features, and word features within sentences. The method for feature extraction of adult audio information can also be the same as that for children's audio information. Audio information that does not contain children's identifiers can be considered adult audio information and extracted using adult feature extraction methods.

[0051] Through the above embodiments of this disclosure, a children's mode can be added to the voice assistant, so as to not only meet the user's basic needs, but also meet the user's need to use the voice assistant to accompany the child.

[0052] According to embodiments of this disclosure, setting a child mode in a navigation voice assistant can meet the needs of children while allowing users to use map navigation functions normally, in scenarios such as family travel or travel with children.

[0053] Through the above embodiments of this disclosure, while ensuring driving safety, users' perception of the intelligence and humanization of the voice assistant can be increased, improving the user experience of navigation and thus enhancing product competitiveness.

[0054] Voice assistants are an important way for users to interact with navigation products. The level of intelligence and user-friendliness of voice assistants also affects users' judgment of the intelligence and user-friendliness of navigation products. Therefore, improving the intelligence and user-friendliness of voice assistants is an important evaluation indicator for intelligent navigation products.

[0055] According to embodiments of this disclosure, the current environment information may include driving state information. Based on the current environment information, determining the second response information corresponding to the second voice request may include: in response to detecting that the driving state information indicates the current driving state is stationary, determining the second response information as target information related to the content requested by the second voice request; and in response to detecting that the driving state information indicates the current driving state is in motion, determining the second response information as predefined information unrelated to the content requested by the second voice request.

[0056] According to embodiments of this disclosure, driving status information can be determined based on onboard sensors. This determined driving status information can be recorded in real-time in navigation products and can indicate whether the user is driving or not. A stationary state can indicate that the user is not driving. Target information may include the content requested by the second voice request. A moving state can indicate that the user is driving. Predefined information may include pre-set reassuring messages, such as "Little friend, we are currently driving, we will reply to you later," etc.

[0057] For example, when a navigation voice assistant identifies child-related audio messages while providing navigation information, it can first determine the current driving status. Then, if the current state is determined to be stationary, the child-related audio messages can be preprocessed, features extracted, and similarity measured using a product-specific corpus to determine the appropriate response—the aforementioned target information—which is then output through the voice assistant. If the current state is determined to be in motion, reassuring information—the aforementioned predefined information—can be determined and output through the voice assistant.

[0058] Through the embodiments disclosed above, a scheme can be implemented to prioritize and broadcast corresponding content based on current driving status information and voice input information. This scheme can effectively utilize the capabilities of voice assistants, enhance users' perception of the intelligence and humanization of voice assistants, and improve users' experience and satisfaction with map products. Especially for family travel, the voice assistant can provide effective driving information for drivers and also meet the needs of children, reducing the impact of children on driving.

[0059] According to embodiments of this disclosure, the first response information may include navigation information. Current environment information may include the driving route information represented by the navigation information. Based on the current environment information, determining the second response information corresponding to the second voice request may include: determining the historical number of trips of the current driving route represented by the driving route information. In response to detecting that the historical number of trips is greater than a first preset value, determining that the second response information is target information related to the content requested by the second voice request. In response to detecting that the historical number of trips is less than or equal to the first preset value, determining that the second response information is predefined information unrelated to the content requested by the second voice request.

[0060] According to embodiments of this disclosure, the driving route information may include full route information determined based on current navigation information, and partial route information to be traversed within the current time period determined based on current location information and full route information. The historical driving count can be determined based on historical records used to record historical trips in navigation products, and can indicate whether the current driving route is a familiar or unfamiliar route to the user. If the historical driving count is greater than a first preset value, the current driving route can be determined to be a familiar route to the user. If the historical driving count is less than or equal to the first preset value, the current driving route can be determined to be an unfamiliar route to the user. The first preset value can be determined based on the user's sensitivity to the route and the complexity of the route. Sensitivity can be determined through user settings. Complexity can be determined through information such as the number of intersections and turns.

[0061] For example, when a navigation voice assistant identifies child-related audio messages while providing navigation information, it can first determine the current driving route. Then, if the route is familiar to the user, the child-related audio messages can be preprocessed, features extracted, and similarity measured using a product-specific corpus to determine the appropriate response—the aforementioned target information—which is then output through the voice assistant. If the route is unfamiliar to the user, reassuring messages—the aforementioned predefined information—can be determined and output through the voice assistant.

[0062] The embodiments described above provide a scheme for prioritizing and broadcasting content based on current driving route information and voice input information. This scheme effectively utilizes the capabilities of voice assistants, enhancing users' perception of the intelligence and user-friendliness of the voice assistant, and improving user experience and satisfaction with map products. Especially for family travel, the voice assistant can provide effective driving information to the driver while also meeting the needs of children, reducing the impact of children on driving.

[0063] According to embodiments of this disclosure, the first response information may include navigation information. Current environment information may include traffic condition information related to the current driving route, and the current driving route may include the route represented by the navigation information. Determining the second response information corresponding to the second voice request based on the current environment information may include: determining at least one of the number of traffic lights, the number of turns, and the number of cameras in the current driving route based on the traffic condition information. In response to detecting that at least one of the number of traffic lights, the number of turns, and the number of cameras is greater than a second preset value, the second response information is determined to be target information related to the content requested by the second voice request. In response to detecting that at least one of the number of traffic lights, the number of turns, and the number of cameras is less than or equal to the second preset value, the second response information is determined to be predefined information unrelated to the content requested by the second voice request.

[0064] According to embodiments of this disclosure, the traffic condition information may further include at least one of the following: traffic flow information for the corresponding road segment, information that may require route changes due to unforeseen events. The second preset value can be determined based on the number of traffic lights, the number of turns, and the number of cameras, etc. The first preset value and the second preset value may be the same or different.

[0065] For example, when a navigation voice assistant identifies child-related audio messages while broadcasting navigation information, it can first determine the current road conditions. Then, if the road conditions are relatively simple (e.g., few traffic lights, cameras, or turns; sparse traffic; no emergencies), the child-related audio messages can be preprocessed, features extracted, and similarity measured using a product-specific corpus to determine the appropriate response—the target information—which is then output through the voice assistant. Conversely, if the road conditions are more complex (e.g., numerous traffic lights, cameras, or turns; complex traffic; emergencies), reassuring information—the predefined information—can be determined and output through the voice assistant.

[0066] Through the embodiments disclosed above, priority analysis can be performed based on current road conditions and voice input information to determine the appropriate content playback scheme. This scheme can effectively utilize the capabilities of the voice assistant, enhance users' perception of the voice assistant's intelligence and user-friendliness, and improve users' experience and satisfaction with map products. Especially for family travel, the voice assistant can provide effective driving information for the driver while also meeting the needs of children, reducing the impact of children on driving.

[0067] According to embodiments of this disclosure, the current environmental information may include at least one of the following: driving status information, driving route information, and driving road condition information. By identifying various types of environmental information, corresponding content can be output.

[0068] According to embodiments of this disclosure, if at least one of the following conditions is met—that the current time is a stationary state, the current driving route is a familiar route to the user, and the current road conditions are relatively simple—and target information needs to be output via the voice assistant, then the necessary navigation information can still be output first. Then, the target information is output after the navigation information output is completed.

[0069] For example, when the user is driving, the route is familiar, and the road conditions are not complex, the navigation announcements can be simplified, only providing necessary information such as turns, prioritizing the needs of children, such as displaying destination information. When the user is driving but the route is unfamiliar and the road conditions are complex, the navigation announcements can be prioritized, downplaying the needs of children and reassuring them, such as displaying predefined information. When the user is not driving, the needs of children can be prioritized, such as displaying destination information.

[0070] According to embodiments of this disclosure, the second voice request may include a voice request related to a child. Outputting first response information and second response information may include: determining a response role related to the child; and outputting the second response information based on the audio information of the response role.

[0071] According to embodiments of this disclosure, a voice pack specifically designed for children can be installed in the voice assistant, and the child-related response role can be determined based on the roles included in the voice pack. The determination method may include at least one of determination based on a person's settings or determination through manual selection. For the output content of adult voice requests and child-related voice requests, corresponding voices can be used for playback based on different voice packs.

[0072] For example, for voice requests from children, voice packs specific to children, such as those featuring cartoon characters that children like, can be used to read the content.

[0073] Through the above embodiments of this disclosure, feedback on the needs of child users can be provided using children's voices, and audio output results can be output for children after analyzing the children's voice input content, thereby improving the user experience.

[0074] Figure 3 A flowchart illustrating the application of the voice interaction method according to this disclosure in a navigation voice assistant is shown.

[0075] like Figure 3 As shown, the method includes operations S310 to S390.

[0076] While operating the S310, during the navigation voice assistant's broadcast of navigation information, a voice request was received from a child user.

[0077] In operation S320, determine the driving status information corresponding to the moment when a voice request from a child user is received.

[0078] In operation S330, is the current driving state a stationary state? If yes, then execute operation S380; if no, then execute operation S390.

[0079] In operation S340, determine the driving route information corresponding to the moment when the voice request from a child user is received.

[0080] In operation S350, is the historical number of trips for the current route greater than the first preset value? If yes, then execute operation S380; if no, then execute operation S390.

[0081] When operating the S360, determine the driving traffic information corresponding to the moment when a voice request from a child user is received.

[0082] When operating S370, is the current road condition complex? If yes, then execute operation S390; if no, then execute operation S380.

[0083] In operation S380, target information related to the content requested by the second voice request is output.

[0084] In operation S390, predefined information unrelated to the content requested by the second voice request is output.

[0085] Through the above embodiments of this disclosure, the implemented voice assistant can fulfill other user needs beyond basic map functions such as reporting traffic conditions and navigation. For example, it can broadcast more intelligent and user-friendly content to help relieve driving fatigue or anxiety and facilitate better driving.

[0086] Figure 4 A block diagram of a voice interaction device according to an embodiment of the present disclosure is shown schematically.

[0087] like Figure 4 As shown, the voice interaction device 400 includes a first determining module 410, a second determining module 420, and an output module 430.

[0088] The first determining module 410 is used to determine the current environment information in response to receiving a second voice request during the process of outputting a first response information corresponding to the first voice request.

[0089] The second determining module 420 is used to determine the second response information corresponding to the second voice request based on the current environment information.

[0090] Output module 430 is used to output the first response information and the second response information.

[0091] According to embodiments of this disclosure, the current environmental information includes driving status information. The second determining module includes a first determining unit and a second determining unit.

[0092] The first determining unit is configured to determine, in response to detecting that the driving state information indicates that the current driving state is stationary, the second response information is target information related to the content requested by the second voice request.

[0093] The second determining unit is configured to determine, in response to detecting that the driving state information indicates the current driving state is in motion, the second response information is predefined information unrelated to the content requested by the second voice request.

[0094] According to embodiments of this disclosure, the first response information includes navigation information. Current environment information includes the driving route information represented by the navigation information. The second determining module includes a third determining unit, a fourth determining unit, and a fifth determining unit.

[0095] The third determining unit is used to determine the historical number of trips on the current route represented by the route information.

[0096] The fourth determining unit is used to determine, in response to detecting that the number of historical trips is greater than a first preset value, that the second response information is target information related to the content requested by the second voice request.

[0097] The fifth determining unit is used to determine, in response to detecting that the number of historical trips is less than or equal to a first preset value, that the second response information is predefined information unrelated to the content requested by the second voice request.

[0098] According to embodiments of this disclosure, the first response information includes navigation information. The current environment information includes road condition information related to the current driving route, and the current driving route includes the route represented by the navigation information. The second determining module includes a sixth determining unit, a seventh determining unit, and an eighth determining unit.

[0099] The sixth determining unit is used to determine, based on road condition information, at least one of the following: the number of traffic lights, the number of turns, and the number of cameras in the current driving route.

[0100] The seventh determining unit is configured to determine, in response to detecting that at least one of the number of traffic lights, the number of turns, and the number of cameras is greater than a second preset value, that the second response information is target information related to the content requested by the second voice request.

[0101] The eighth determining unit is configured to determine, in response to detecting that at least one of the number of traffic lights, the number of turns, and the number of cameras is less than or equal to a second preset value, that the second response information is predefined information unrelated to the content requested by the second voice request.

[0102] According to embodiments of this disclosure, the second voice request includes a voice request related to a child. The voice interaction device also includes an adding module and an obtaining module.

[0103] An additional module is added to add child-related target identification information to the audio information related to the second voice request in response to receiving a second voice request during the output of the first response information.

[0104] The acquisition module is used to extract features from audio information including target identification information, and obtain feature extraction results so as to determine target information related to the content requested by the second voice request based on the feature extraction results.

[0105] According to embodiments of this disclosure, the output module includes a first output unit.

[0106] The first output unit is used to output the second response information during the output interval of the first response information.

[0107] According to embodiments of this disclosure, the second voice request includes a voice request related to a child. The output module includes a ninth determining unit and a second output unit.

[0108] The ninth determining unit is used to determine the response role related to the child.

[0109] The second output unit is used to output second response information based on the audio information of the responding role.

[0110] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0111] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the voice interaction method of the present disclosure.

[0112] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the voice interaction method of the present disclosure.

[0113] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the voice interaction method of this disclosure.

[0114] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0115] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0116] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0117] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as voice interaction methods. For example, in some embodiments, the voice interaction method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the voice interaction method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the voice interaction method by any other suitable means (e.g., by means of firmware).

[0118] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0119] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0120] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0121] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0122] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0123] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0124] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0125] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A voice interaction method, comprising: In response to receiving a second voice request during the output of a first response information corresponding to the first voice request, the current environment information is determined; Based on the current environment information, determine the second response information corresponding to the second voice request; as well as Output the first response information and the second response information; The current environmental information includes driving status information; The step of determining the second response information corresponding to the second voice request based on the current environment information includes: In response to detecting that the driving state information indicates the current driving state is stationary, the second response information is determined to be target information related to the content requested by the second voice request; and In response to detecting that the driving state information indicates that the current driving state is in motion, the second response information is determined to be predefined information that is unrelated to the content requested by the second voice request; The first response information includes navigation information; the second voice request includes a voice request related to a child; and the predefined information includes pre-set soothing information used to comfort a child's needs. The method further includes: in response to receiving the second voice request during the output of the first response information, adding target identification information related to the child to the child-type audio information related to the second voice request, the target identification information being used to characterize the user type as a child-type user; and performing feature extraction on the audio information including the target identification information to obtain a feature extraction result, so as to determine target information related to the content requested by the second voice request based on the feature extraction result.

2. The method according to claim 1, wherein, The first response information includes navigation information; the current environment information includes the driving route information represented by the navigation information; The step of determining the second response information corresponding to the second voice request based on the current environment information includes: Determine the historical number of trips on the current route represented by the route information; In response to detecting that the number of historical trips is greater than a first preset value, the second response information is determined to be target information related to the content requested by the second voice request; and In response to detecting that the number of historical trips is less than or equal to the first preset value, the second response information is determined to be predefined information unrelated to the content requested by the second voice request.

3. The method according to claim 1, wherein, The first response information includes navigation information; the current environment information includes road condition information related to the current driving route, and the current driving route includes the route represented by the navigation information; The step of determining the second response information corresponding to the second voice request based on the current environment information includes: Based on the driving road condition information, determine at least one of the following: the number of traffic lights, the number of turns, and the number of cameras in the current driving route; In response to detecting that at least one of the number of traffic lights, the number of turns, and the number of cameras is less than a second preset value, the second response information is determined to be target information related to the content requested by the second voice request; and In response to detecting that at least one of the number of traffic lights, the number of turns, and the number of cameras is greater than or equal to the second preset value, the second response information is determined to be predefined information unrelated to the content requested by the second voice request.

4. The method according to claim 1, wherein, The output of the first response information and the second response information includes: During the output interval of the first response information, the second response information is output.

5. The method according to claim 1, wherein, The second voice request includes voice requests related to children; The output of the first response information and the second response information includes: Determine the response role associated with the child; as well as Based on the audio information of the responding role, the second response information is output.

6. A voice interaction device, comprising: The first determining module is used to determine the current environment information in response to receiving a second voice request during the process of outputting a first response information corresponding to the first voice request; The second determining module is used to determine the second response information corresponding to the second voice request based on the current environment information; as well as The output module is used to output the first response information and the second response information; The current environmental information includes driving status information; The second determining module includes: The first determining unit is configured to, in response to detecting that the driving state information indicates the current driving state is stationary, determine that the second response information is target information related to the content requested by the second voice request; and The second determining unit is configured to, in response to detecting that the driving state information indicates that the current driving state is a motion state, determine that the second response information is predefined information unrelated to the content requested by the second voice request; The first response information includes navigation information; the second voice request includes a voice request related to a child; and the predefined information includes pre-set soothing information used to comfort a child's needs. The device further includes: The module is configured to, in response to receiving the second voice request during the output of the first response information, add target identification information related to the child to the child-related audio information associated with the second voice request, wherein the target identification information is used to characterize the user type as a child user; and The acquisition module is used to extract features from audio information including the target identification information to obtain feature extraction results, so as to determine target information related to the content requested by the second voice request based on the feature extraction results.

7. The apparatus according to claim 6, wherein, The first response information includes navigation information; The current environmental information includes the driving route information represented by the navigation information; The second determining module includes: The third determining unit is used to determine the historical number of trips of the current driving route represented by the driving route information; The fourth determining unit is configured to, in response to detecting that the number of historical trips is greater than a first preset value, determine that the second response information is target information related to the content requested by the second voice request; and The fifth determining unit is configured to determine, in response to detecting that the number of historical trips is less than or equal to the first preset value, that the second response information is predefined information unrelated to the content requested by the second voice request.

8. The apparatus according to claim 6, wherein, The first response information includes navigation information; the current environment information includes road condition information related to the current driving route, and the current driving route includes the route represented by the navigation information; The second determining module includes: The sixth determining unit is used to determine, based on the driving road condition information, at least one of the following: the number of traffic lights, the number of turns, and the number of cameras in the current driving route; The seventh determining unit is configured to, in response to detecting that at least one of the number of traffic lights, the number of turns, and the number of cameras is less than a second preset value, determine that the second response information is target information related to the content requested by the second voice request; and The eighth determining unit is configured to determine, in response to detecting that at least one of the number of traffic lights, the number of turns, and the number of cameras is greater than or equal to the second preset value, that the second response information is predefined information unrelated to the content requested by the second voice request.

9. The apparatus according to claim 6, wherein, The output module includes: The first output unit is used to output the second response information during the output interval of the first response information.

10. The apparatus according to claim 6, wherein, The second voice request includes voice requests related to children; The output module includes: The ninth determining unit is used to determine the response role associated with the child; as well as The second output unit is used to output the second response information based on the audio information of the responding role.

11. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.

13. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Method and device for controlling sound play in vehicular information entertainment system

    CN102496376A

  • Early education robot speech interaction education system and method

    CN108109622A

  • Broadcast control method in electronic equipment and electronic equipment

    CN113934397A