Voice interaction method, voice interaction device, server and readable storage medium

By integrating parallel command response sentences, a smooth and concise voice broadcast that conforms to user perception habits is generated, solving the problem of lengthy or unclear responses in in-vehicle voice assistants and improving user experience.

CN116110384BActive Publication Date: 2026-03-24GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-07
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

In existing technologies, when in-vehicle voice assistants process multiple user commands, the response content is prone to being lengthy or unclear, which reduces the user experience.

Method used

By receiving parallel instructions, the system merges response sentences according to the sentence structure of natural language to generate fluent and concise voice broadcasts. This includes adjusting the timing of response sentences and adding connecting words to ensure that the voice broadcasts conform to the user's perceptual habits.

Benefits of technology

It improved user satisfaction, reduced waiting time, ensured clear and smooth feedback, and enhanced the practicality of voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116110384B_ABST
    Figure CN116110384B_ABST
Patent Text Reader

Abstract

The application discloses a voice interaction method, a voice interaction device, a server and a readable storage medium. The voice interaction method comprises the following steps: receiving parallel instructions of a user in a cabin forwarded by a vehicle; obtaining a plurality of reply sub-sentences arranged in time sequence corresponding to a plurality of instructions in the parallel instructions; fusing the reply sub-sentences needing adjustment according to the sentence type and time sequence of each of the reply sub-sentences according to the sentence organization mode of natural language to obtain a voice broadcast; and delivering the voice broadcast to the vehicle so that the vehicle feeds back to the user according to the voice broadcast. Thus, by fusing the reply sub-sentences needing adjustment according to the sentence organization mode of natural language, a fluent and concise voice broadcast can be obtained, which is different from piecewise reply or omitted reply to the multiple parallel instructions of the user, the voice broadcast is more in line with the perception habit of the user, the user obtains clear and clear feedback, and the satisfaction of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of vehicles, in particular to a voice interaction method, a voice interaction device, a server and a readable storage medium. BACKGROUND

[0002] With the development of technology, voice interaction has been increasingly widely applied in multiple fields. Among them, in the field of automobiles, the dependence of driving behavior on both hands provides a large number of landing scenarios for voice interaction technology. At present, intelligentization has become an important development direction of the automobile industry, and as one of the important manifestations of the intelligentization of automobiles, the market's requirements and expectations for the intelligent experience that the vehicle voice assistant can bring are also higher and higher.

[0003] In the related art, when a user continuously issues multiple instructions, the general technical solution is to reply to the multiple instructions in turn, which will result in lengthy reply content and may cause the same type of information to be repeatedly broadcast, thereby reducing the user's experience; or a unified brief broadcast is performed, which may result in unclear feedback information due to excessive simplification, leaving room for improvement. SUMMARY

[0004] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, one object of the present application is to propose a voice interaction method that can generate voice broadcast in accordance with the user's perception habits, so that the user can obtain clear and clear feedback, thereby improving the user's satisfaction.

[0005] According to the voice interaction method of the present application, the method comprises: receiving parallel instructions of a user in a cabin forwarded by a vehicle; obtaining a plurality of reply sentences arranged in time sequence corresponding to each instruction in the parallel instructions; fusing the reply sentences to be adjusted according to the sentence type and time sequence of each reply sentence in accordance with the sentence organization mode of natural language, to obtain a voice broadcast; and issuing the voice broadcast to the vehicle, so that the vehicle feeds back to the user according to the voice broadcast.

[0006] According to the voice interaction method of the present application, by fusing the reply sentences to be adjusted in accordance with the sentence organization mode of natural language, a fluent and concise voice broadcast can be obtained, which is different from the piece-by-piece reply or omission reply to the user's multiple parallel instructions, and the voice broadcast is more in line with the user's perception habits, so that the user can obtain clear and clear feedback, thereby improving the user's satisfaction.

[0007] According to the voice interaction method, the method further comprises: in a case that the first reply sentence is not the fusable reply, determining the first reply sentence as a reply sentence that does not need to be adjusted, and issuing the first reply sentence to the vehicle so that the vehicle feeds back to the user according to the first reply sentence. Thus, the user can be fed back in time, and the waiting time of the user is reduced.

[0008] According to the voice interaction method, the method further comprises: in a case that the first reply sentence is not the fusable reply, determining the first reply sentence as a reply sentence that does not need to be adjusted, and issuing the first reply sentence to the vehicle so that the vehicle feeds back to the user according to the first reply sentence. Thus, the user can be fed back in time, and the waiting time of the user is reduced.

[0009] According to the voice interaction method, the method further comprises: in a case that the first reply sentence is not the fusable reply, determining the first reply sentence as a reply sentence that does not need to be adjusted, and issuing the first reply sentence to the vehicle so that the vehicle feeds back to the user according to the first reply sentence. Thus, the user can be fed back in time, and the waiting time of the user is reduced.

[0010] According to the voice interaction method, the method further comprises: after determining whether the first reply sentence is the fusable reply, determining whether a subsequent reply sentence is a query type reply, determining a reply sentence before the first query type reply as a reply sentence that needs to be adjusted, and issuing the query type reply. Thus, the vehicle can reply to the user's question as soon as possible, the waiting time of the user is reduced, and the satisfaction of the user is improved.

[0011] According to the voice interaction method, the method further comprises: after determining whether the first reply sentence is the fusable reply, determining whether a subsequent reply sentence is a query type reply, determining a reply sentence before the first query type reply as a reply sentence that needs to be adjusted, and issuing the query type reply. Thus, the vehicle can reply to the user's question as soon as possible, the waiting time of the user is reduced, and the satisfaction of the user is improved.

[0012] According to the voice interaction method, the method further comprises: in the case that the reply waiting time exceeds the target time length, fusing the current determined reply sentences that need to be adjusted according to a sentence organization mode of natural language, and splicing other reply sentences to obtain a voice broadcast. In this way, the reply waiting time can be limited, and the time for the user to wait can be avoided from being too long.

[0013] According to the voice interaction method, the fusing of the reply sentences that need to be adjusted according to the sentence organization mode of natural language comprises: adjusting the time sequence of at least part of the reply sentences; and / or adding a conjunction between at least part of the reply sentences; and / or fusing a plurality of the reply sentences according to semantics into one sentence. In this way, the connection of the voice broadcast can be smoother, and the length of the voice broadcast can be reduced.

[0014] According to the voice interaction method, the adjusting of the time sequence of at least part of the reply sentences comprises: in the case that the reply sentences that need to be adjusted are multiple, adjusting the time sequence of the reply sentences that need to be adjusted to all the fusable reply sentences being arranged continuously, and the relative time sequence between all the fusable reply sentences being kept unchanged. In this way, the reply sentences can maintain the original time sequence as much as possible, and the satisfaction of the user can be improved.

[0015] According to the voice interaction method, the adjusting of the time sequence of at least part of the reply sentences comprises: in the case that a first fusable reply is located at a first or a second of the multiple reply sentences, keeping the time sequence of the first fusable reply unchanged; and in the case that a first fusable reply is located after a second of the multiple reply sentences, adjusting the first fusable reply to the second of the multiple reply sentences. Through the above setting, the fusable reply with more information can be timely broadcasted, and the satisfaction of the user can be improved.

[0016] According to the voice interaction method, the adjusting of the time sequence of at least part of the reply sentences comprises: in the case that there are non-fusable reply sentences between the multiple fusable reply sentences, adjusting the non-fusable reply sentences between the multiple fusable reply sentences to be after the last fusable reply sentence, and keeping the relative time sequence between the adjusted non-fusable reply sentences unchanged. In this way, the voice broadcast can be more in line with the perception habits of the user.

[0017] According to the voice interaction method, the adding of the conjunction between at least part of the reply sentences comprises: adding a first conjunction between a fusable reply and a general non-fusable reply; and in the case that there is one fusable reply in the reply sentences that need to be adjusted, the fusable reply is broadcasted through a detailed version of TTS, and a second conjunction is added between a non-fusable reply and the detailed version of TTS. In this way, the connection of the voice broadcast can be smoother and more natural.

[0018] The application further provides a voice interaction device.

[0019] The voice interaction device according to the application comprises a receiving module configured to receive parallel instructions of a user in a cabin forwarded by a vehicle; a processing module configured to obtain a plurality of reply clauses arranged in sequence corresponding to each instruction in the parallel instructions; a judging module configured to fuse the reply clauses to be adjusted according to a sentence type and a sequence of each of the reply clauses, and obtain a voice broadcast according to a sentence organization mode of a natural language; and a sending module configured to send the voice broadcast to the vehicle so that the vehicle feeds back to the user according to the voice broadcast.

[0020] According to the voice interaction device, the reply clauses to be adjusted are fused according to the sentence organization mode of the natural language, and the fluent and concise voice broadcast can be obtained. Unlike the reply to the parallel instructions of the user in a piece-by-piece manner or the omission of the reply, the voice broadcast is more in line with the perception habit of the user, so that the user can obtain the explicit and clear feedback, and the user satisfaction is improved, and the practicability of the voice interaction device is improved.

[0021] The application further provides a server.

[0022] The server according to the application comprises a memory and a processor, and the memory stores a computer program, and the computer program is executed by the processor to implement any one of the methods.

[0023] According to the server, the processor executes the computer program stored on the memory, and the fluent and concise voice broadcast can be obtained. Unlike the reply to the parallel instructions of the user in a piece-by-piece manner or the omission of the reply, the voice broadcast is more in line with the perception habit of the user, so that the user can obtain the explicit and clear feedback, and the user satisfaction is improved.

[0024] The application further provides a nonvolatile computer readable storage medium of a computer program.

[0025] The nonvolatile computer readable storage medium of the computer program according to the application implements any one of the methods when the computer program is executed by one or more processors.

[0026] According to the nonvolatile computer readable storage medium of the computer program, the processor executes the computer program stored on the readable storage medium, and the fluent and concise voice broadcast can be obtained. Unlike the reply to the parallel instructions of the user in a piece-by-piece manner or the omission of the reply, the voice broadcast is more in line with the perception habit of the user, so that the user can obtain the explicit and clear feedback, and the user satisfaction is improved.

[0027] Additional aspects and advantages of the application will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following and / or by practice of the application. BRIEF DESCRIPTION OF DRAWINGS

[0028] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, including the accompanying drawings and appended claims.

[0029] Figure 1 is one of flowcharts of the voice interaction method of the present application;

[0030] Figure 2 is one of flowcharts of the voice interaction method of the present application;

[0031] Figure 3 is one of flowcharts of the voice interaction method of the present application;

[0032] Figure 4 is one of flowcharts of the voice interaction method of the present application;

[0033] Figure 5 is one of flowcharts of the voice interaction method of the present application;

[0034] Figure 6 is one of flowcharts of the voice interaction method of the present application;

[0035] Figure 7 is one of flowcharts of the voice interaction method of the present application;

[0036] Figure 8 is one of flowcharts of the voice interaction method of the present application;

[0037] Figure 9 is one of flowcharts of the voice interaction method of the present application;

[0038] Figure 10 is one of flowcharts of the voice interaction method of the present application;

[0039] Figure 11 is one of flowcharts of the voice interaction method of the present application;

[0040] Figure 12 is one of flowcharts of the voice interaction method of the present application; DETAILED DESCRIPTION

[0041] The voice interaction method of the present application is described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference signs represent the same or similar elements or elements having the same or similar functions throughout. The voice interaction method described below by reference to the drawings is exemplary and is only used to explain the present application and cannot be understood as a limitation of the present application.

[0042] Below, the voice interaction method according to the present application is described with reference to the accompanying drawings.

[0043] As shown in Figure 1 , the voice interaction method according to the present application, the method comprises:

[0044] S10: receiving the parallel instructions of the user in the cabin forwarded by the vehicle. That is, during the operation of the vehicle, when the user in the cabin issues a parallel instruction to the vehicle, the radio device in the vehicle can collect the parallel instruction of the user, and the vehicle can forward the collected parallel instruction, so that the receiving module can receive the parallel instruction of the user in the cabin forwarded by the vehicle, and the parallel instruction usually includes multiple instructions. Among them, the above-mentioned parallel instruction can be issued by one user, or can be issued by multiple different users.

[0045] S20: obtaining a plurality of reply clauses arranged in time sequence corresponding to each instruction in the parallel instruction. That is, after receiving the parallel instruction forwarded by the vehicle, the parallel instruction can be identified to divide the parallel instruction into multiple instructions according to the time sequence and determine the content of the instruction, so as to obtain a plurality of reply clauses arranged in time sequence corresponding to each instruction in the parallel instruction.

[0046] S30: according to the sentence type and time sequence of each reply clause, the reply clause to be adjusted is fused according to the sentence organization mode of natural language to obtain a voice broadcast.

[0047] That is, after determining the plurality of reply clauses corresponding to each instruction in the parallel instruction, the reply clause to be adjusted can be determined, and after the reply clause to be adjusted is determined, the reply clause to be adjusted can be fused according to the sentence type and time sequence of the reply clause according to the sentence organization mode of natural language (wherein the sentence organization mode of natural language means that a conjunction is provided between two reply clauses, and the time sequence of the reply clause tends to be consistent with each instruction in the parallel instruction), so as to obtain a smooth and concise voice broadcast.

[0048] S40: issuing the voice broadcast to the vehicle so as to feedback to the user according to the voice broadcast. That is, after generating the voice broadcast, the voice broadcast can be delivered to the vehicle, and the broadcasting device in the vehicle issues a reminder according to the voice broadcast to deliver a brief and clear reply to the user, so that the user can obtain effective feedback. Therefore, it is beneficial to improve the satisfaction of the user.

[0049] According to the voice interaction method, the fluent and concise voice broadcast can be obtained by fusing the reply sentences to be adjusted according to the sentence organization mode of the natural language, unlike the piecewise reply or the omitted reply to the multiple parallel instructions of the user, the voice broadcast is more in line with the perception habit of the user, so that the user can obtain clear and explicit feedback, and the user's satisfaction is improved.

[0050] In the present application, as shown in Figure 2 According to the sentence type and the time sequence of each reply sentence, the reply sentences to be adjusted are fused according to the sentence organization mode of the natural language, including:

[0051] S31: In the case that the first reply sentence is a fusible reply, the first reply sentence is determined as the reply sentence to be adjusted, and the reply sentences to be adjusted are fused according to the sentence organization mode of the natural language; wherein the fusible reply means that the reply sentence includes the subject pointing to the specific control object and the verb for executing the target action.

[0052] It should be noted that the sentence type of the reply sentence includes the fusible reply and the non-fusible reply, the fusible reply means that the reply sentence includes the subject pointing to the specific control object and the verb for executing the target action, such as the main driver window is opened, and the non-fusible reply includes the query type reply and the general type non-fusible reply (such as not supporting the function).

[0053] That is to say, after determining the multiple reply sentences corresponding to each instruction in the parallel instruction, the first reply sentence is first judged, if the first reply sentence is a fusible reply, the first reply sentence is determined as the reply sentence to be adjusted, so that the first reply sentence can be fused with the subsequent reply sentences to be adjusted according to the sentence organization mode of the natural language. Therefore, the fluency of the reply is improved.

[0054] In the present application, as shown in Figure 3 According to the sentence type and the time sequence of each reply sentence, the reply sentences to be adjusted are fused according to the sentence organization mode of the natural language, including:

[0055] S32: In the case that the first reply sentence is a fusible reply, the first reply sentence is determined as the reply sentence to be adjusted, the reply sentences to be adjusted are fused according to the sentence organization mode of the natural language, and the first voice broadcast is issued, the first voice broadcast includes the affirmative response word.

[0056] That is, after determining the plurality of reply clauses corresponding to each instruction in the parallel instruction, the first reply clause is first judged. If the first reply clause is a fusable reply, the first reply clause is determined as a reply clause that needs to be adjusted, so that the first reply clause can be fused with the subsequent reply clause that needs to be adjusted according to the sentence organization mode of natural language. At this time, the first voice broadcast can be issued to the vehicle, so that the vehicle can first play the first voice broadcast, and the first voice broadcast includes an affirmative response word such as "OK" and "No problem". Thus, the waiting time of the user is reduced, the overall fluency of the reply is improved, and the satisfaction of the user is improved.

[0057] In the present application, as shown in Figure 4 the method further comprises:

[0058] S33: In the case that the first reply clause is not a fusable reply, the first reply clause is determined as a reply clause that does not need to be adjusted, and the first reply clause is issued to the vehicle so as to feedback to the user according to the first reply clause.

[0059] That is, after determining the plurality of reply clauses corresponding to each instruction in the parallel instruction, the first reply clause is first judged. If the first reply clause is a fusable reply, the first reply clause is determined as a reply clause that needs to be adjusted, so that the first reply clause can be fused with the subsequent reply clause that needs to be adjusted according to the sentence organization mode of natural language. At this time, the first voice broadcast can be issued to the vehicle, so that the vehicle can first play the first voice broadcast, and the first voice broadcast includes an affirmative response word such as "OK" and "No problem". Thus, the waiting time of the user is reduced, the overall fluency of the reply is improved, and the satisfaction of the user is improved.

[0060] In the present application, as shown in Figure 5 the method further comprises:

[0061] S34: After determining whether the first reply clause is a fusable reply, it is determined whether the subsequent reply clause is a query type reply. The reply clause before the first query type reply is determined as a reply clause that needs to be adjusted, and the query type reply is issued.

[0062] That is, after determining whether the first reply clause is a fusable reply, the subsequent reply clauses can be judged one by one in time sequence. When it is judged that there is a query type reply in the subsequent reply clause, the time sequence of the first query type reply is determined, and the reply clause before the first query type reply is determined as a reply clause that needs to be adjusted. The reply clauses before the first query type reply are fused according to the sentence organization mode of natural language to obtain a voice broadcast. The voice broadcast and the query type reply are issued to the vehicle, so that the vehicle can play the voice broadcast and the query type reply in turn.

[0063] Then, if the first query type reply is not the last reply sentence, the reply sentences after the first query type reply can be sequentially judged according to the time sequence, if there is a second query type reply, the reply sentences after the first query type reply and before the second query type reply can be determined as the reply sentences to be adjusted, and the determination is sequentially performed.

[0064] Through the above setting, the vehicle can reply to the user's question as soon as possible, reducing the user's waiting time, and improving the user's satisfaction.

[0065] In the present application, as shown in Figure 5 After determining whether the first reply sentence is a fusible reply, whether the subsequent reply sentence is a query type reply is determined, comprising:

[0066] S341: After determining whether the first reply sentence is a fusible reply, whether the subsequent reply sentence is a query type reply is determined in the case that the reply waiting time does not exceed the target time length. It should be noted that the target time length can be designed according to the experience and experimental data of the technical personnel, for example, the target time length can be set to 1000ms.

[0067] That is, after receiving the parallel instruction of the user in the cockpit forwarded by the vehicle, the reply waiting time can be timed. After determining whether the first reply sentence is a fusible reply, the subsequent reply sentences can be sequentially judged according to the time sequence, so as to judge whether the subsequent reply sentence is a query type reply in the case that the reply waiting time does not exceed the target time length, and when the reply waiting time exceeds the target time length, the judgment of whether the subsequent reply sentence is a query type reply is stopped. In this way, in the case that the number of reply sentences is large and there is no query type reply in the reply sentences or the time sequence of the query type reply is late, the waiting time caused by the need to judge the query type reply can be avoided. Therefore, the user's satisfaction is improved.

[0068] In the present application, as shown in Figure 5 The method further comprises:

[0069] S35: In the case that the reply waiting time exceeds the target time length, the currently determined reply sentences to be adjusted are fused according to the sentence organization mode of natural language, and other reply sentences are spliced to obtain voice broadcast.

[0070] That is, after determining whether the first reply sentence is a fusable reply, the subsequent reply sentences can be sequentially judged in time sequence to determine whether the subsequent reply sentences are query type replies. In the case where the reply waiting time exceeds the target time length, the judgment on the subsequent reply sentences is stopped, the currently determined reply sentences that need to be adjusted are fused according to the natural language sentence organization mode, and the other reply sentences that are not determined are spliced to obtain the voice broadcast. It should be noted that the currently determined reply sentences that need to be adjusted do not include the reply sentences that have been issued.

[0071] Through the above setting, the reply waiting time is limited to avoid the user waiting for too long, which is beneficial to improve the user's satisfaction.

[0072] In the specific working process, as shown in Figure 11 , the first reply sentence can be first judged. If the first reply sentence is not a fusable reply, the first reply sentence is directly issued. If the first reply sentence is a fusable reply, the first voice broadcast is issued. Then, the second reply sentence can be judged. In the case where the number of current reply sentences is less than the total number of sentences and the reply waiting time does not exceed the target time length, it is determined whether the second reply sentence is a query type reply. If yes, the reply sentences before the second reply sentence (excluding the reply sentences that have been issued) are fused according to the natural language sentence organization mode to obtain the voice broadcast, and the query type reply is issued. If not, the second reply sentence is determined as a reply sentence that needs to be adjusted. Then, the third reply sentence can be judged, and so on. Until the number of current reply sentences is greater than the total number of sentences or the reply waiting time exceeds the target time length, the currently determined reply sentences that need to be adjusted (excluding the reply sentences that have been issued) are fused according to the natural language sentence organization mode, and the other reply sentences are spliced to obtain the voice broadcast.

[0073] In the present application, as shown in Figure 6 , the reply sentences that need to be adjusted are fused according to the natural language sentence organization mode, which includes:

[0074] S36: Adjusting the time sequence of at least part of the reply sentences. That is, when the reply sentences that need to be adjusted are determined, the time sequence of at least part of the reply sentences can be adjusted, such as arranging a plurality of fusable replies in time sequence in succession, so that the plurality of fusable replies can be fused.

[0075] S37: And / or, adding a conjunction between at least part of the reply sentences. That is, when the reply sentences that need to be adjusted are determined, a conjunction can be added between at least part of the reply sentences, such as adding a conjunction between the fusable replies and the non-fusable replies, so that the connection of the voice broadcast is more fluent.

[0076] S38: and / or, fuse multiple reply sentences into one sentence according to semantics. That is, when the reply sentences to be adjusted are determined, multiple fusable replies can be fused into one reply sentence according to semantics to reduce the length of voice broadcast. For example, when the sentences of the parallel instruction include "Driver's window open" and "Driver's heating open", the corresponding reply sentences are "Driver's window open" and "Driver's heating open", which can be fused according to semantics to "Driver's window and heating open".

[0077] In the present application, as shown in Figure 7 S36: adjusting the time sequence of at least part of the reply sentences, including:

[0078] S361: in the case of multiple reply sentences to be adjusted, adjusting the time sequence of the reply sentences to be adjusted to arrange all fusable replies in sequence, and keeping the relative time sequence between all fusable replies unchanged.

[0079] That is, in the case of multiple reply sentences to be adjusted, the time sequence of the reply sentences to be adjusted can be adjusted to arrange all fusable replies in sequence and fuse them, and the relative time sequence between the fusable replies remains unchanged, that is, the fusable reply with an earlier time sequence before adjustment is still in the front after adjustment. Thus, the reply sentence can maintain the original time sequence as much as possible, which is beneficial to improve user satisfaction.

[0080] In the present application, as shown in Figure 8 S36: adjusting the time sequence of at least part of the reply sentences, including:

[0081] S362: in the case of the first fusable reply being located at the first or second of the multiple reply sentences, keeping the time sequence of the first fusable reply unchanged. That is, in the case of the first fusable reply being located at the first or second of the multiple reply sentences, the fusable replies in the subsequent reply sentences can be adjusted to arrange the fusable replies in the subsequent reply sentences in sequence after the first fusable reply to keep the time sequence of the first fusable reply unchanged.

[0082] S363: in the case of the first fusable reply being located after the second of the multiple reply sentences, adjusting the first fusable reply to the second of the multiple reply sentences. That is, in the case of the first fusable reply being located after the second of the multiple reply sentences, the first fusable reply can be adjusted to the second of the multiple reply sentences, and the fusable replies in the subsequent reply sentences can be adjusted to arrange the fusable replies in the subsequent reply sentences in sequence after the first fusable reply.

[0083] Through the above setting, in the case of small timing adjustment of the reply sentence, the fusible reply with more information can be timely reported, thereby improving the user's satisfaction.

[0084] In the present application, as shown in S36: adjusting the timing of at least part of the reply sentence, comprising: Figure 9

[0085] S364: In the case of including non-fusible replies between the plurality of fusible replies, the non-fusible replies located between the plurality of fusible replies are adjusted to be after the last fusible reply, and the relative timing between the adjusted non-fusible replies remains unchanged.

[0086] That is, in the case of including non-fusible replies between the plurality of fusible replies, the non-fusible replies located between the plurality of fusible replies can be adjusted to be after the last fusible reply, so that the plurality of fusible replies can be continuously arranged before the non-fusible replies with the relative timing unchanged, so that the content after the fusion of the plurality of fusible replies can be reported first, and the relative timing between the adjusted non-fusible replies can be set to remain unchanged, reducing the timing change of the reply sentence. Thus, the voice broadcast is more in line with the user's perception habits, thereby improving the user's satisfaction.

[0087] In the present application, as shown in S37: adding a conjunction between at least part of the reply sentences, comprising: Figure 10

[0088] S371: Adding a first conjunction between the fusible reply and the general non-fusible reply. That is, if the reply sentence after the fusible reply is a general non-fusible reply, a first conjunction can be added between the fusible reply and the general non-fusible reply. The first conjunction is set to "but" or other words with similar semantics.

[0089] S372: In the case of one fusible reply in the reply sentence to be adjusted, the fusible reply is broadcast through the detailed TTS, and a second conjunction is added between the non-fusible reply and the detailed TTS. That is, in the case of one fusible reply in the reply sentence to be adjusted, the fusible reply does not need to be fused, and the fusible reply can be directly broadcast through the detailed TTS. If there is a non-fusible reply before the detailed TTS, a second conjunction can be added between the non-fusible reply and the detailed TTS, and the second conjunction is set to "in addition" or other words with similar semantics.

[0090] Through the above setting, the connection of the voice broadcast is more smooth and natural, thereby improving the user's satisfaction.

[0091] ​​It should be noted that when the reply sentence is 4, according to whether the 4 sentences are fusable replies, there are multiple examples as shown in the following table.

[0092]

[0093] The application further provides a voice interaction device.

[0094] The voice interaction device according to the application comprises a receiving module, a processing module, a judging module and a sending module.

[0095] The receiving module is configured to receive parallel instructions of a user in a vehicle cabin forwarded by the vehicle;

[0096] The processing module is configured to obtain a plurality of reply sentences arranged in time sequence corresponding to each instruction in the parallel instructions;

[0097] The judging module is configured to fuse the reply sentences that need to be adjusted according to the sentence type and time sequence of each reply sentence, according to the sentence organization mode of natural language, to obtain a voice broadcast,

[0098] The sending module is configured to send the voice broadcast to the vehicle, so that the vehicle feeds back to the user according to the voice broadcast.

[0099] According to the voice interaction device, the reply sentences that need to be adjusted are fused according to the sentence organization mode of natural language, so that a fluent and concise voice broadcast can be obtained, which is different from the one-by-one reply or omitted reply to the user's multiple parallel instructions, and the voice broadcast is more in line with the user's perception habits, so that the user can obtain clear feedback, which is beneficial to improve the user's satisfaction and improve the practicality of the voice interaction device.

[0100] In the application, the judging module is further configured to, in the case that the first reply sentence is a fusable reply, determine the first reply sentence as a reply sentence that needs to be adjusted, and fuse the reply sentence that needs to be adjusted according to the sentence organization mode of natural language; wherein the fusable reply indicates that the reply sentence includes a subject pointing to a specific control object and a verb for executing a target action.

[0101] In the application, the judging module is further configured to, in the case that the first reply sentence is a fusable reply, determine the first reply sentence as a reply sentence that needs to be adjusted, fuse the reply sentence that needs to be adjusted according to the sentence organization mode of natural language, and send a first voice broadcast, wherein the first voice broadcast includes an affirmative response word.

[0102] In the application, the judging module is further configured to, in the case that the first reply sentence is not a fusable reply, determine the first reply sentence as a reply sentence that does not need to be adjusted, and send the first reply sentence to the vehicle, so that the vehicle feeds back to the user according to the first reply sentence.

[0103] In the present application, the judging module is further configured to determine whether the subsequent reply sentence is a query type reply after determining whether the first reply sentence is a fusable reply, determine the reply sentence before the first query type reply as a reply sentence needing adjustment, and issue the query type reply.

[0104] In the present application, the judging module is further configured to determine whether the subsequent reply sentence is a query type reply in the case that the reply waiting time does not exceed the target time length after determining whether the first reply sentence is a fusable reply.

[0105] In the present application, the judging module is further configured to fuse the currently determined reply sentence needing adjustment according to the sentence organization mode of natural language and splice other reply sentences to obtain a voice broadcast in the case that the reply waiting time exceeds the target time length.

[0106] In the present application, the judging module is further configured to adjust the time sequence of at least part of the reply sentences, and / or add conjunctions between at least part of the reply sentences, and / or fuse a plurality of the reply sentences according to semantics into one sentence.

[0107] In the present application, the judging module is further configured to adjust the time sequence of the reply sentences needing adjustment to all the fusable replies being arranged continuously and the relative time sequence between all the fusable replies being kept unchanged in the case that the reply sentences needing adjustment are multiple.

[0108] In the present application, the judging module is further configured to keep the time sequence of the first fusable reply unchanged in the case that the first fusable reply is located at the first or second of the multiple reply sentences, and adjust the first fusable reply to the second of the multiple reply sentences in the case that the first fusable reply is located after the second of the multiple reply sentences.

[0109] In the present application, the judging module is further configured to adjust the non-fusable reply located between the multiple fusable replies to be after the last fusable reply in the case that the non-fusable reply is included between the multiple fusable replies, and the relative time sequence between the adjusted non-fusable replies is kept unchanged.

[0110] In the present application, the judging module is further configured to add a first conjunction between the fusable reply and the general non-fusable reply, and add a second conjunction between the fusable reply and the detailed version TTS in the case that there is one fusable reply in the reply sentence needing adjustment, and the fusable reply is broadcasted through the detailed version TTS.

[0111] According to the voice interaction device, the fluent and concise voice broadcast can be obtained by fusing the reply sentences to be adjusted according to the natural language sentence organization mode, unlike the piecewise reply or the omitted reply to the multiple parallel instructions of the user, the voice broadcast is more in line with the perception habit of the user, so that the user can obtain clear and explicit feedback, and the user satisfaction is improved, and the practicability of the voice interaction device is improved.

[0112] The voice interaction device in the application can be a device, a component, an integrated circuit or a chip in a mobile terminal. For example, the mobile terminal can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA).

[0113] The voice interaction device in the application can be a device with an operating system. The operating system can be an Android operating system, an ios operating system or other possible operating systems, and the application does not make specific limitations.

[0114] The voice interaction device provided in the application can realize Figures 1 to 10 the voice interaction method, and each process realized by the voice interaction device is not repeated here.

[0115] The application further provides a server.

[0116] The server according to the application comprises a memory and a processor, the memory stores a computer program, and the computer program is executed by the processor to realize the method.

[0117] The server according to the application can obtain the fluent and concise voice broadcast by executing the computer program stored in the memory by the processor, unlike the piecewise reply or the omitted reply to the multiple parallel instructions of the user, the voice broadcast is more in line with the perception habit of the user, so that the user can obtain clear and explicit feedback, and the user satisfaction is improved.

[0118] The application further provides a nonvolatile computer readable storage medium of a computer program.

[0119] Referring to Figure 10 Fig. 1 shows that the nonvolatile computer readable storage medium 100 of the computer program 101 according to the application realizes the method when the computer program 101 is executed by one or more processors 200.

[0120] The processor is any of the processors in the electronic devices described above. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0121] The nonvolatile computer readable storage medium of the computer program according to the application can obtain fluent and concise voice broadcast by executing the computer program stored on the readable storage medium by the processor, unlike piecewise reply or omission reply to multiple parallel instructions of the user, the voice broadcast is more in line with the perception habit of the user, so that the user can obtain clear and clear feedback, which is beneficial to improve the satisfaction of the user.

[0122] In the description of the application, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore cannot be understood as a limitation of the application.

[0123] In the description of the application, "first feature" and "second feature" can include one or more features.

[0124] In the description of the application, "a plurality of" means two or more.

[0125] In the description of the application, "above" or "below" the first feature of the second feature can include direct contact between the first and second features, or can include indirect contact between the first and second features through another feature therebetween.

[0126] In the description of the application, "above", "above" and "above" of the first feature of the second feature include that the first feature is directly above and obliquely above the second feature, or only indicates that the horizontal height of the first feature is higher than that of the second feature.

[0127] In the description of the application within this specification, the terms "a voice interaction method", "some voice interaction methods", "an illustrative voice interaction method", "an example", "a specific example", or "some examples" and the like mean that a particular feature, structure, material, or characteristic being referred to is included in at least one voice interaction method or example of the application. The appearances of the above terms in various places in the specification are not necessarily all referring to the same voice interaction method or example. Furthermore, the particular features, structures, materials, or characteristics can be combined in any suitable manner in one or more voice interaction methods or examples.

[0128] Although the voice interaction methods of the present application have been shown and described with respect to certain embodiments, it is to be understood that certain modifications, substitutions, alternatives, and variations can be made thereto by those skilled in the art without departing from the spirit and scope of the application, which is defined by the claims and their equivalents.

Claims

1. A voice interaction method, characterized in that, The method includes: Receive parallel commands from users in the cockpit forwarded by the vehicle; Obtain multiple response clauses arranged in time sequence corresponding to each instruction in the parallel instructions; Based on the sentence type and timing of each response sentence, the response sentences that need to be adjusted are merged according to the sentence organization method of natural language to obtain the voice broadcast; The voice broadcast is sent to the vehicle so that the vehicle can provide feedback to the user based on the voice broadcast; the step of integrating the reply sentences that need adjustment according to the sentence type and timing of each reply sentence, and in accordance with the sentence organization method of natural language, includes: If the first response clause is a mergeable response, the first response clause is identified as the response clause that needs adjustment, and the response clauses that need adjustment are merged according to the sentence structure of natural language; wherein, the mergeable response means that the response clause includes a subject pointing to a specific controlled object and a verb used to perform the target action; If the first reply clause is a mergeable reply, determine whether the subsequent reply clauses are query-type replies. If it is determined that there are query-type replies in the subsequent reply clauses, identify the reply clauses before the first query-type reply as reply clauses that need to be adjusted, and issue the query-type reply.

2. The voice interaction method according to claim 1, characterized in that, The process of integrating the adjusted response sentences according to the sentence type and timing of each response sentence, following the natural language sentence organization method, includes: If the first response clause is a mergeable response, the first response clause is identified as the response clause that needs to be adjusted. The response clause that needs to be adjusted is merged according to the sentence structure of natural language, and the first voice broadcast is issued. The first voice broadcast includes response words used to express affirmation.

3. The voice interaction method according to claim 1, characterized in that, The method further includes: If the first response clause is not a mergeable response, the first response clause is determined as a response clause that does not need to be adjusted, and the first response clause is sent to the vehicle so that the vehicle can provide feedback to the user based on the first response clause.

4. The voice interaction method according to claim 1, characterized in that, When the first response clause is a mergeable response, after determining whether subsequent response clauses are query-type responses, if it is determined that a query-type response exists in a subsequent response clause, determining whether a subsequent response clause is a query-type response includes: After determining that the first response clause is a mergeable response, and provided that the response waiting time does not exceed the target duration, determine whether subsequent response clauses are query-type responses.

5. The voice interaction method according to claim 4, characterized in that, The method further includes: If the response waiting time exceeds the target duration, the currently determined response sentences that need adjustment are merged according to the sentence structure of natural language, and other response sentences are spliced ​​together to obtain the voice broadcast.

6. The voice interaction method according to any one of claims 2-5, characterized in that, The adjusted response sentences, organized according to the sentence structure of natural language, include: Adjust the timing of at least some of the response clauses; And / or, Add conjunctions between at least some of the aforementioned response clauses; And / or, Multiple response clauses are merged into one clause according to their semantics.

7. The voice interaction method according to claim 6, characterized in that, The adjustment of the timing of at least some of the response clauses includes: When there are multiple response clauses that need adjustment, the timing of the response clauses that need adjustment is adjusted so that all the mergeable responses are arranged consecutively, and the relative timing between all the mergeable responses remains unchanged.

8. The voice interaction method according to claim 6, characterized in that, The adjustment of the timing of at least some of the response clauses includes: If the first mergeable response is located in the first or second of the plurality of response clauses, the timing of the first mergeable response remains unchanged; If the first mergeable response is located after the second of the plurality of response clauses, the first mergeable response is adjusted to the second of the plurality of response clauses.

9. The voice interaction method according to claim 6, characterized in that, The adjustment of the timing of at least some of the response clauses includes: In the case where there are non-fusionable responses among multiple fusionable responses, the non-fusionable responses located among the multiple fusionable responses are adjusted to be after the last fusionable response, and the relative timing between the adjusted non-fusionable responses remains unchanged.

10. The voice interaction method according to claim 6, characterized in that, The addition of connecting words between at least some of the response clauses includes: Add a first connecting word between mergeable replies and general non-mergeable replies; If there is only one mergeable reply in the reply clause that needs adjustment, the mergeable reply is broadcast through the detailed version of TTS, and a second connecting word is added between the non-mergeable reply and the detailed version of TTS.

11. A voice interaction device, implementing the method according to any one of claims 1-10, characterized in that, include: The receiving module is used to receive parallel commands from the user in the cockpit forwarded by the vehicle. The processing module is used to obtain multiple response sentences arranged in time sequence corresponding to each instruction in the parallel instructions; The judgment module is used to merge the reply sentences that need to be adjusted according to the sentence type and timing of each reply sentence, and to obtain the voice broadcast. The sending module is used to send the voice broadcast to the vehicle so that the vehicle can provide feedback to the user based on the voice broadcast.

12. A server, characterized in that, The server includes a memory and a processor, the memory storing a computer program that, when executed by the processor, implements the method according to any one of claims 1-10.

13. A non-volatile computer-readable storage medium for computer programs, characterized in that, When the computer program is executed by one or more processors, it implements the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Voice interaction method, vehicle and storage medium

    CN114898752A