Voice interaction method, voice interaction device, server and readable storage medium
By merging the parallel commands of the in-vehicle voice assistant, a concise and clear voice broadcast is generated, which solves the problem of lengthy or unclear responses from multiple commands in the in-vehicle voice assistant and improves the user experience.
Patent Information
- Application Number
- CN202211100018.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-09-07
AI Technical Summary
In existing technologies, when in-vehicle voice assistants receive multiple user commands, their responses are often lengthy or unclear, leading to a degraded user experience.
By receiving parallel instructions, it generates voice broadcasts that integrate multiple instructions, merges responses to the same functions or controls, handles dependencies and conflicts, and merges voice regions to generate concise and clear feedback.
It reduced user waiting time, improved the accuracy and clarity of feedback, and enhanced user satisfaction.
Smart Images

Figure CN115527535B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of vehicles, in particular to a voice interaction method, a voice interaction device, a server and a readable storage medium. BACKGROUND
[0002] With the development of technology, voice interaction has been increasingly widely applied in multiple fields. Among them, in the field of automobiles, the dependence of driving behavior on both hands provides a large number of landing scenarios for voice interaction technology. At present, intelligentization has become an important development direction of the automobile industry, and as one of the important manifestations of the intelligentization of automobiles, the market's requirements and expectations for the intelligent experience that the vehicle voice assistant can bring are also higher and higher.
[0003] In the related art, when a user continuously issues multiple instructions, the general technical solution is to reply to the multiple instructions in turn, which will result in lengthy reply content and may cause the same type of information to be repeatedly broadcast, thereby reducing the user's experience; or unified brief broadcast, which may cause the feedback information to be unclear due to excessive simplification, leaving room for improvement. SUMMARY
[0004] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, one object of the present application is to propose a voice interaction method, which has a short user waiting time, accurate and clear reply content, so that the user can obtain clear feedback, thereby improving the user's satisfaction.
[0005] According to the voice interaction method of the present application, the method comprises: receiving multiple parallel instructions of a user in a cabin forwarded by a vehicle; generating a voice broadcast that integrates the replies corresponding to the multiple parallel instructions according to the vehicle function or vehicle-mounted system control pointed to in each instruction, the operation type of the instruction and the pointing audio area; and issuing the voice broadcast to the vehicle so as to feed back to the user according to the voice broadcast.
[0006] According to the voice interaction method of the present application, by receiving and recognizing multiple parallel instructions of a user in a cabin, a voice broadcast that integrates the replies corresponding to the multiple parallel instructions is generated to provide a short and clear reply to the user, without replying to the user's multiple instructions one by one, thereby reducing the user's waiting time and providing accurate and clear reply content, so that the user can obtain clear feedback, thereby improving the user's satisfaction.
[0007] According to the voice interaction method, the voice broadcast fused with the replies corresponding to the multiple parallel instructions is generated according to the vehicle function or the vehicle-mounted system control pointed to in each instruction, the operation type of the instruction, and the sound area pointed to by the instruction, and includes: obtaining multiple replies corresponding to multiple parallel instructions; performing vehicle function or vehicle-mounted system control merging processing on the replies pointing to the same vehicle function or the same vehicle-mounted system control; performing sound area merging processing on the replies in which the vehicle function or the vehicle-mounted system control does not have a dependent relationship and does not have a conflict relationship with an intersection sound area after the vehicle function or the vehicle-mounted system control merging processing; and performing fusion based on the operation type of the reply after the sound area merging processing to obtain the voice broadcast. Thus, the length of the voice broadcast can be effectively reduced.
[0008] According to the voice interaction method, the vehicle function or vehicle-mounted system control merging processing on the replies pointing to the same vehicle function or the same vehicle-mounted system control includes: de-duplicating the replies with the same vehicle function or vehicle-mounted system control, and taking the union of each sound area pointed to by the replies with the same vehicle function or vehicle-mounted system control; and in the case where there are adjacent sound areas, fusing the adjacent sound areas by using a directional word to reduce the number of sound areas and not change the range of the sound areas. Thus, the accuracy of the voice broadcast is improved.
[0009] According to the voice interaction method, the sound area merging processing on the replies in which the vehicle function or the vehicle-mounted system control does not have a dependent relationship and does not have a conflict relationship with an intersection sound area after the vehicle function or the vehicle-mounted system control merging processing includes: de-duplicating the replies with the same sound area, and concatenating the corresponding vehicle function or vehicle-mounted system control after the sound area. Thus, the corresponding reply can be performed in the order of the instructions, and the user satisfaction is improved.
[0010] According to the voice interaction method, the fusion based on the operation type of the reply after the sound area merging processing includes: in the case where the operation types of the replies after the sound area merging processing are the same, merging the replies while keeping the operation type unchanged; and in the case where the operation types of the replies after the sound area merging processing are different, merging the replies by changing the operation type to a general type expression. Thus, the operation type is uniformly replied, and the length of the voice broadcast is reduced.
[0011] According to the voice interaction method, the method further includes: in the case where the operation type of the reply is not opening, issuing the reply to the vehicle in order to perform concatenation broadcast by the vehicle for the replies in which the vehicle function or the vehicle-mounted system control has a dependent relationship after the vehicle function or the vehicle-mounted system control merging processing. Thus, the user can obtain accurate feedback.
[0012] According to the voice interaction method, the method further comprises: for the reply of the vehicle function or the vehicle-mounted system control after the merging processing, in the case that the operation type is opening, a first voice broadcast is issued to the vehicle so as to be broadcast by the vehicle, and the first voice broadcast is used to indicate the dependency relationship. Thus, the vehicle can express the dependency relationship when replying, so as to improve the perception of the user.
[0013] According to the voice interaction method, the method further comprises: for the reply of the vehicle function or the vehicle-mounted system control after the merging processing, in the case that the operation type is opening, a first voice broadcast is issued to the vehicle so as to be broadcast by the vehicle, and the first voice broadcast is used to indicate the dependency relationship. Thus, the vehicle can express the dependency relationship when replying, so as to improve the perception of the user.
[0014] The application further provides a voice interaction device.
[0015] The voice interaction device comprises: a receiving module, which is used to receive multiple parallel instructions of a user in a cabin forwarded by a vehicle; a processing module, which is used to generate a voice broadcast fused with replies corresponding to the multiple parallel instructions according to the vehicle function or the vehicle-mounted system control pointed to in each instruction, the operation type of the instruction and the source sound area of the instruction; and a sending module, which is used to issue the voice broadcast to the vehicle so as to feed back to the user according to the voice broadcast.
[0016] According to the voice interaction device, multiple parallel instructions of a user in a cabin are received and identified, a voice broadcast fused with replies corresponding to the multiple parallel instructions is generated, so as to give a short and clear reply to the user, and it is not necessary to reply to multiple instructions of the user one by one, the waiting time of the user is reduced, the reply content is accurate and clear, the user can obtain clear feedback, the satisfaction of the user is improved, and the practicability of the voice interaction device is improved.
[0017] The application further provides a server.
[0018] The server comprises a memory and a processor, the memory stores a computer program, and the computer program is executed by the processor to realize the method.
[0019] The server according to the present application can generate a voice broadcast that integrates the replies corresponding to the multiple parallel instructions, so as to make a brief and clear reply to the user, without replying to the multiple instructions of the user one by one, thereby reducing the waiting time of the user, and the reply content is accurate and clear, so that the user can obtain clear feedback, and the satisfaction of the user is improved.
[0020] The present application further provides a non-volatile computer readable storage medium of a computer program.
[0021] The non-volatile computer readable storage medium of the computer program according to the present application, when the computer program is executed by one or more processors, implements any of the above-mentioned methods.
[0022] The non-volatile computer readable storage medium of the computer program according to the present application, when the computer program is executed by one or more processors, implements any of the above-mentioned methods.
[0023] Additional aspects and advantages of the present application will be made apparent by the following description and the appended claims. BRIEF DESCRIPTION OF DRAWINGS
[0024] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description, taken in conjunction with the accompanying drawings, in which:
[0025] Figure 1 is one of flowcharts of the voice interaction method of the present application;
[0026] Figure 2 is one of flowcharts of the voice interaction method of the present application;
[0027] Figure 3 is one of flowcharts of the voice interaction method of the present application;
[0028] Figure 4 is one of flowcharts of the voice interaction method of the present application;
[0029] Figure 5 is one of flowcharts of the voice interaction method of the present application;
[0030] Figure 6 is one of flowcharts of the voice interaction method of the present application;
[0031] Figure 7 is the seventh flowchart of the voice interaction method of the present application;
[0032] Figure 8 is the eighth flowchart of the voice interaction method of the present application;
[0033] Figure 9 is the working flowchart of the voice interaction method of the present application;
[0034] Figure 10 is the connection state schematic diagram of the non-volatile computer readable storage medium and the processor of the voice interaction method of the present application. DETAILED DESCRIPTION
[0035] The voice interaction method of the present application is described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The voice interaction method described below by referring to the accompanying drawings is exemplary and is only used to explain the present application and cannot be understood as a limitation of the present application.
[0036] Hereinafter, the voice interaction method according to the present application is described with reference to the accompanying drawings.
[0037] As shown in Figure 1 , according to the voice interaction method of the present application, the method comprises:
[0038] S10: receiving multiple parallel instructions of the user in the cabin forwarded by the vehicle. That is, during the operation of the vehicle, when the user in the cabin issues an instruction to the vehicle, the radio device in the vehicle can collect the instruction of one user, if the instruction is one, the vehicle can directly execute the corresponding operation according to the instruction and make the corresponding reply; if the instruction is multiple parallel instructions, the above multiple parallel instructions can be issued by one user or multiple different users, the vehicle can collect the multiple parallel instructions and forward them to the outside, so that the receiving module can receive the multiple parallel instructions of the user in the cabin forwarded by the vehicle.
[0039] S20: generating a voice broadcast that integrates the replies corresponding to the multiple parallel instructions according to the vehicle function or the vehicle-mounted system control pointed to in each instruction, the operation type of the instruction and the pointing audio area. It should be noted that the vehicle function refers to functions such as heating seat, ventilation, etc., the vehicle-mounted system control refers to controls such as vehicle-mounted screen, window, air conditioner, etc., the operation type of the instruction includes opening, closing, adjusting, and the pointing audio area refers to the in-vehicle space area such as the main driver, the co-driver, the rear row, the front row, etc.
[0040] That is, when receiving multiple parallel instructions forwarded by the vehicle, the instructions can be identified to determine the vehicle function or vehicle-mounted system control pointed by the instructions, the operation type of the instructions and the sound area pointed by the instructions, to determine the reply content required by the multiple parallel instructions. At this time, the associated parts in the reply content corresponding to the multiple parallel instructions can be integrated, such as merging the associated vehicle functions or merging the sound areas pointed, to generate a voice broadcast that integrates the replies corresponding to the multiple parallel instructions.
[0041] S30: issuing the voice broadcast to the vehicle, so that the vehicle feeds back to the user according to the voice broadcast. That is, when the voice broadcast that integrates the replies corresponding to the multiple parallel instructions is generated, the voice broadcast can be delivered to the vehicle, and the public address equipment in the vehicle issues a reminder according to the voice broadcast to deliver a brief and clear reply to the user, so that the user can obtain feedback. Therefore, it is beneficial to improve the satisfaction of the user.
[0042] According to the voice interaction method, by receiving and identifying multiple parallel instructions of the user in the cabin, a voice broadcast that integrates the replies corresponding to the multiple parallel instructions is generated to reply to the user briefly and clearly, without replying to multiple instructions of the user one by one, which reduces the waiting time of the user and makes the reply content accurate and clear, so that the user can obtain clear feedback, which is beneficial to improve the satisfaction of the user.
[0043] In the present application, as shown in Figure 2 S20: generating a voice broadcast that integrates the replies corresponding to the multiple parallel instructions according to the vehicle function or vehicle-mounted system control pointed by each instruction, the operation type of the instruction and the sound area pointed by the instruction, including:
[0044] S21: obtaining multiple replies corresponding to the multiple parallel instructions. That is, when receiving multiple parallel instructions forwarded by the vehicle, the instructions can be segmented and identified to determine the content of the instructions to obtain multiple replies respectively corresponding to the multiple parallel instructions.
[0045] S22: performing vehicle function or vehicle-mounted system control merging processing on the replies pointing to the same vehicle function or the same vehicle-mounted system control. That is, the content of the multiple replies can be determined to determine whether the multiple replies have content pointing to the same vehicle function or the same vehicle-mounted system control, and if so, the corresponding content can be merged.
[0046] For example, when the parallel instructions of the user are "main driver seat heating is turned on, right rear seat heating is turned on", the corresponding replies are "main driver seat heating is turned on, right rear seat heating is turned on", at this time, the two replies have the same vehicle function, and "main driver seat heating" and "right rear seat heating" can be merged into "main driver seat and right rear seat heating".
[0047] S23: performing the merging processing of the audio zones for the reply of the vehicle function or the vehicle control system control after the merging processing of the vehicle function or the vehicle control system control, wherein the vehicle function or the vehicle control system control has no dependency relationship and no conflict relationship with intersection audio zones. It should be noted that the dependency relationship and the conflict relationship are pre-stored programs designed by engineers, the dependency relationship refers to the switch of the vehicle control system control and the function that can be adjusted after the vehicle control system control is turned on, such as the screen switch and the screen brightness adjustment; the conflict relationship refers to the vehicle function that cannot be turned on at the same time, such as the one-way wind and the mirror wind.
[0048] That is, after the merging processing of the vehicle function or the vehicle control system control is completed, if the vehicle function or the vehicle control system control has no dependency relationship and no conflict relationship with intersection audio zones, the corresponding audio zones can be merged. For example, when the user's parallel instruction is "main driver seat heating is turned on, main driver seat ventilation is turned on", the corresponding reply is "main driver seat heating is turned on, main driver seat ventilation is turned on", at this time, the audio zones can be fused, and only one main driver seat in the reply is retained.
[0049] S24: based on the operation type pointed by the reply after the merging processing of the audio zones, fusing to obtain the voice broadcast. That is, after the merging processing of the audio zones, the operation type pointed by the reply can be fused to unify the reply. Therefore, the length of the voice broadcast is effectively reduced.
[0050] For example, when the user's parallel instruction is "main driver seat heating is turned on, main driver seat ventilation is turned on", the corresponding reply is "main driver seat heating is turned on, main driver seat ventilation is turned on", at this time, the operation type can be fused, that is, "main driver seat heating is turned on" and "main driver seat ventilation is turned on" are merged into "main driver seat heating and ventilation are turned on".
[0051] In the present application, as shown in Figure 3 S22: performing the merging processing of the vehicle function or the vehicle control system control for the reply pointing to the same vehicle function or the same vehicle control system control, comprising:
[0052] S221: removing the duplicate replies of the same vehicle function or the same vehicle control system control, and taking the union of the audio zones pointed by the replies of the same vehicle function or the same vehicle control system control; wherein in the case that there are adjacent audio zones, the adjacent audio zones are fused by the direction words to reduce the number of audio zones and not to change the range of the audio zones.
[0053] That is, after the same vehicle function or vehicle system control reply is de-duplicated, the respective sound zones to which the same vehicle function or vehicle system control reply points can be taken and collected to reduce the number of sound zones in the reply and reduce the length of the voice broadcast. In the case of adjacent sound zones (such as the main driver seat and the co-driver seat), the adjacent sound zones can be fused by using directional words, such as fusing the main driver seat and the co-driver seat into the front seats, or fusing the main driver seat and the left rear seat into the left side seat, thereby reducing the number of sound zones while ensuring that the range of the sound zones remains unchanged. Thus, the accuracy of the voice broadcast is improved.
[0054] Specifically, as shown in Table 1, when the functions in the reply are the same and the two pointing sound zones are non-adjacent sound zones, such as the main driver and the right rear, they can be directly represented as “main driver and right rear”; and when the functions in the reply are the same and the two pointing sound zones are adjacent sound zones, such as the main driver and the co-driver, they can be merged into the front row; such as the main driver and the left rear, they can be merged into the left side; such as the co-driver and the right rear, they can be merged into the right side; such as the left rear and the right rear, they can be merged into the rear row. When the functions in the reply are the same and include all sound zones, the corresponding sound zones are represented as “all”.
[0055] Table 1
[0056]
[0057] For example, when the user's parallel instruction is “main driver seat heating on, co-driver seat heating on”, the corresponding reply is “main driver seat heating on, co-driver seat heating on”, at this time, the vehicle functions in the reply are the same, and the main driver and the co-driver are adjacent sound zones, so “main driver seat heating” and “co-driver seat heating” can be fused into “front seat heating”.
[0058] For example, when the user's parallel instruction is “main driver window open, front window open”, the corresponding reply is “main driver window open, front window open”, at this time, the vehicle functions in the reply are the same, and there is a sound zone overlap, so “main driver window” and “front window” can be fused into “front window”.
[0059] It needs to be explained that in order to make the fusion order of the reply more consistent with the fusion order when the user issues the instruction, the instruction involving three fusion points is further regulated. When three points are said at the same time, first see if the first two can be fused: if the first two sound areas can be fused, broadcast the fusion of the first two sound areas and add the third sound area, such as the driver seat, the front passenger seat and the left rear seat can be combined into the front row and the left rear; if the second sound area cannot be fused, broadcast the sound area formed by the fusion of the first sound area and the last two sound areas, such as the driver seat, the right rear seat and the front passenger seat can be combined into the driver seat and the right side.
[0060] For example, when the user's parallel instruction is "the driver seat heating is turned on, the front passenger seat heating is turned on, and the right rear seat heating is turned on", the corresponding reply is "the driver seat heating is turned on, the front passenger seat heating is turned on, and the right rear seat heating is turned on", at this time, the vehicle functions in the reply are the same, and the driver and the front passenger are adjacent sound areas, and "the driver seat heating", "the front passenger seat heating" and "the rear seat heating" can be combined into "the front row and the right rear seat heating".
[0061] In the present application, as shown in S23, the reply after the merging processing of the vehicle function or the vehicle-mounted system control is performed, the reply in which the vehicle function or the vehicle-mounted system control does not have a dependent relationship and does not have a conflict relationship with the intersection sound area, the merging processing of the sound area is performed, including: Figure 4
[0062] S231: The same sound area of the reply is de-duplicated, and the corresponding vehicle function or vehicle-mounted system control is spliced after the sound area. That is to say, after the same sound area of the reply is de-duplicated, the corresponding vehicle function or vehicle-mounted system control can be spliced after the sound area, and the splicing order can be set according to the order of the reply. Thus, the corresponding vehicle function or vehicle-mounted system control can be replied according to the order of the instruction, improving the satisfaction of the user. For example, when the user's parallel instruction is "the driver seat heating is turned on, and the driver seat ventilation is turned on", the corresponding reply is "the driver seat heating is turned on, and the driver seat ventilation is turned on", at this time, the sound area can be fused, that is, "the driver seat heating" and "the driver seat ventilation" are combined into "the driver seat heating and ventilation".
[0063] In the present application, as shown in S24, based on the operation type pointed by the reply after the merging processing of the sound area, the fusion is performed, including: Figure 5
[0064] S241: In the case where the operation types of the replies after the merging processing of the sound area are the same, the replies are merged while keeping the operation type unchanged. That is to say, after the sound area merging processing, if the operation types are the same, such as both are opening or both are closing, the replies can be merged while keeping the operation type unchanged.
[0065] For example, when the user's parallel instruction is "main driver seat heating is on, the heating of the co-driver seat is on, and the right rear seat ventilation is on", the corresponding reply is "the main driver seat heating is on, the co-driver seat heating is on, and the right rear seat ventilation is on", at this time, the operation types are all opening, and "the main driver seat heating is on", "the co-driver seat heating is on", and "the right rear seat ventilation is on" can be fused into "the front seat heating and the right rear seat ventilation are both on".
[0066] S242: In the case where the operation types of the replies after the merging processing of the sound areas are different, the replies are merged in the case where the operation types are changed into the generic type expression. That is, in the case where the operation types are different after the sound area merging processing, the replies can be unified by using a more general reply, such as "good" or "finished". Thus, the unified reply of the operation types is realized, which is beneficial to reducing the length of the voice broadcast.
[0067] For example, when the user's parallel instruction is "main driver seat heating is on, the heating of the co-driver seat is on, the right rear window is closed, and the air conditioner is set to 28 degrees", the corresponding reply is "the main driver seat heating is on, the co-driver seat heating is on, the right rear window is closed, and the air conditioner is set to 28 degrees", at this time, the operation types are different, and "the main driver seat heating is on" and "the co-driver seat heating is on", "the right rear window is closed", and "the air conditioner is set to 28 degrees" can be fused into "the front seat heating, the right rear window, and the air conditioner are all set".
[0068] In the present application, as shown in Figure 6 The method further includes:
[0069] S25: In the case where the operation types of the replies of the vehicle functions or the vehicle-mounted system controls after the merging processing of the vehicle functions or the vehicle-mounted system controls are not opening, the reply of the vehicle functions or the vehicle-mounted system controls having a dependent relationship is issued to the vehicle, so that the vehicle performs splicing broadcast.
[0070] That is, in the case where the operation types of the replies of the vehicle functions or the vehicle-mounted system controls after the merging of the vehicle functions or the vehicle-mounted system controls are not opening, the reply of the vehicle functions or the vehicle-mounted system controls having a dependent relationship can be directly issued to the vehicle, so that the vehicle can directly perform splicing broadcast. Thus, the user can obtain accurate feedback.
[0071] For example, when the user's parallel instruction is "temperature is adjusted to 26 degrees, air conditioner is turned off, and front row window is opened", although there is a dependency relationship between the air conditioner and the temperature adjustment, the implementation of multiple instructions does not interfere, at this time, "temperature is adjusted to 26 degrees, air conditioner is turned off, and front row window is opened" can be directly issued to the vehicle, so that the vehicle can execute the instructions in turn and splice and broadcast.
[0072] In the present application, as shown in Figure 7 The method further comprises:
[0073] S26: In the reply after the merging of the vehicle functions or the vehicle-mounted system controls, the reply of the vehicle functions or the vehicle-mounted system controls with a dependency relationship, in the case of the operation type being opening, a first voice broadcast is issued to the vehicle for broadcasting by the vehicle, and the first voice broadcast is used to indicate the dependency relationship.
[0074] That is, in the reply after the merging of the vehicle functions or the vehicle-mounted system controls, if there is a reply of the vehicle functions or the vehicle-mounted system controls with a dependency relationship, in the case of the operation type being opening, a first voice broadcast can be generated, and the first voice broadcast is used to indicate the dependency relationship, such as "vehicle-mounted system control A is opened, and helps to complete function B". In this way, the first voice broadcast can be issued to the vehicle, so that the vehicle can express the dependency relationship when replying. Thus, the perception of the user is improved.
[0075] For example, when the user's parallel instruction is "open the main screen and adjust the main screen to the brightest", the corresponding reply is "the main screen is opened and the main screen is adjusted to the brightest", at this time, since the main screen needs to be opened first and then the brightness of the main screen is adjusted, there is a dependency relationship between the opening of the main screen and the adjustment of the brightness of the main screen, and a first broadcast voice can be generated, which is "the main screen is opened, and helps to adjust the brightness of the main screen to the best".
[0076] In the present application, as shown in Figure 8 The method further comprises:
[0077] S27: In the reply after the merging of the vehicle functions or the vehicle-mounted system controls, the reply of the vehicle functions or the vehicle-mounted system controls without a dependency relationship, in the case of the vehicle functions or the vehicle-mounted system controls having a conflict relationship and the corresponding audio zones having an intersection, the vehicle function or the vehicle-mounted system control behind is overlaid on the vehicle function or the vehicle-mounted system control in front, and a second voice broadcast is issued to the vehicle for broadcasting by the vehicle, and the second voice broadcast is used to indicate the conflict relationship and the execution state of the vehicle function or the vehicle-mounted system control behind.
[0078] That is, when the merging processing of the vehicle functions or the vehicle system controls is completed, in the reply without the dependency relationship, if there is a vehicle function or a conflict relationship and the corresponding audio area has an intersection, the latter vehicle function or vehicle system control covers the former vehicle function or vehicle system control, that is, the vehicle only executes the latter vehicle function or vehicle system control. At this time, the second voice broadcast can be generated, and the second voice broadcast can express the conflict relationship and indicate that the latter vehicle function or vehicle system control is executed, such as A function and B function cannot be operated at the same time, help you execute B function. Therefore, by issuing the second voice broadcast to the vehicle, the vehicle can express the conflict relationship when replying, so as to improve the perception of the user.
[0079] For example, when the user's parallel instruction is "turn on the air conditioner, set the mirror image wind for the driver, and turn on the front defogger", the corresponding reply is "the air conditioner is turned on, the mirror image wind for the driver is turned on, and the front defogger is turned on". Since there is a conflict between the mirror image wind and the front defogger, the vehicle only executes the latter operation, that is, turns on the front defogger. At this time, the second broadcast voice can be generated, and the second broadcast voice is "the air conditioner has been turned on, the mirror image wind and the front defogger cannot be operated at the same time, and help you turn on the front defogger".
[0080] In the specific working process, as shown in Figure 9 When the multiple replies corresponding to the multiple parallel instructions are obtained, the vehicle functions or vehicle system controls in the multiple replies can be judged. If there are the same vehicle functions or vehicle system controls, the same vehicle functions or vehicle system controls can be merged. Then, the dependency relationship of the vehicle functions or vehicle system controls in the multiple replies is judged. If there is a dependency relationship, it is further judged whether the operation type in the dependency relationship is opening. If yes, the corresponding reply is generated as the first voice broadcast. If not, the corresponding reply is spliced and broadcast. Then, the conflict relationship in the reply without the dependency relationship is judged. If there is a conflict relationship and the corresponding audio area has an intersection, the corresponding reply is generated as the second voice broadcast. Next, the audio area in the reply without the dependency relationship and the conflict relationship with the intersection audio area is judged. If there is the same audio area, the same audio area is removed. If there are the same adjacent audio areas of the vehicle functions or vehicle system controls, the direction words can be fused to reduce the number of audio areas. Then, the operation type of the reply after the merging processing of the audio area can be judged. If the operation types are the same, the operation types remain unchanged. If the operation types are different, the operation types are changed to a general direction class expression to generate a voice broadcast. Finally, the generated voice broadcast can be issued to the vehicle.
[0081] The voice interaction method of the present application will be described below according to some examples.
[0082] When the user's parallel commands are "Driver's seat ventilation on, front seat ventilation on, driver's side window open, air conditioning on", the corresponding response is "Driver's seat ventilation is on, front seat ventilation is on, driver's side window open, air conditioning is on for you". At this time, they can be merged to generate the voice broadcast "Air conditioning, front seat ventilation and driver's side window are all on".
[0083] When a user's parallel command is "Turn on the front seat heating, turn on the right rear seat heating, close the driver's side window, and turn the ambient lighting to purple," the corresponding response is "The front seat heating is on, the right rear seat heating is on, the driver's side window is closed, and the ambient lighting is adjusted for you." At this point, the commands can be merged to generate the voice announcement "The ambient lighting, the front and right rear seat heating, and the driver's side window are all adjusted."
[0084] When the user's parallel commands are "Driver's seat ventilation open, driver's side window open, intelligent suspension adjustment open, ambient light open", the corresponding response is "Driver's seat ventilation is open, driver's side window is open, intelligent suspension adjustment is open, ambient light open". At this time, they can be merged to generate the voice broadcast "Intelligent suspension adjustment, ambient light, driver's seat ventilation and window are all open".
[0085] When a user's parallel command is "Raise the driver's seat heating, close the rear windows, move the driver's seat all the way back, and turn on the air conditioning," the corresponding response is "The driver's seat heating has been raised, the rear windows are closed, the driver's seat is all the way back, and the air conditioning has been turned on for you." At this point, the commands can be merged to generate a voice announcement that "The air conditioning, driver's seat heating, seat position, and rear windows are all adjusted."
[0086] When the user's parallel command is "Adjust the driver's seat back, raise the passenger seat heater, set the air conditioning to 28 degrees, and raise all windows," the corresponding response is "The driver's seat is adjusted, the passenger seat heater is raised, the air conditioning temperature is set to 28 degrees, and all windows are raised." At this point, the commands can be merged to generate a voice announcement that "The temperature, driver's seat, passenger seat heater, and all windows are adjusted."
[0087] When a user's parallel command is "Adjust the driver's seat back, raise the passenger seat heater, turn on the air conditioning, and close all windows," the corresponding response is "The driver's seat is adjusted, the passenger seat heater is raised, the air conditioning is turned on, and all windows are closed." At this point, the commands can be merged to generate a voice announcement that says "The air conditioning, driver's seat, passenger seat heater, and all windows are adjusted."
[0088] When a user's parallel command is "Set the driver's side to mirror mode and open the air blower", the corresponding response is "The driver's side is set to mirror mode and the air blower is open". At this time, there is a conflict in the response, so a second voice message can be generated, namely "Mirror mode and air blower cannot be opened at the same time. I'll open the air blower for you".
[0089] When a user's parallel command is "Open the home screen, brighten the home screen, turn the home screen to the brightest," the corresponding response is "The home screen is open, the home screen is brightened, the home screen is turned to the brightest." At this time, there is a dependency relationship in the response, which can generate the first voice broadcast, namely, "The home screen is open, and I have adjusted the brightness of the home screen for you."
[0090] When the user's parallel command is "Turn on the air conditioning, set the driver's side to mirror mode, and turn on the front defroster", the corresponding response is "The air conditioning is on, the driver's side mirror mode is on, and the front defroster is on". At this time, there is a dependency relationship in the response, which can generate the first voice broadcast, so that the voice broadcast content is "The air conditioning is on. Mirror mode and front defroster cannot be operated at the same time. I've turned on the front defroster for you."
[0091] It should be noted that the voice interaction method provided in this application can be executed by a voice interaction device or a processing module in the voice interaction device for executing the voice interaction method.
[0092] The present invention also proposes a voice interaction device.
[0093] The voice interaction device according to the present invention includes: a receiving module, a processing module, and a sending module;
[0094] The receiving module is used to receive multiple parallel commands from the user in the cockpit relayed by the vehicle.
[0095] The processing module is used to generate a voice broadcast that integrates the responses to multiple parallel commands, based on the vehicle function or in-vehicle system control pointed to in each command, the operation type of the command, and the source audio region of the command.
[0096] The sending module is used to send voice broadcasts to the vehicle so that the vehicle can provide feedback to the user based on the voice broadcasts.
[0097] According to the voice interaction device of the present invention, by receiving and recognizing multiple parallel commands from the user in the cockpit, a voice broadcast that integrates the responses corresponding to the multiple parallel commands is generated to provide the user with a short and clear reply. This eliminates the need to reply to each of the user's multiple commands one by one, reducing the user's waiting time. Moreover, the reply content is accurate and clear, allowing the user to receive clear feedback, which helps to improve user satisfaction and enhances the practicality of the voice interaction device.
[0098] In this invention, the processing module is further configured to acquire multiple responses corresponding to multiple parallel instructions; merge responses pointing to the same vehicle function or the same in-vehicle system control; merge responses after merging vehicle functions or in-vehicle system controls where there is no dependency between the vehicle functions or in-vehicle system controls and no conflicting relationship with overlapping audio regions; and fuse responses based on the operation type pointed to by the responses after merging audio regions to obtain voice broadcast.
[0099] In this invention, the processing module is also used to deduplicate responses with the same vehicle function or in-vehicle system control, and to take the union of the sound regions pointed to by responses with the same vehicle function or in-vehicle system control; wherein, in the case of adjacent sound regions, adjacent sound regions are merged by locative words to reduce the number of sound regions without changing the range of the sound regions.
[0100] In this invention, the processing module is also used to deduplicate replies with the same audio range and to append the corresponding vehicle function or in-vehicle system control after the audio range.
[0101] In this invention, the processing module is also used to merge the responses while keeping the operation type unchanged when the operation types of the responses after merging the sound regions are the same; and to merge the responses by changing the operation type to a generic expression when the operation types of the responses after merging the sound regions are different.
[0102] In this invention, the processing module is also used to send the reply to the vehicle if there is a dependency relationship between the vehicle function or the vehicle system control in the reply after merging the vehicle function or the vehicle system control, when the operation type is not open, so that the vehicle can splice and broadcast the reply.
[0103] In this invention, the processing module is also used to send a first voice broadcast to the vehicle when the operation type is "on" in the response after merging the vehicle functions or in-vehicle system controls, so that the vehicle can broadcast the response. The first voice broadcast is used to indicate the dependency relationship.
[0104] In this invention, the processing module is also used to handle responses after merging vehicle functions or in-vehicle system controls where there is no dependency between the vehicle functions or in-vehicle system controls. In cases where there is a conflict between vehicle functions or in-vehicle system controls and their corresponding voice regions overlap, the module overwrites the preceding vehicle functions or in-vehicle system controls with the subsequent vehicle functions or in-vehicle system controls and sends a second voice broadcast to the vehicle so that the vehicle can broadcast the message. The second voice broadcast is used to indicate the conflict relationship and the execution status of the subsequent vehicle functions or in-vehicle system controls.
[0105] According to the voice interaction device of the present invention, by receiving and recognizing multiple parallel commands from the user in the cockpit, a voice broadcast that integrates the responses corresponding to the multiple parallel commands is generated to provide the user with a short and clear reply. This eliminates the need to reply to each of the user's multiple commands one by one, reducing the user's waiting time. Moreover, the reply content is accurate and clear, allowing the user to receive clear feedback, which helps to improve user satisfaction and enhances the practicality of the voice interaction device.
[0106] The voice interaction device in this invention can be a device, or a component, integrated circuit, or chip in a mobile terminal. For example, the mobile terminal can be a mobile phone, tablet computer, laptop computer, PDA, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc.
[0107] The voice interaction device in this invention can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this invention does not specifically limit it.
[0108] The voice interaction device provided by this invention can achieve Figures 1 to 8 To avoid repetition, the various processes implemented by the voice interaction device in the voice interaction method will not be described in detail here.
[0109] This invention also proposes a server.
[0110] The server according to the present invention includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements any of the methods described above.
[0111] According to the server of the present invention, by executing the computer program stored in the memory through the processor, a voice broadcast that integrates the responses corresponding to multiple parallel instructions can be generated to provide a short and clear reply to the user. This eliminates the need to reply to each of the user's multiple instructions one by one, reducing the user's waiting time. Moreover, the reply content is accurate and clear, allowing the user to receive clear feedback and improving user satisfaction.
[0112] The present invention also proposes a non-volatile computer-readable storage medium for computer programs.
[0113] Reference Figure 10As shown, the non-volatile computer-readable storage medium 100 of the computer program 101 according to the present invention implements any of the methods described above when the computer program 101 is executed by one or more processors 200.
[0114] The processor mentioned above is the processor in the electronic device described in the example above. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0115] According to the present invention, a non-volatile computer-readable storage medium for a computer program can generate a voice broadcast that integrates responses to multiple parallel instructions by executing the computer program stored on the storage medium through a processor. This provides a brief and clear response to the user without having to respond to each of the user's multiple instructions individually, reducing the user's waiting time. Furthermore, the response content is accurate and clear, allowing the user to receive clear feedback and improving user satisfaction.
[0116] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0117] In the description of this invention, "first feature" and "second feature" may include one or more of the features.
[0118] In the description of this invention, "a plurality of" means two or more.
[0119] In the description of this invention, the first feature being "above" or "below" the second feature may include the first and second features being in direct contact, or it may include the first and second features not being in direct contact but being in contact through another feature between them.
[0120] In the description of this invention, the terms "above," "over," and "on top" for the first feature and the second feature include the first feature being directly above or diagonally above the second feature, or simply indicating that the first feature is at a higher horizontal level than the second feature.
[0121] In the description of this specification, references to terms such as "a voice interaction method," "some voice interaction methods," "illustrative voice interaction method," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that voice interaction method or example is included in at least one voice interaction method or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same voice interaction method or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more voice interaction methods or examples.
[0122] Although the voice interaction methods of the present invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and variations can be made to these voice interaction methods without departing from the principles and spirit of the present invention, the scope of which is defined by the claims and their equivalents.
Claims
1. A voice interaction method, characterized in that, The method includes: Receive multiple parallel commands from the user in the cockpit forwarded by the vehicle; Based on the vehicle function or in-vehicle system control pointed to in each instruction, the operation type of the instruction, and the directional tone range, a voice broadcast that integrates the responses corresponding to multiple parallel instructions is generated. The voice broadcast is sent to the vehicle so that the vehicle can provide feedback to the user based on the voice broadcast; The process of generating a voice broadcast that integrates responses to multiple parallel commands, based on the vehicle function or in-vehicle system control pointed to in each command, the operation type of the command, and the audio region pointed to by the command, includes: Obtain multiple responses corresponding to multiple parallel instructions; For responses that point to the same vehicle function or the same in-vehicle system control, merge the vehicle function or in-vehicle system control. For responses that, after merging vehicle functions or in-vehicle system controls, do not have dependencies on vehicle functions or in-vehicle system controls and do not have conflicting relationships with overlapping audio regions, audio region merging processing will be performed. Based on the operation type of the response after merging the sound regions, the voice is fused to obtain the voice broadcast; The overlapping sound zone refers to the sound zone corresponding to the same seat, and the dependency relationship refers to the on / off state of the vehicle system control and the adjustable function after the vehicle system control is turned on.
2. The voice interaction method according to claim 1, characterized in that, The process of merging vehicle functions or vehicle system controls in responses that point to the same vehicle function or the same in-vehicle system control includes: Deduplicat responses with the same vehicle function or in-vehicle system control, and take the union of the sound zones pointed to by responses with the same vehicle function or in-vehicle system control; wherein, in the case of adjacent sound zones, adjacent sound zones are merged by locative words to reduce the number of sound zones without changing the range of the sound zones, wherein the adjacent sound zones are the sound zones corresponding to adjacent seats.
3. The voice interaction method according to claim 1, characterized in that, The process of merging the audio regions of responses where there are no dependencies between vehicle functions or in-vehicle system controls and no conflicts with overlapping audio regions after merging vehicle function or in-vehicle system controls includes: Remove duplicate responses with the same audio range and append the corresponding vehicle function or in-vehicle system control after that audio range.
4. The voice interaction method according to claim 1, characterized in that, The fusion based on the operation type of the response pointer after merging the sound regions includes: If the operation types of the responses after merging the audio regions are the same, merge the responses while keeping the operation types unchanged; If the operation types of the responses after merging the sound regions are different, the responses will be merged if the operation type is changed to a generic expression.
5. The voice interaction method according to claim 1, characterized in that, The method further includes: For responses that have dependencies on vehicle functions or in-vehicle system controls after merging vehicle functions or controls, if the operation type is not "open", the response is sent to the vehicle so that the vehicle can splice and broadcast it.
6. The voice interaction method according to claim 1, characterized in that, The method further includes: For responses that involve dependencies between vehicle functions or in-vehicle system controls after merging vehicle functions or controls, if the operation type is "on", a first voice broadcast is sent to the vehicle so that the vehicle can broadcast the message. The first voice broadcast is used to indicate the dependency.
7. The voice interaction method according to claim 1, characterized in that, The method further includes: In responses to vehicle functions or in-vehicle system controls that do not have dependencies after merging the vehicle functions or controls, if there are conflicts between the vehicle functions or in-vehicle system controls and their corresponding voice registers overlap, the subsequent vehicle function or in-vehicle system control overwrites the previous one, and a second voice broadcast is sent to the vehicle so that the vehicle can broadcast the message. The second voice broadcast is used to indicate the conflict relationship and the execution status of the subsequent vehicle function or in-vehicle system control.
8. A voice interaction device, implementing the method according to any one of claims 1-7, characterized in that, include: The receiving module is used to receive multiple parallel commands from the user in the cockpit forwarded by the vehicle. The processing module is used to generate a voice broadcast that integrates the responses to multiple parallel commands based on the vehicle function or in-vehicle system control pointed to in each command, the operation type of the command, and the source audio region of the command. The sending module is used to send the voice broadcast to the vehicle so that the vehicle can provide feedback to the user based on the voice broadcast.
9. A server, characterized in that, The server includes a memory and a processor, the memory storing a computer program that, when executed by the processor, implements the method according to any one of claims 1-7.
10. A non-volatile computer-readable storage medium for a computer program, characterized in that, When the computer program is executed by one or more processors, it implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Voice interaction method, vehicle and storage medium
CN114898752A