Method and system for supporting multi-session mode of vehicle and recording medium

By using surround sound and sound directionality in vehicles, and allocating sound space and outputting speech based on user information, the problem that traditional vehicle speech recognition devices cannot distinguish the voice output of multiple users is solved, and conversation partners can be easily identified and speech can be output in a directional manner.

CN113851122BActive Publication Date: 2026-03-27HYUNDAI MOTOR CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-01
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Traditional vehicle speech recognition devices cannot effectively distinguish the voice output of multiple users while driving, making it difficult to identify conversation partners in teleconferences.

Method used

By utilizing surround sound and sound directionality functions, sound space is allocated based on user information, and speech is output to the corresponding space, thereby achieving directional allocation and output of speech.

Benefits of technology

It enables easy identification of conversation partners in multi-conversation mode, ensuring that the speech output corresponds to the user, and improving the interaction efficiency of drivers in teleconferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113851122B_ABST
    Figure CN113851122B_ABST
Patent Text Reader

Abstract

The disclosure relates to a method and system for supporting a multi-session mode of a vehicle and a recording medium. The method for supporting a multi-session mode of a vehicle of the disclosure can include receiving user information of a multi-session mode and at least one of a message and an utterance from a session partner participating in the multi-session mode, allocating a sound space based on the user information, and allocating directionality to an utterance generated based on at least one of the message and the utterance and outputting the utterance to the allocated space.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to a method and system for supporting a multi-conversation mode of a vehicle. BACKGROUND

[0002] Generally, a speech recognition device for a vehicle allows a conference call using a multi-conversation mode as well as a 1:1 conversion, and also allows multiple users to transmit / receive text through a messenger application.

[0003] However, in a conventional speech recognition device, when a user on a driver's seat calls a speech recognition system and initiates a teleconference, when speech is output with respect to text between multiple users through a messenger application during driving, the speech is output without distinguishing the users, and thus the users cannot be easily recognized during the teleconference. SUMMARY

[0004] An object of the disclosure is to provide a method and system for supporting a multi-conversation mode of a vehicle, which provides a function of easily recognizing a conversation partner by outputting sound with directionality through a surround sound function and a sound directionality function during a voice call between multiple users.

[0005] Those skilled in the art will recognize that the disclosure can achieve objects other than those specifically described herein, and that the disclosure can achieve the above and other objects in the manner set forth herein or as modified in any manner by the skilled practitioner, and the above and other objects realized from the following detailed description.

[0006] To achieve these objects and other advantages and in accordance with the purpose of the disclosure, as embodied and broadly described herein, a method for supporting a multi-conversation mode of a vehicle can include receiving user information of a multi-conversation mode and at least one of a message and speech from a conversation partner participating in the multi-conversation mode, allocating a sound space based on the user information, and allocating directionality to speech generated based on at least one of the message and the speech and outputting the speech to the allocated space.

[0007] According to an embodiment, the user information can include at least one of a number, a gender, an age, an ID, and a name of the conversation partner.

[0008] According to an embodiment, the allocating a sound space based on the user information can include allocating a sound space in response to the number of the conversation partner.

[0009] According to an embodiment, the method can further include grouping the conversation partners based on the user information when the number of the conversation partners is greater than the number of the allocated spaces, and allocating the generated groups to the sound spaces.

[0010] According to an embodiment, assigning the directionality to the utterance and outputting the utterance to the assigned space can include arranging a conversation partner corresponding to a message in a farthest space among the assigned spaces based on a message reception order, and assigning the directionality to an utterance corresponding to a message received from the conversation partner and outputting the utterance.

[0011] According to an embodiment, assigning the directionality to the utterance corresponding to the message received from the conversation partner and outputting the utterance can include setting a predetermined utterance output time for the assigned space, determining whether the utterance output corresponding to the received message is within the predetermined utterance output time, excluding special characters and modifiers from the message when the utterance output is not within the predetermined utterance output time, determining whether output of the message excluding the special characters and modifiers as the utterance is within the predetermined utterance output time, controlling an utterance output speed so that the utterance output speed becomes a predetermined value when the output of the message excluding the special characters and modifiers as the utterance is not within the predetermined utterance output time, and determining whether the utterance output is within the predetermined utterance output time, and controlling the message so that the message is output in an overlapped manner to the sound space when the utterance output according to the utterance output speed control is not within the predetermined utterance output time.

[0012] According to an embodiment, determining whether the utterance output corresponding to the received message is within the predetermined utterance output time can include determining whether a number of sentences of the received message is equal to or greater than a predetermined value when the utterance output is not within the predetermined utterance output time, and dividing an utterance output interval when the number of sentences is equal to or greater than the predetermined value.

[0013] According to an embodiment, the method can further include determining a face direction of the driver through a camera when the driver speaks, determining whether the face direction of the driver coincides with the directionality of the sound, and identifying a conversation partner corresponding to a corresponding sound space when the face direction of the driver coincides with the directionality of the sound.

[0014] According to an embodiment, the method can further include receiving an utterance of the driver when the driver speaks to the conversation partner, and transmitting a message or the utterance of the driver to the conversation partner.

[0015] According to the method and system for supporting a multi-conversation mode of a vehicle according to the disclosure, a conversation partner can be easily identified, and a specific partner can be selected for a conversation in the multi-conversation mode.

[0016] In addition, a text transmitted from a plurality of users is converted into an utterance corresponding to a corresponding user or an utterance to which directionality is assigned, so that a user who listens to the utterance can easily identify a user who transmits the text in the multi-conversation mode.

[0017] Those skilled in the art will appreciate that effects which can be achieved by using the present disclosure are not limited to what has been particularly described hereinabove and other advantages of the present disclosure will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this application, illustrate embodiment(s) of the present disclosure and together with the description serve to explain the principles of the present disclosure. In the drawings:

[0019] Figure 1A is a block diagram of a system for supporting a multi-session mode of a vehicle according to an embodiment of the present disclosure;

[0020] Figure 1B is a block diagram of a system for supporting a multi-session mode of a vehicle according to an embodiment of the present disclosure;

[0021] Figure 2 is a diagram illustrating an example of a method for supporting a multi-session mode using directivity of sound according to an embodiment of the present disclosure;

[0022] Figure 3 is a flowchart of a method of supporting a multi-session mode according to an embodiment of the present disclosure;

[0023] Figure 4 is a diagram for describing a spatial allocation method using directivity of sound according to an embodiment of the present disclosure; and

[0024] Figure 5 is a diagram illustrating a voice output method in a method for supporting a multi-session mode according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] Hereinafter, devices and various methods to which embodiments of the present disclosure are applied will be described in greater detail with reference to the accompanying drawings. In the entire description, the terms "module" and "part" are used for convenience of description and thus can be used interchangeably, and do not have any distinguishable meaning or function.

[0026] In the description of the embodiments, it will be understood that, when a component is referred to as being "on" or "under" another component, it can be directly on the other component, or an intermediate component can also be present.

[0027] It will be understood that, although the terms first, second, A, B, (a), (b), etc. can be used herein to describe various elements, these terms are merely used to distinguish one element from another, and are not necessarily used to describe a sequence or ordering among the elements. It will be understood that, when an element is referred to as being "connected to" or "accessed" another element, the element can be directly connected to or directly accessed by the other element, or the element can be connected to or accessed by the other element via another element.

[0028] In addition, unless explicitly defined otherwise, terms such as "include", "comprise", and "have" should be construed as inclusive or open-ended, rather than exclusive or closed-ended. Unless explicitly defined otherwise, all technical, scientific or other aspects of terms are consistent with the meaning understood by those skilled in the art. Unless explicitly defined otherwise in the disclosure, commonly used terms found in dictionaries should be interpreted in the context of the relevant technical writings without being too idealistic or unrealistic.

[0029] Figure 1A is a block diagram of a system for supporting a multi-session mode of a vehicle according to an embodiment of the disclosure.

[0030] Referring to Figure 1A , the system for supporting a multi-session mode of a vehicle can include a data receiver 110, a user classifier 120, a speech direction assigner 130, a speech modulator 140, and a speech output unit 150.

[0031] The data receiver 110 can receive at least one of a message and a speech from a session partner through a communication application for implementing a multi-session mode. The data receiver 110 can receive a speech of a session partner during a plurality of sessions through a telephone call. The data receiver 110 can receive a message from a session partner during a plurality of sessions. The data receiver 110 can receive user information of a multi-session mode. In one example, the data receiver 110 can include an antenna or a transceiver configured to perform wireless / wired communication to receive data from a server or an external device, for example, a mobile device including a mobile phone.

[0032] The user classifier 120 can classify a session partner of an output message through circuitry or protocol analysis. Accordingly, the user classifier 120 can classify a user based on user information obtained from a telephone number / session message protocol. Here, the user information can include at least one of the number of session partners, gender, age, ID, and name.

[0033] The speech direction assigner 130 can assign a sound space based on the user information.

[0034] To this end, the utterance direction assigner 130 can divide a sound space in response to a user classified by the user classifier 120. Then, the utterance direction assigner 130 can assign a direction in response to the classified user.

[0035] When an utterance corresponding to text is output, the utterance modulator 140 can modulate the utterance using user information so that the utterance corresponds to a male / female voice or a user age voice of a corresponding user, and output the modulated utterance so that a conversation partner can be easily recognized. To this end, the utterance modulator 140 can be used for utterance modulation suitable for outputting an utterance using user information in the case of text. The utterance modulator 140 can be selectively set ON / OFF.

[0036] The utterance output unit 150 can output the utterance with directionality to the assigned space.

[0037] The utterance output unit 150 can assign directionality to an utterance generated based on at least one of a message and an utterance received from the data receiver 110, and output the utterance with directionality to a space assigned by the utterance direction assigner 130. In an example, the utterance output unit 150 can include one or more speakers. The utterance output unit 150 can be configured to adjust respective volumes of the one or more speakers, for example, based on directionality of a space assigned by the utterance direction assigner 130.

[0038] Here, the utterance can be output so that a direction of an output conversation partner's utterance corresponds to a direction of a sound space assigned by the utterance direction assigner 130.

[0039] Figure 1B is a block diagram of a system for supporting a multi-conversation mode of a vehicle according to an embodiment of the disclosure.

[0040] Referring to Figure 1B , the system for supporting a multi-conversation mode of a vehicle can include a conversation partner recognizer 160, an utterance receiver 170, and a data transmitter 180.

[0041] The conversation partner recognizer 160 can determine, through a camera, a space that a driver faces for a conversation with a user from a space assigned by the utterance direction assigner 130. When a direction of the driver's face coincides with directionality of an utterance, the conversation partner recognizer 160 can recognize a conversation partner corresponding to a sound space corresponding to the utterance. In an example, the direction of the driver's face coinciding with the directionality of the utterance can mean that the driver's face faces the sound space corresponding to the utterance.

[0042] Accordingly, the conversation partner recognizer 160 can recognize the conversation partner of the driver among the users of the communication application through the camera.

[0043] When the driver outputs the utterance to the conversation partner, the utterance receiver 170 can receive the utterance of the driver. In one example, the utterance receiver 170 can include a microphone configured to receive voice data and convert the voice data into a message or an utterance. In one example, the message or the utterance can be in the form of text after conversion through the utterance receiver 170.

[0044] The data transmitter 180 can confirm the conversation partner through the conversation partner recognizer 160 and the utterance receiver 170. Here, the data transmitter 180 can optionally check whether the selected conversation partner is correct.

[0045] The data transmitter 180 can transmit the message or the utterance of the driver to the confirmed conversation partner. In one example, the data transmitter 180 can include an antenna or a transceiver configured to perform wireless / wired communication to transmit data to a server or an external device (for example, a mobile device including a mobile phone).

[0046] Figure 2 FIG. 1 is a diagram illustrating an example of a method for supporting a multi-conversation mode using directivity of sound according to an embodiment of the disclosure.

[0047] Referring to Figure 2 , the data receiver 110 can receive information about users participating in a multi-conversation mode or a chat room from the communication application 210.

[0048] The utterance direction allocator 130 can allocate the sound space 220 based on the user information. The utterance direction allocator 130 can allocate the sound space 220 in response to the number of conversation partners included in the user information.

[0049] Although five sound spaces 220 of a sector are allocated based on the driver 10 and the conversation partners are allocated to each space in this diagram, the shape of the sound space 220 and the number of sound spaces 220 are not limited thereto.

[0050] When the number of conversation partners is greater than the number of allocated spaces, the utterance direction allocator 130 can group the conversation partners based on the user information received from the user classifier 120 and allocate the group to the sound space 220.

[0051] Although two conversation partners are grouped based on the user information and the group is allocated to one of the five allocated spaces in this diagram, the number of groups and the number of conversation partners in the group are not limited thereto.

[0052] The utterance output unit 150 can output an utterance to a space allocated by the utterance direction allocator 130 in response to a message and an utterance received from a plurality of users through a communication application.

[0053] When a sound is output to the inside of the vehicle, the utterance output unit 150 can output a sound having directivity 230 to a space allocated by the utterance direction allocator 130. Here, the utterance output unit 150 can output a sound having directivity 230 using a surround sound effect.

[0054] When a message is received from a conversation partner, the utterance output unit 150 can allocate directivity to an utterance corresponding to the received message and output the utterance having directivity.

[0055] When many messages are received at the same time, the utterance output unit 150 can allocate a predetermined utterance output time to each space allocated by the utterance direction allocator 130 in order to output an utterance corresponding to a message. According to an embodiment, the utterance output unit 150 can allocate 3 seconds for each allocated space.

[0056] When the utterance output unit fails to output an utterance to each allocated space within a predetermined utterance output time, the utterance output unit 150 can control the utterance output speed. To this end, the utterance output unit 150 can increase the utterance output speed. When the utterance output unit 150 fails to output an utterance even after increasing the utterance output speed, the utterance output unit 150 can simultaneously output an utterance to the sound space 220 in a superimposed manner.

[0057] If the utterance output unit 150 fails to output an utterance to each allocated space within a predetermined utterance output time, the utterance output unit 150 can not output prosodic words, modifiers, etc. of a message.

[0058] When the number of sentences of a message received from a single conversation partner is equal to or greater than a predetermined value, the utterance output unit 150 can stop utterance output after a predetermined utterance output time elapses and output an utterance to another allocated space. Thereafter, when the utterance output to the other allocated space ends, the utterance output unit 150 can continue to output the stopped message as an utterance. Here, the predetermined number of sentences of a message can be 2 or more.

[0059] When there is no message from other conversation partners for a set time, the utterance output unit 150 can extend utterance output in an allocated space in which an utterance is currently output for a predetermined utterance output time.

[0060] When a phone call is performed through the multi-session mode, the utterance output unit 150 can allocate a sound space 220 corresponding to a circuit participating in a session, and control an utterance to be output to the allocated space corresponding to a phone call circuit.

[0061] Meanwhile, when a phone call is performed through the multi-session mode, the utterance output unit 150 can allocate a sound space 200 to a session partner of a phone call in advance, and when a call from a user corresponding to the allocated space is received, the utterance output unit 150 can output a ring tone to the allocated space.

[0062] The session partner recognizer 160 can recognize a face direction 250 of the driver detected through the camera 240. The session partner recognizer 160 can determine whether the face direction 250 of the driver coincides with a directivity 230 of a sound with respect to a space allocated by the utterance direction allocator 130.

[0063] When the face direction 250 of the driver coincides with the directivity 230 of the sound, the session partner recognizer 160 can select a session partner corresponding to a space 220 corresponding to the directivity 230 of the sound.

[0064] Then, the utterance output unit 150 can output an utterance to the allocated space based on the utterance of the session partner selected by the session partner recognizer 160.

[0065] Figure 3 is a flowchart of a method of supporting a multi-session mode according to an embodiment of the disclosure.

[0066] Referring to Figure 3 , the data receiver 110 can acquire information about users participating in a multi-session mode (S310).

[0067] After step S310, the utterance direction allocator 130 can allocate a sound space 200 in a vehicle based on the obtained user information (S320). Here, when it is determined based on the user information that the number of session partners of the multi-session mode is greater than the number of allocated spaces, the utterance direction allocator 130 can group the multi-session mode users based on the user classification information received by the user classifier 120, and allocate the group to the allocated space.

[0068] After step S320, the data receiver 110 can receive a message or an utterance from a session partner participating in the multi-session mode (S330).

[0069] After step S330, the utterance output unit 150 can allocate a directivity 230 of a sound to a message or an utterance corresponding to each session partner, and output it similarly to the utterance.

[0070] After step S320, when the driver speaks, the conversation partner identifier 160 can identify the direction in which the driver is facing and conversing when the driver speaks through the camera, and when the direction in which the driver is facing and conversing corresponds to the space allocated by the speech direction allocator 130, the conversation partner identifier 160 can select the conversation partner corresponding to that space (S350).

[0071] After step S350, the speech receiver 170 can receive the driver's speech, and the data transmitter 180 can send the message or the driver's speech to the selected conversation partner (S360).

[0072] Figure 4 This is a diagram illustrating a spatial allocation method using the directionality of sound according to embodiments of the present disclosure.

[0073] Reference Figure 4 When conversation partners are assigned to the allocated spaces, the speech direction allocator 130 can allocate the sound spaces 220 in multi-conversation mode according to the order in which messages are received. Here, spaces can be allocated such that, based on the order in which messages are received, the conversation partner corresponding to the message is assigned to the furthest space in the allocated spaces, so that the driver can easily identify speech in the vehicle.

[0074] like Figure 4 As shown in (a), when the first message 401 to the fifth message 405 are received in multi-session mode, as follows: Figure 4 As shown in (b), the speech direction allocator 130 can arrange conversation partners in five allocated spaces 211 to 215 based on the order in which messages are received.

[0075] The speech direction allocator 130 can arrange the session partner corresponding to the first message 401 in the edge space. In this figure, the session partner corresponding to the first message 401 can be arranged in the leftmost space 211.

[0076] The speech direction allocator 130 can arrange the conversation partner corresponding to the second message 402 in the space furthest from the conversation partner corresponding to the first message 401. In this figure, the conversation partner corresponding to the second message 402 can be arranged in the rightmost space 215.

[0077] The speech direction allocator 130 can arrange the conversation partner corresponding to the third message 403 in the space furthest from the conversation partner corresponding to the second message 402. In this figure, the conversation partner corresponding to the third message 403 can be arranged in space 212 to the left of the central space.

[0078] The utterance direction assigner 130 can arrange the conversation partner corresponding to the fourth message 404 in the space farthest from the conversation partner corresponding to the third message 403. In this figure, the conversation partner corresponding to the fourth message 404 can be arranged in the space 214 on the right of the center space.

[0079] The utterance direction assigner 130 can arrange the conversation partner corresponding to the fifth message 405 in the space farthest from the conversation partner corresponding to the fourth message 404. In this figure, the conversation partner corresponding to the fifth message 405 can be arranged in the center space 213.

[0080] Figure 5 FIG. 2 is a diagram illustrating a voice output method in a method for supporting a multi-conversation mode according to an embodiment of the disclosure.

[0081] Referring to Figure 5 When many messages are received at the same time, the utterance output unit 150 can set a predetermined utterance output time for the allocated space (S510).

[0082] After step S510, the utterance output unit 150 can determine whether the utterance corresponding to the received message can be output within the predetermined utterance output time, when the utterance cannot be output within the predetermined utterance output time, determine whether the number of sentences of the received message is equal to or greater than a predetermined value, and when the number of sentences is equal to or greater than the predetermined value, divide the utterance output interval (S520). Thereafter, the utterance output unit 150 can output the message based on the divided output interval when the utterance can be output within the predetermined utterance output time.

[0083] After step S520, when the utterance cannot be output within the predetermined utterance output time, the utterance output unit 150 can exclude special characters, modifiers, etc. from the message, determine whether the message from which the special characters, modifiers, etc. have been excluded can be output as an utterance within the predetermined utterance output time, and output the message when the message can be output as an utterance within the predetermined utterance output time (S530). Then, the utterance output unit 150 outputs the message from which the special characters, modifiers, etc. have been excluded.

[0084] After step S530, when the message from which the special characters, modifiers, etc. have been excluded cannot be output as an utterance within the predetermined output time, the utterance output unit 150 can control the utterance output speed (S540). Thereafter, the utterance output unit 150 can increase the utterance output speed to a predetermined speed, and determine whether the utterance can be output within the predetermined utterance output time.

[0085] After step S540, when the utterance output according to the utterance output utterance control cannot be performed within a predetermined utterance output time, the utterance output unit 150 can perform control so that the message is simultaneously output in a superimposed manner into the space (S550).

[0086] The above-described methods according to the embodiments can be implemented as programs executable in computers and stored in computer-readable recording media (e.g., non-transitory computer-readable recording media). For example, the methods or operations performed by each of the components including but not limited to the user classifier 120, the utterance direction assigner 130, the utterance modulator 140, and the conversation partner recognizer 160 can be implemented as computer-readable codes stored on a memory implemented by, for example, a computer-readable recording medium such as a non-transitory computer-readable recording medium. Examples of the computer-readable recording medium include ROM, RAM, CD-ROM, magnetic tapes, floppy disks, optical data storage devices, etc. The computer-readable recording medium can be distributed over network-coupled computer systems so that the computer-readable code is stored and executed in a distributed manner. In addition, those skilled in the art can easily infer functional programs, codes, and code segments for implementing the above-described methods. The system can include a computer, a processor, or a microprocessor. Alternatively, each of the user classifier 120, the utterance direction assigner 130, the utterance modulator 140, and the conversation partner recognizer 160, or together, can be implemented as a computer, a processor, or a microprocessor. When the computer, the processor, or the microprocessor reads and executes the computer-readable codes stored in the computer-readable recording medium, the computer, the processor, or the microprocessor can be configured to perform the above-described operations / methods.

[0087] It will be apparent to those skilled in the art that various modifications and variations can be made in the present disclosure without departing from the spirit or scope of the disclosure. Thus, it is intended that the present disclosure cover the modifications and variations of this disclosure provided they come within the scope of the appended claims and their equivalents.

Claims

1. A method for supporting a multi-session mode of a vehicle, comprising: receiving user information of the multi-session mode and at least one of a message and an utterance from a session partner participating in the multi-session mode; allocating a sound space to the session partner based on the user information, respectively; and allocating directionality to an utterance generated based on at least one of the message and the utterance, and outputting the utterance to a sound space allocated to a session partner corresponding to the utterance; determining a face direction of a driver through a camera while the driver is speaking; receiving an utterance of the driver; determining an utterance having sound directionality matching the face direction of the driver; and transmitting the utterance of the driver to a session partner corresponding to the utterance having matching directionality, wherein the directionality allocated to the utterance corresponds to a direction in which the sound space is set with respect to the driver.

2. The method of claim 1, wherein, The user information includes at least one of a number, a gender, an age, an ID, and a name of the session partner.

3. The method of claim 1, wherein, Allocating directionality to an utterance and outputting the utterance to an allocated sound space includes: based on a message reception order, the session partner corresponding to the message is arranged in a farthest sound space among the allocated sound spaces; and allocating directionality to an utterance corresponding to the message received from the session partner and outputting the utterance.

4. The method of claim 3, wherein, Allocating directionality to an utterance corresponding to the message received from the session partner and outputting the utterance includes: setting a predetermined utterance output time for the allocated sound space; determining whether output of the utterance corresponding to the received message is within the predetermined utterance output time; when utterance output is not within the predetermined utterance output time, excluding special characters and modifiers from the message, and determining whether output of the message excluding the special characters and the modifiers as an utterance is within the predetermined utterance output time; when output of the message excluding the special characters and the modifiers as an utterance is not within the predetermined utterance output time, controlling an utterance output speed so that the utterance output speed becomes a predetermined value, and determining whether utterance output is within the predetermined utterance output time; and when utterance output according to the utterance output speed control is not within the predetermined utterance output time, controlling the message so that the message is output to the sound space in a superimposed manner.

5. The method of claim 4, wherein, Determining whether output of the utterance corresponding to the received message is within the predetermined utterance output time includes, when utterance output is not within the predetermined utterance output time, determining whether a number of sentences of the received message is equal to or greater than a predetermined value, and when the number of sentences is equal to or greater than the predetermined value, dividing an utterance output interval. 6.The method of claim 3, further comprising: determining whether the face direction of the driver coincides with the directionality of sound; and when the face direction of the driver coincides with the directionality of sound, identifying a session partner corresponding to a corresponding sound space. 7.The method of claim 6, further comprising: checking whether the session partner is correct; and transmitting a message or utterance of the driver to the session partner. 8.A non-transitory computer-readable recording medium storing a program for implementing the method of claim 1. 9.A system for supporting a multi-session mode of a vehicle, comprising: a data receiver for receiving user information of the multi-session mode and at least one of a message and an utterance from a session partner participating in the multi-session mode; an utterance direction allocator for allocating a sound space to the session partner respectively based on the user information; and an utterance outputter for allocating directionality to an utterance generated based on at least one of the message and the utterance and outputting the utterance to a sound space allocated to a session partner corresponding to the utterance; an utterance receiver for receiving an utterance of a driver when the driver speaks to the session partner; a session partner recognizer for determining a face direction of the driver through a camera when the driver speaks and determining an utterance having sound directionality matching the face direction of the driver; and a data transmitter for transmitting the utterance of the driver to a session partner corresponding to the utterance having matching directionality, wherein the directionality allocated to the utterance corresponds to a direction in which the sound space is set with respect to the driver.

10. The system of claim 9, wherein, The user information includes at least one of a number, a gender, an age, an ID, and a name of the session partner.

11. The system of claim 9, wherein, The utterance outputter arranges a session partner corresponding to a message in a farthest sound space among the allocated sound spaces based on a message reception order, allocates directionality to an utterance corresponding to the message received from the session partner, and outputs the utterance.

12. The system of claim 11, wherein, The utterance outputter sets a predetermined utterance output time for the allocated sound space, The utterance outputter determines whether output of the utterance corresponding to the received message is within the predetermined utterance output time, when utterance output is not within the predetermined utterance output time, the utterance outputter excludes special characters and modifiers from the message and determines whether output of the message excluding the special characters and the modifiers as an utterance is within the predetermined utterance output time, when output of the message excluding the special characters and the modifiers as an utterance is not within the predetermined utterance output time, the utterance outputter controls an utterance output speed so that the utterance output speed becomes a predetermined value, determines whether utterance output is within the predetermined utterance output time, and when utterance output according to the utterance output speed control is not within the predetermined utterance output time, the utterance outputter controls the message so that the message is output to a sound space in a superimposed manner.

13. The system of claim 12, wherein, When utterance output is not within the predetermined utterance output time, the utterance outputter determines whether a number of sentences of the received message is equal to or greater than a predetermined value, and when the number of sentences is equal to or greater than the predetermined value, divides an utterance output interval.

14. The system of claim 11, the conversation partner identifier configured to determine whether the face direction of the driver is consistent with the directionality of sound; and identify a conversation partner corresponding to a corresponding sound space when the face direction of the driver is consistent with the directionality of sound.

15. The system of claim 14, further comprising: the data transmitter further configured to transmit a message or utterance of the driver to the conversation partner after receiving from the driver whether the conversation partner intended by the driver matches the conversation partner.

Citation Information

Patent Citations

  • Information processing device, control method, and program

    CN107408027A

  • Audio processing method, processing equipment, terminal and computer readable storage medium

    CN110035250A

  • Speech dialogue device

    JP2009261010A