System for providing interactive artificial intelligence-based car infotainment, and control method therefor
An AI-driven system provides interactive 3D car graphics and functions via voice and text commands, addressing the challenge of complex car information comprehension by enabling intuitive visualization and understanding.
Patent Information
- Application Number
- PCT/KR2024/019999
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-03
AI Technical Summary
Users face difficulty in understanding and relating fragmented automobile information to the actual appearance and functions of a car, especially as vehicle types and features become more complex, without being restricted by space and time.
An interactive artificial intelligence-based system that uses conversational AI to provide three-dimensional graphic images of a car's exterior and interior, allowing users to explore and learn about its functions through a user terminal and server interaction, including voice and text commands for enhanced understanding.
Enables users to visualize and comprehend car features and functions intuitively without spatial or temporal limitations, enhancing user experience and knowledge through interactive, real-time graphical simulations.
Smart Images

Figure KR2024019999_03072025_PF_FP_ABST
Abstract
Description
System for providing interactive artificial intelligence-based automotive infotainment and its control method
[0001] The present disclosure relates to a conversational artificial intelligence-based automotive infotainment system, and more particularly, to a system that provides images and function information of a vehicle corresponding to a user's voice or text through conversational artificial intelligence, and a control method thereof.
[0002] This technology was developed through the Artificial Intelligence Industry Convergence Business Unit's project (2023 AI+X Technology Development Support Project / Project Name: Technology for Building Content for Conversational AI-Based Future Mobility Car Infotainment).
[0003] An artificial intelligence system is a computer system that implements human-level intelligence. It is a system in which the machine learns and makes judgments on its own, and its recognition rate improves with use.
[0004] Artificial intelligence technology consists of machine learning and deep learning technologies that utilize algorithms that classify and learn the characteristics of input data on their own, as well as elemental technologies that mimic the cognitive and judgment functions of the human brain by utilizing machine learning and deep learning algorithms.
[0005] Element technologies include, for example, linguistic comprehension technology that recognizes human language / characters, visual comprehension technology that recognizes objects as if they were human vision, and inference / prediction technology that judges information and logically infers and predicts.
[0006] Linguistic understanding is the technology of recognizing and applying / processing human language / characters, including natural language processing, machine translation, dialogue systems, question-answering, and speech recognition / synthesis.
[0007] Specifically, it is possible to automatically determine the semantic role of a sentence by having the computer determine what the subject and object are in the sentence and what their semantic relationship is, and to analyze the sentence structure and dependency structure.
[0008] Furthermore, with the advancement of computer graphics technology, the appearance of real-world objects can now be rendered as computer graphics-based 3D images and presented to users. Rendering of 3D computer graphics images and videos can be performed to create realistic appearances, resembling real objects.
[0009] As the types of automobiles increase and their functions become more complex, there is a problem in that users are unable to understand and connect this information to the actual appearance and functions of the automobile even if they are provided with fragmentary and one-off automobile information.
[0010] Therefore, there is a need to find a way for users to view the actual appearance of a car in a simple way without being restricted by space and time, and to learn the complex functions of a car through a virtual graphic simulation.
[0011] The purposes of the present disclosure are not limited to those mentioned above, and other purposes and advantages of the present disclosure not mentioned above can be understood through the following description and will be more clearly understood through the embodiments of the present disclosure. Furthermore, it will be readily apparent that the purposes and advantages of the present disclosure can be realized by the means and combinations thereof set forth in the claims.
[0012] According to one embodiment of the present disclosure, a method for controlling a conversational artificial intelligence-based car infotainment system including a user terminal and a server may include: when user command data is acquired while the user terminal outputs a car GUI, the user terminal transmits the user command data to a server; the server identifies text data corresponding to the user command data received from the user terminal; the server identifies the meaning of the text data; the server generates an output control signal of a car GUI corresponding to the text data based on the identified meaning; the server generates answer data corresponding to the text data based on the identified meaning; the server transmits the output control signal and the answer data to the user terminal; and the user terminal outputs and provides the car GUI and the answer data corresponding to the output control signal to a user.
[0013] Meanwhile, the automobile GUI may be a three-dimensional graphic image of the exterior or interior of a preset automobile model.
[0014] Meanwhile, the step of identifying the meaning of the text data includes a step in which the server identifies a part or a group of parts of an automobile included in the text data based on the identified meaning; and the step of generating an output control signal of the GUI includes a step in which the server generates an output control signal of the GUI for the user terminal to output an enlarged image of the part or the group of parts of the identified automobile.
[0015] Meanwhile, the step of identifying the meaning of the text data includes a step in which the server identifies a function of the vehicle corresponding to the identified vehicle part or the identified vehicle part group; and the step of generating an output control signal of the GUI includes a step in which the server generates an output control signal of the GUI for outputting an image in which the function of the vehicle corresponding to the identified vehicle part or the identified vehicle part group is performed.
[0016] Meanwhile, the step of identifying the meaning of the text data includes a step of the server identifying change information for at least one of a first-person view angle of the user and a first-person view distance of the user corresponding to the automobile GUI being output through the user terminal based on the identified meaning; and the step of generating an output control signal of the GUI includes a step of the server generating an output control signal of the GUI that changes at least one of a first-person view angle of the user and a first-person view distance of the user corresponding to the automobile GUI being output through the user terminal based on the identified change information.
[0017] Meanwhile, the step of identifying the meaning of the text data includes a step in which the server identifies a part or a group of parts of an automobile included in the text data based on the identified meaning; and the step of generating the answer data includes a step in which the server can generate answer data including a description of the identified part or the group of parts of the automobile.
[0018] Meanwhile, the step of identifying the meaning of the text data includes a step in which the server identifies a function of the vehicle corresponding to a part of the identified vehicle or a group of parts identified; and the step of generating the answer data includes a step in which the server can generate answer data including a description of the function of the identified vehicle.
[0019] Meanwhile, the step of outputting the automobile GUI and the answer data and providing them to the user may include the user terminal outputting the automobile GUI through a display and providing it to the user, and the user terminal outputting an answer voice corresponding to the answer data through a speaker and providing it to the user.
[0020] According to one embodiment of the present disclosure, a method for controlling an electronic device providing interactive artificial intelligence-based car infotainment may include, when user command data is acquired while outputting a car GUI through a display, a step of identifying text data corresponding to the user command data; a step of identifying the meaning of the text data; a step of generating an output control signal of a car GUI corresponding to the text data based on the identified meaning; a step of generating answer data corresponding to the text data based on the identified meaning; and a step of outputting the car GUI and the answer data corresponding to the output control signal through at least one of a display and a speaker to provide the same to a user.
[0021] A non-transitory computer-readable recording medium according to one embodiment of the present disclosure may store at least one instruction that is executed by a processor of an electronic device to perform a control method of the electronic device.
[0022] A system that combines conversational artificial intelligence and car infotainment presented in 3D images allows users to view the exterior of a car and learn about its functions without being restricted by time or space.
[0023] Aspects, features and advantages of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings.
[0024] FIG. 1 is a diagram illustrating an electronic device and a server constituting an interactive artificial intelligence-based automotive infotainment system according to one embodiment of the present disclosure.
[0025] FIG. 2 is a diagram for explaining the overall operation of an interactive artificial intelligence-based automotive infotainment system according to one embodiment of the present disclosure.
[0026] FIG. 3a is a block diagram illustrating a configuration of a user terminal according to an embodiment of the present disclosure.
[0027] FIG. 3b is a block diagram illustrating a configuration of a server according to an embodiment of the present disclosure.
[0028] FIG. 4 is a sequence diagram illustrating the operation of a system according to an embodiment of the present disclosure.
[0029] FIG. 5A is a drawing for explaining a screen in which a three-dimensional image of the exterior of a car is output through a user terminal according to one embodiment of the present disclosure.
[0030] FIG. 5b is a drawing for explaining a screen in which a three-dimensional image of the exterior of a car is output through a user terminal according to an embodiment of the present disclosure.
[0031] FIG. 6A is a drawing for explaining a screen in which a three-dimensional image of the interior of a car is output through a user terminal according to one embodiment of the present disclosure.
[0032] FIG. 6b is a drawing for explaining a screen in which a three-dimensional image of the interior of a car is output through a user terminal according to one embodiment of the present disclosure.
[0033] FIG. 7 is a flowchart illustrating the operation of an electronic device according to an embodiment of the present disclosure.
[0034] The present embodiments may be modified and have various embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope to specific embodiments, but should be understood to encompass various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0035] In describing the present disclosure, if it is determined that a specific description of a related known function or configuration may unnecessarily obscure the gist of the present disclosure, a detailed description thereof will be omitted.
[0036] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concepts of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to further faithfully and completely convey the technical concepts of the present disclosure to those skilled in the art.
[0037] The terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the scope of the rights. Singular expressions include plural expressions unless the context clearly dictates otherwise.
[0038] In this disclosure, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a corresponding feature (e.g., a component such as a number, function, operation, or part), and do not exclude the presence of additional features.
[0039] In this disclosure, expressions such as “A or B,” “at least one of A and / or B,” or “one or more of A or / and B” can include all possible combinations of the listed items. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” can all refer to (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
[0040] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0041] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that the component may be directly coupled to the other component, or may be connected through another component (e.g., a third component).
[0042] On the other hand, when it is said that a component (e.g., a first component) is "directly connected" or "directly connected" to another component (e.g., a second component), it can be understood that no other component (e.g., a third component) exists between the component and the other component.
[0043] The expression "configured to" as used in the present disclosure may be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" may not necessarily mean only "specifically designed to" in terms of hardware.
[0044] Instead, in some contexts, the phrase "a device configured to" may mean that the device, in conjunction with other devices or components, is "capable of" performing A, B, and C. For example, the phrase "a processor configured (or set) to perform A, B, and C" may refer to a dedicated processor (e.g., an embedded processor) for performing those operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform those operations by executing one or more software programs stored in a memory device.
[0045] In the embodiments, a 'module' or 'part' performs at least one function or operation, and may be implemented as hardware or software, or as a combination of hardware and software. Furthermore, a plurality of 'modules' or 'parts' may be integrated into at least one module and implemented as at least one processor, except for a 'module' or 'part' that needs to be implemented as a specific hardware.
[0046] Meanwhile, the various elements and areas in the drawings are schematically drawn. Therefore, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0047] Hereinafter, with reference to the attached drawings, embodiments according to the present disclosure will be described in detail so that a person having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement the present disclosure.
[0048] FIG. 1 is a diagram illustrating an electronic device and a server constituting an interactive artificial intelligence-based automotive infotainment system according to one embodiment of the present disclosure.
[0049] Referring to FIG. 1, the user terminal (100) may include at least one of a desktop personal computer, a tablet personal computer, a laptop personal computer, a netbook computer, a mobile device, a smartphone, and a wearable device, but is not limited thereto.
[0050] The server (200) may be, for example, a computer that provides services to clients via a network. The server (200) may be an FTP server (200), a web server (200), a database server (200), or a cloud-based server (200), and the server (200) may be built with an operating system such as Linux.
[0051] FIG. 2 is a diagram for explaining the overall operation of an interactive artificial intelligence-based automotive infotainment system according to one embodiment of the present disclosure.
[0052] Referring to FIG. 2, the user terminal (100) can output an automobile GUI to provide an interactive artificial intelligence-based automobile infotainment service to the user.
[0053] The user terminal (100) can obtain the user's voice data through a microphone while the automobile GUI is output.
[0054] In addition, the user terminal (100) can input text corresponding to a user command through the user interface while the automobile GUI is output.
[0055] Here, the user's voice data or text data may be for a user command for controlling the automotive infotainment system. The user terminal (100) may transmit the acquired voice data or text data to the server (200).
[0056] The server (200) can perform a voice recognition operation to identify text data corresponding to voice data received from the user terminal (100).
[0057] The server (200) can identify the meaning of text data corresponding to the user's voice, generate a control signal for controlling the vehicle GUI output of the vehicle infotainment system, and generate response data for the text data and transmit the same to the user terminal (100). In this case, the user terminal (100) can output the vehicle GUI based on the control signal received from the server (200) and provide it to the user, and output a response voice corresponding to the response data and provide it to the user.
[0058] The interactive artificial intelligence-based automobile infotainment system including the user terminal (100) and server (200) described above can be utilized in a web-environment and web-based manner without downloading separate programs, solutions, and data to the user terminal (100), and thus has the advantage of being able to provide services to users easily and quickly.
[0059] FIG. 3a is a block diagram illustrating the configuration of a user terminal (100) according to one embodiment of the present disclosure.
[0060] Referring to FIG. 3a, the user terminal (100) may include a display (110), a speaker (120), a microphone (130), a communication interface (140), a memory (150), and a processor (160).
[0061] However, the device configuration of the user terminal (100) is not limited to that described above, and other device configurations, for example, a user interface may be additionally included or some configurations may be omitted.
[0062] The display (110) may include various types of display (110) panels, such as a Liquid Crystal Display (LCD) panel, an Organic Light Emitting Diodes (OLED) panel, an Active-Matrix Organic Light-Emitting Diode (AM-OLED), a Liquid Crystal on Silicon (LCoS), a Quantum dot Light-Emitting Diode (QLED), a Digital Light Processing (DLP), a Plasma Display Panel (PDP) panel, an inorganic LED panel, and a Micro LED panel, but is not limited thereto. Meanwhile, the display (110) may also configure a touch screen together with a touch panel, and may be formed of a flexible panel.
[0063] The display (110) may be implemented in a 2D square or rectangular shape, but is not limited thereto and may be implemented in various shapes such as a circle, polygon, or 3D solid shape.
[0064] The display (110) may be placed on one area of the surface of the electronic device, but is not limited thereto, and may be implemented as a three-dimensional hologram projected on a three-dimensional space projected on space, or as a projection projected on a two-dimensional plane.
[0065] The display (110) may be included as a component of the electronic device, but is not limited thereto, and a separately provided display (110) device may be connected wirelessly / wirelessly to the electronic device through a communication interface (140) or an input / output interface to output images, videos, and GUI according to signals from the processor (160). In this case, the processor (160) may perform a connection with the display (110) device wirelessly / wirelessly through the communication interface (140) or the input / output interface to transmit signals for outputting images, videos, and GUI.
[0066] The processor (160) can output automobile infotainment consisting of three-dimensional images, videos, and GUIs of the exterior or interior of the automobile to the user through the display (110).
[0067] The processor (160) can output text containing information on the configuration-specific functions of the vehicle and provide it to the user through the display (110).
[0068] In addition, the processor (160) can output a response text to the user's voice through the display (110) and provide it to the user.
[0069] The speaker (120) may be composed of a tweeter for reproducing high-frequency sounds, a midrange for reproducing mid-frequency sounds, a woofer for reproducing low-frequency sounds, a subwoofer for reproducing ultra-low-frequency sounds, an enclosure for controlling resonance, a crossover network for dividing the frequency of an electric signal input to the speaker (120) into bands, etc.
[0070] The speaker (120) can output audio signals to the outside of the electronic device. The speaker (120) can output multimedia playback, recording playback, various notification sounds, voice messages, etc. The electronic device may include an audio output device such as the speaker (120), but may also include an output device such as an audio output terminal. In particular, the speaker (120) can provide acquired information, information processed and produced based on acquired information, response results to user voice, operation results, etc. in voice form.
[0071] The processor (160) can output a response voice corresponding to the user's voice through the speaker (120) and provide it to the user.
[0072] In addition, the processor (160) can output various guidance voices related to the operation of interactive artificial intelligence-based automotive infotainment to the user through the speaker (120).
[0073] A microphone (130) may refer to a module that acquires sound and converts it into an electrical signal, and may be a condenser microphone, a ribbon microphone, a moving coil microphone, a piezoelectric element microphone, a carbon microphone, or a MEMS (Micro Electro Mechanical System) microphone. In addition, it may be implemented in an omnidirectional, bidirectional, unidirectional, subcardioid, supercardioid, or hypercardioid manner.
[0074] The processor (160) can identify a user voice command based on the acquired user voice information. Specifically, the processor (160) can include a STT (Speech to Text) module, and can identify an electrical signal corresponding to the user voice in the form of text through the STT module. The processor (160) can perform an operation corresponding to the identified user voice command. Here, when identifying the user voice command, the processor (160) can identify the user voice command included in the electrical signal corresponding to the user voice or the text corresponding to the electrical signal by using a language model (Lagauge Mode, LM) and an ASR (Automatic Sound Recognition) model.
[0075] The processor (160) can acquire user voice data through the microphone (130). Here, the user voice data may be a user command for controlling an interactive artificial intelligence-based automotive infotainment system.
[0076] The user interface may include a button, a lever, a switch, a touch interface, etc., and the touch interface may be implemented in a way that receives input by the user's touch on the display (110) screen.
[0077] The processor (160) can receive text data through a user interface. Here, the text data can correspond to a user command for GUI control and functional information acquisition of an interactive artificial intelligence-based automotive infotainment system.
[0078] Additionally, the processor (160) can receive GUI control commands for enlarging, reducing, moving, switching, etc. of the automobile GUI being output through the display via a touch-type user interface.
[0079] The communication interface (140) may include a wireless communication interface, a wired communication interface, or an input interface. The wireless communication interface may communicate with various external devices using wireless communication technology or mobile communication technology. Examples of such wireless communication technologies may include Bluetooth, Bluetooth Low Energy, CAN communication, Wi-Fi, Wi-Fi Direct, ultrawide band (UWB), Zigbee, infrared Data Association (IrDA), or near field communication (NFC). Examples of mobile communication technologies may include 3GPP, Wi-Max, Long Term Evolution (LTE), 5G, and the like.
[0080] A wireless communication interface can be implemented using an antenna, a communication chip, a substrate, etc. that can transmit electromagnetic waves to the outside or receive electromagnetic waves transmitted from the outside.
[0081] A wired communication interface can communicate with various devices based on a wired communication network. Here, the wired communication network can be implemented using physical cables such as paired cables, coaxial cables, fiber optic cables, or Ethernet cables.
[0082] Depending on the embodiment, either the wireless communication interface or the wired communication interface may be omitted. Accordingly, the electronic device may include only the wireless communication interface or only the wired communication interface. Furthermore, the electronic device may include an integrated communication interface (140) that supports both wireless connections via the wireless communication interface and wired connections via the wired communication interface.
[0083] The electronic device is not limited to including one communication interface (140) that performs one type of communication connection, but may include multiple communication interfaces (140) that perform multiple types of communication connections.
[0084] The processor (160) can transmit the user's voice data acquired through the microphone (130) to the server (200) through the communication interface (140).
[0085] In addition, the processor (160) can transmit text data corresponding to the user's voice data acquired through the microphone (130) to the server (200) through the communication interface (140).
[0086] The processor (160) can receive a control signal for outputting a vehicle GUI corresponding to a user's voice command from the server (200) through a communication interface (140).
[0087] The processor (160) can receive response data corresponding to the user's voice command from the server (200) through the communication interface (140).
[0088] However, in addition, the processor (160) can perform a communication connection with the server (200) through the communication interface (140) to transmit or receive various information related to the operation of the conversational artificial intelligence-based automobile infotainment system.
[0089] The memory (150) temporarily or non-temporarily stores various programs or data, and transmits the stored information to the processor (160) upon a call from the processor (160). In addition, the memory (150) can store various information necessary for operations, processing, or control operations of the processor (160) in an electronic format.
[0090] The memory (150) may include, for example, at least one of a main memory and an auxiliary memory. The main memory may be implemented using a semiconductor storage medium such as ROM and / or RAM. The ROM may include, for example, a conventional ROM, EPROM, EEPROM, and / or MASK-ROM. The RAM may include, for example, DRAM and / or SRAM. The auxiliary memory may be implemented using at least one storage medium capable of permanently or semi-permanently storing data, such as a flash memory device, an SD (Secure Digital) card, a solid state drive (SSD), a hard disk drive (HDD), an optical media such as a magnetic drum, a compact disc (CD), a DVD, or a laser disc, a magnetic tape, a magneto-optical disc, and / or a floppy disk.
[0091] The memory (150) can store the user's voice data acquired through the microphone (130). The memory (150) can store text data corresponding to the user's voice data.
[0092] The memory (150) can store a control signal for outputting a vehicle GUI corresponding to a user's voice command. The memory (150) can store response data corresponding to the user's voice command.
[0093] The memory (150) can store various information related to the operation of an interactive artificial intelligence-based automotive infotainment system.
[0094] The processor (160) controls the overall operation of the electronic device. Specifically, the processor (160) is connected to the configuration of the electronic device, including the memory (150) as described above, and can control the overall operation of the electronic device by executing at least one instruction stored in the memory (150) as described above. In particular, the processor (160) may be implemented as a single processor or as multiple processors.
[0095] The processor (160) may be implemented in various ways. For example, one or more processors (160) may include one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), a Many Integrated Core (MIC), a Digital Signal Processor (DSP), a Neural Processing Unit (NPU), a hardware accelerator, or a machine learning accelerator. The one or more processors (160) may control one or any combination of other components of the electronic device, and may perform operations related to communication or data processing. The one or more processors (160) may execute one or more programs or instructions stored in the memory (150). For example, the one or more processors (160) may perform a method according to an embodiment of the present disclosure by executing one or more instructions stored in the memory (150).
[0096] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one processor (160) or may be performed by a plurality of processors (160). For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by the first processor, or the first operation and the second operation may be performed by the first processor (e.g., a general-purpose processor) and the third operation may be performed by the second processor (e.g., an artificial intelligence-only processor).
[0097] One or more processors (160) may be implemented as a single core processor including one core, or may be implemented as one or more multicore processors including multiple cores (e.g., homogeneous multicore or heterogeneous multicore). When one or more processors (160) are implemented as a multicore processor, each of the multiple cores included in the multicore processor may include an internal memory of the processor, such as an on-chip memory (150), and a common cache shared by the multiple cores may be included in the multicore processor (160). In addition, each of the multiple cores (or some of the multiple cores) included in the multicore processor (160) may independently read and execute a program instruction for implementing a method according to an embodiment of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to an embodiment of the present disclosure.
[0098] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core among the plurality of cores included in a multi-core processor, or may be performed by the plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in the multi-core processor, or the first operation and the second operation may be performed by a first core included in the multi-core processor, and the third operation may be performed by a second core included in the multi-core processor.
[0099] In embodiments of the present disclosure, the processor (160) may mean a system on a chip (SoC) in which one or more processors (160) and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, a GPU, an APU, a MIC, a DSP, an NPU, a hardware accelerator, or a machine learning accelerator, but embodiments of the present disclosure are not limited thereto.
[0100] The processor (160) is connected to the display (110), speaker (120), microphone (130), communication interface (140), and memory (150) described above, and can identify / acquire data, information, and signals related to the operation of the conversational artificial intelligence-based automobile infotainment system, and can perform an operation to control the above-described configuration by executing at least one instruction.
[0101] FIG. 3b is a block diagram illustrating the configuration of a server (200) according to one embodiment of the present disclosure.
[0102] Referring to FIG. 3b, the server (200) may include a communication interface (210), memory (220), and processor (230).
[0103] The general detailed configuration and functions of the communication interface (210), memory (220), and processor (230) are described above with reference to FIG. 3a.
[0104] The processor (230) can perform a communication connection with the user terminal (100) through the communication interface (210) to receive the user's voice data or text data corresponding to the user's voice data.
[0105] When user voice data is received through the communication interface (210), the processor (230) can identify a user voice command based on the acquired user voice information. Specifically, the processor (230) can include a STT (Speech to Text) module, and can identify an electrical signal corresponding to the user voice in the form of text through the STT module. The processor (230) can perform an operation corresponding to the identified user voice command. Here, when identifying the user voice command, the processor (230) can identify the user voice command included in the electrical signal corresponding to the user voice or the text corresponding to the electrical signal by using a language model (Lagauge Mode, LM) and an ASR (Automatic Sound Recognition) model.
[0106] The processor (230) can transmit control signals for automobile GUI output and response data for user voice to the user terminal (100) through the communication interface (210).
[0107] In addition, the processor (230) can perform a communication connection with a user terminal (100) through a communication interface (210) to transmit or receive various information related to the operation of an interactive artificial intelligence-based automobile infotainment system.
[0108] The memory (220) can store the user's voice data acquired through a microphone. The memory (220) can store text data corresponding to the user's voice data.
[0109] The memory (220) can store a control signal for outputting a vehicle GUI corresponding to a user's voice command. The memory (220) can store response data corresponding to the user's voice command.
[0110] The memory (220) can store various information related to the operation of an interactive artificial intelligence-based automotive infotainment system.
[0111] The processor (230) is connected to a communication interface (210) and a memory (220) to identify / acquire data, information, and signals related to the operation of an interactive artificial intelligence-based automotive infotainment system, and can perform an operation to control the above-described configuration by executing at least one instruction.
[0112] A more specific method of controlling the system is described with reference to FIGS. 4 to 6.
[0113] FIG. 4 is a sequence diagram illustrating the operation of a system according to an embodiment of the present disclosure.
[0114] Referring to FIG. 4, the user terminal (100) can output an automobile GUI and provide it to the user.
[0115] FIGS. 5a and 5b are drawings for explaining a screen in which a three-dimensional image of the exterior of a car is output through a user terminal (100) according to one embodiment of the present disclosure.
[0116] Referring to FIGS. 5a and 5b, the vehicle GUI may be a three-dimensional graphic image of the appearance of a preset vehicle model.
[0117] The 3D image of the car's exterior is a realistic representation of a preset car model, and the color and texture of the car's exterior surface can be expressed differently depending on the light.
[0118] FIGS. 6A and 6B are drawings for explaining a screen on which a three-dimensional image of the interior of a car is output through a user terminal (100) according to one embodiment of the present disclosure.
[0119] Referring to FIGS. 6a and 6b, the automobile GUI may be a three-dimensional graphic image of the interior of a preset automobile model.
[0120] The three-dimensional graphic image of the interior may be a dashboard, air vents, levers, buttons, touch pad, instrument panel, center monitor, center fascia, gear, rear-view mirror, side mirror, etc. that are visible to the user when looking forward from the driver's seat of the car.
[0121] Additionally, the three-dimensional graphic images of the interior may include buttons located under the driver's seat, such as the brake pedal, accelerator pedal, front passenger seat, rear passenger seat, and buttons located inside the car doors.
[0122] The user terminal (100) can acquire user command data while outputting the automobile GUI (S410). Here, the user command data may be user voice data acquired through a microphone (130), text data acquired through a user interface, or GUI touch control command data acquired through a touch interface.
[0123] When user command data is obtained, the user terminal (100) can transmit the user command data to the server (200) (S420).
[0124] However, without being limited thereto, when the user's voice data is acquired, the user terminal (100) may input the user's voice data into a pre-trained voice recognition model to acquire text data corresponding to the voice data, and transmit the acquired text data to the server (200).
[0125] Here, the speech recognition model may be an automatic sound recognition (ASR) model.
[0126] A speech recognition model may include a preprocessing module that performs preprocessing operations such as signal amplification and feature value extraction on speech data.
[0127] A speech recognition model may include a pattern identification module that recognizes phonemes, syllables, and words, which are elements necessary for constructing sentences, based on features obtained through preprocessing of a speech signal to produce results from features.
[0128] The speech recognition model may include a post-processing module that performs post-processing operations to obtain sentences by reconstructing phonemes, syllables, and words for vector values output from a language module or a speech recognition module.
[0129] The server (200) can identify text data corresponding to user command data, for example, voice data or text data (S430). Here, the server (200) can also identify text data by inputting voice data into the above-described voice recognition model.
[0130] The server (200) can identify the meaning of text data corresponding to user command data received from the user terminal (100) (S440). Based on the identified meaning, the server (200) can generate an output control signal of the automobile GUI corresponding to the text data (S450), and based on the identified meaning, can generate response data corresponding to the text data (S460).
[0131] Here, the server (200) can input text data into a language model to identify the meaning of the text data and generate an output control signal and response data corresponding to the text data.
[0132] A language model is a model that assigns probabilities to word sequences, and is also called a probabilistic language model.
[0133] A language model can be, for example, a Large Language Model (LLM), which is a deep learning model trained on unlabeled text through self-supervised or semi-self-supervised learning based on a large amount of data.
[0134] Natural language processing (NLP), which recognizes and interprets input natural language based on a large language model, and natural language generation (NLG), which generates, reinforces, and reconstructs words and sentences, can be performed.
[0135] More specifically, large-scale language models can be used in machine translation, text summarization, question-answering systems, conversational systems, and content generation.
[0136] A large-scale language model can convert input text data into vector values in a latent space through text embedding, and based on the vector values in the latent space, it can identify the context of words and phrases with similar meanings as well as other relationships between words such as parts of speech, and output results for natural language processing, natural language generation, etc. based on the relationships between these vector values.
[0137] Specifically, the server (200) can identify a part or group of parts of an automobile included in the text data based on the identified meaning.
[0138] A part of an automobile may be any component, accessory, or component included in the automobile. Examples include, but are not limited to, a dashboard, steering wheel, gear shift, center fascia, wipers, side mirrors, brake pedal, air vents, air conditioning control buttons, window buttons, trunk lid, headlamps, and tail lamps.
[0139] A component group may be a group of components that perform a single function or the same function (e.g., driving, speed control, air conditioning, window opening, etc.) related to the components included in the text data. For example, if the component included in the identified text is a wheel, the component group for the wheel may consist of a wheel, a tire, a brake pad, and a brake caliper. Furthermore, if the component included in the identified text is a steering wheel, the component group for the steering wheel may consist of a steering wheel, a turn signal lever, a wiper lever, and a steering wheel position adjustment lever. If the component included in the identified text is a window, the component group for the window may consist of window operation buttons, windows, etc. If the component included in the identified text is an air conditioning system, the component group for the air conditioning system may consist of air conditioning / heater operation buttons, vents, etc. The component group may be a speed control group. If the component included in the identified text is a pedal, the component group for the pedal may consist of a brake pedal, an accelerator pedal, and a clutch pedal.
[0140] And, the server (200) can identify a function of the vehicle corresponding to an identified vehicle part or an identified vehicle part group.
[0141] The functions of a car may include, but are not limited to, for example, controlling the interior temperature using an air conditioning system, accelerating by pressing an accelerator pedal, opening and closing a sunroof, opening and closing windows, wiper function, music playback function, interior lighting operation function, steering function, and gear operation function.
[0142] An output control signal of a GUI for causing a user terminal (100) to output an enlarged image of an identified automobile part or a group of identified parts can be generated.
[0143] The server (200) can generate an output control signal of a GUI for outputting an image in which a function of a vehicle corresponding to an identified vehicle part or a group of identified parts is performed.
[0144] According to various embodiments, when the text data includes a plurality of automobile parts or parts groups, the server (200) may generate an output control signal of a GUI that causes the user terminal (100) to sequentially output a three-dimensional image for each of the plurality of automobile parts or parts groups for a preset period of time based on the meaning of the identified text.
[0145] Here, the server (200) may generate an output control signal of the GUI so that the user terminal (100) outputs a three-dimensional image of a component for a longer period of time when the number of components of a component group corresponding to an identified component among the plurality of components included in the identified text data is large, or when the information provided to the driver in relation to the performance of a function corresponding to the identified component is large, and the user terminal (100) outputs a three-dimensional image of a component for a shorter period of time when the number of components of a component group corresponding to an identified component among the plurality of components is small, or when the information provided to the driver in relation to the performance of a function corresponding to the identified component is small.
[0146] Here, information provided to the driver in relation to the performance of a function can refer to information provided visually, audibly, quantitatively, or numerically when the component performs its function. For example, if the component is a dashboard, the information provided to the user may include speed, engine RPM, fuel gauge, etc. Furthermore, if the component is an air conditioning unit, the information provided to the user may include temperature, wind direction, and wind speed.
[0147] According to various embodiments, the server (200) may generate a GUI output control signal in which a second time period during which the user terminal (100) outputs a three-dimensional image for a component in which the number of components included in a component group corresponding to a first component included in the identified text data is less than a first preset value and information provided to a driver in relation to the performance of a function corresponding to the component is greater than a second preset value is shorter than a first time period during which the user terminal (100) outputs a three-dimensional image for a component in which the number of components included in a component group corresponding to a component included in the identified text data is greater than a first preset value and information provided to a driver in relation to the performance of a function corresponding to the component is less than a second preset value.
[0148] The server (200) can identify change information for at least one of a user's first-person view angle and a user's first-person view distance corresponding to the automobile GUI being output through the user terminal (100) based on the meaning of the identified text data.
[0149] The server (200) can generate an output control signal of the GUI that changes at least one of the user's first-person view angle and the user's first-person view distance corresponding to the automobile GUI being output through the user terminal (100) based on the identified change information.
[0150] The server (200) can identify a part or a group of parts of an automobile included in the text data based on the identified meaning, and generate response data including a description of the identified part or group of parts of an automobile.
[0151] The server (200) can identify a function of the vehicle corresponding to a part of the identified vehicle or a group of identified parts, and generate response data including a description of the function of the identified vehicle.
[0152] In addition, if the user command data received from the user terminal (100) is a touch command for GUI control for zooming in, zooming out, moving, or switching of the automobile GUI, the server (200) can generate a GUI output control signal for zooming in, zooming out, moving, or switching of the automobile GUI corresponding to the touch command.
[0153] The server (200) can transmit control signals and response data to the user terminal (100) (S470).
[0154] The user terminal (100) can output the automobile GUI and response data and provide them to the user (S480).
[0155] The user terminal (100) can output the automobile GUI through a display and provide it to the user.
[0156] The user terminal (100) may provide a response voice corresponding to the response data to the user by outputting it through a speaker. However, the user terminal (100) is not limited thereto, and may also provide a text corresponding to the response data to the user by outputting it through a display.
[0157] The above-described system can be implemented not only as a system composed of a user terminal (100) and a server (200), but also through the operation of a single electronic device.
[0158] FIG. 7 is a flowchart illustrating the operation of an electronic device according to an embodiment of the present disclosure.
[0159] Referring to FIG. 7, when user command data is acquired while outputting a vehicle GUI through a display, the electronic device can identify text data corresponding to the user command data (S710).
[0160] The electronic device can identify the meaning of text data (S720).
[0161] The electronic device can generate an output control signal of the automobile GUI corresponding to the text data based on the identified meaning (S730).
[0162] The electronic device can generate response data corresponding to the text data based on the identified meaning (S740).
[0163] The electronic device can provide the user with a vehicle GUI and response data corresponding to the output control signal by outputting the vehicle GUI and response data through at least one of a display and a speaker (S750).
[0164] The artificial intelligence-related function according to the present disclosure is operated through the processor (160) (230) and memory (150) (220) of the user terminal (100) or server (200).
[0165] The processor (160)(230) may be composed of one or more processors (160)(230). At this time, the one or more processors (160)(230) may include at least one of a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), and an NPU (Neural Processing Unit), but is not limited to the examples of the processors (160)(230) described above.
[0166] The CPU is a general-purpose processor (160) (230) capable of performing not only general operations but also artificial intelligence operations. It can efficiently execute complex programs through a multi-layer cache structure. The CPU is advantageous in a serial processing method that enables organic linking of previous and subsequent calculation results through sequential calculations. The general-purpose processor (160) (230) is not limited to the examples described above, except in cases where it is specifically referred to as a CPU.
[0167] A GPU is a processor (160) (230) for large-scale operations such as floating point operations used in graphic processing, and can perform large-scale operations in parallel by integrating a large number of cores. In particular, a GPU may be advantageous compared to a CPU in parallel processing methods such as convolution operations. In addition, a GPU may be used as a co-processor (160) (230) to supplement the function of a CPU. The processor (160) (230) for large-scale operations is not limited to the examples described above, except in cases where it is specified as the above-described GPU.
[0168] An NPU is a processor (160) (230) specialized in artificial intelligence operations using an artificial neural network, and each layer constituting the artificial neural network can be implemented with hardware (e.g., silicon). At this time, since the NPU is designed specifically according to the required specifications of the company, it has a lower degree of freedom compared to a CPU or GPU, but it can efficiently process the artificial intelligence operations requested by the company. Meanwhile, as a processor (160) (230) specialized in artificial intelligence operations, the NPU can be implemented in various forms such as a TPU (Tensor Processing Unit), an IPU (Intelligence Processing Unit), a VPU (Vision processing unit), etc. The artificial intelligence processor (160) (230) is not limited to the above-described examples, except in cases where it is specified as the above-described NPU.
[0169] In addition, one or more processors (160)(230) may be implemented as a SoC (System on Chip). At this time, the SoC may further include, in addition to one or more processors (160)(230), a memory (150)(220), and a network interface such as a bus for data communication between the processor (160)(230) and the memory (150)(220).
[0170] When a plurality of processors (160)(230) are included in a SoC (System on Chip) included in a user terminal (100) or a server (200), the user terminal (100) or the server (200) may perform operations related to artificial intelligence (e.g., operations related to learning or inference of an artificial intelligence model) by using some of the processors (160)(230) among the plurality of processors (160)(230). For example, the user terminal (100) or the server (200) may perform operations related to artificial intelligence by using at least one of a GPU, an NPU, a VPU, a TPU, or a hardware accelerator specialized in artificial intelligence operations such as convolution operations or matrix multiplication operations among the plurality of processors (160)(230). However, this is merely an example, and it is of course possible to process operations related to artificial intelligence by using a CPU or a general-purpose processor (160)(230).
[0171] In addition, the user terminal (100) or server (200) can perform operations related to functions related to artificial intelligence by utilizing multiple cores (e.g., dual cores, quad cores, etc.) included in one processor (160) (230). In particular, the user terminal (100) or server (200) can perform artificial intelligence operations such as convolution operations, matrix multiplication operations, etc. in parallel by utilizing multiple cores included in the processor (160) (230).
[0172] One or more processors (160)(230) are controlled to process input data according to predefined operation rules or artificial intelligence models stored in the memory (150)(220). The predefined operation rules or artificial intelligence models are characterized by being created through learning.
[0173] Here, "created through learning" means that a predefined set of behavioral rules or an artificial intelligence model with desired characteristics is created by applying a learning algorithm to a large number of learning data. This learning may be performed on the device itself, where the artificial intelligence according to the present disclosure is performed, or through a separate server (200) / system.
[0174] An artificial intelligence model may be composed of multiple neural network layers. At least one layer has at least one weight value and performs its operation through the operation result of the previous layer and at least one defined operation. Examples of neural networks include a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, and a transformer. The neural networks in the present disclosure are not limited to the above-described examples unless otherwise specified.
[0175] A learning algorithm is a method for training a target device (e.g., a robot) using a large amount of learning data, enabling the target device to make decisions or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Unless otherwise specified, the learning algorithms in this disclosure are not limited to the aforementioned examples.
[0176] According to one embodiment, the method according to the various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0177] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.
Claims
1. A method for controlling an interactive artificial intelligence-based car infotainment system including a user terminal and a server, A step of obtaining user command data while the user terminal outputs the automobile GUI, wherein the user terminal transmits the user command data to a server; A step in which the server identifies text data corresponding to the user command data received from the user terminal; A step in which the server identifies the meaning of the text data; A step in which the server generates an output control signal of an automobile GUI corresponding to the text data based on the identified meaning; A step in which the server generates response data corresponding to the text data based on the identified meaning; The step of the server transmitting the output control signal and the response data to the user terminal; and A control method, comprising: a step of the user terminal outputting a car GUI and the answer data corresponding to the output control signal and providing them to the user.
2. In paragraph 1, The above car GUI is, A control method, wherein the three-dimensional graphic image is the exterior or interior of a preset automobile model.
3. In paragraph 1, The step of identifying the meaning of the above text data is: The step of the server identifying a part or a group of parts of the automobile included in the text data based on the identified meaning; The step of generating the output control signal of the above GUI is: A control method, wherein the server generates an output control signal of a GUI for causing the user terminal to output an enlarged image of a part of the identified automobile or a group of parts identified.
4. In paragraph 3, The step of identifying the meaning of the above text data is: The step of the server identifying a function of the vehicle corresponding to a part of the identified vehicle or a group of parts identified; The step of generating the output control signal of the above GUI is: A control method, wherein the server generates an output control signal of a GUI for outputting an image in which a function of a vehicle corresponding to a part of the identified vehicle or a group of parts identified above is performed.
5. In paragraph 1, The step of identifying the meaning of the above text data is: The step of identifying change information for at least one of a user's first-person view angle and a user's first-person view distance corresponding to a vehicle GUI being output through the user terminal based on the identified meaning by the server; The step of generating the output control signal of the above GUI is: A control method, wherein the server generates an output control signal of a GUI that changes at least one of a user's first-person view angle and a user's first-person view distance corresponding to a car GUI being output through the user terminal based on the identified change information.
6. In paragraph 1, The step of identifying the meaning of the above text data is: The step of the server identifying a part or a group of parts of the automobile included in the text data based on the identified meaning; The steps for generating the above response data are: A control method, wherein the server generates response data including a description of a part of the identified vehicle or a group of parts identified.
7. In paragraph 6, The step of identifying the meaning of the above text data is: The step of the server identifying a function of the vehicle corresponding to a part of the identified vehicle or a group of parts identified; The steps for generating the above response data are: A control method, wherein the server generates response data including a description of the functions of the identified vehicle.
8. In paragraph 1, The step of outputting the above automobile GUI and the above answer data and providing them to the user is as follows. The above user terminal outputs the automobile GUI through a display and provides it to the user, A control method in which the user terminal outputs a response voice corresponding to the response data through a speaker and provides it to the user.
9. A method for controlling an electronic device providing interactive artificial intelligence-based car infotainment, A step of identifying text data corresponding to the user command data when user command data is acquired through a microphone while outputting a car GUI through a display; A step of identifying the meaning of the above text data; A step of generating an output control signal of an automobile GUI corresponding to the text data based on the identified meaning; A step of generating answer data corresponding to the text data based on the identified meaning; and A control method, comprising: a step of providing a vehicle GUI and the answer data corresponding to the output control signal to a user by outputting the same through at least one of a display and a speaker.
10. A non-transitory computer-readable recording medium storing at least one instruction that is executed by a processor of an electronic device to cause the electronic device to perform the control method of claim 9.
Citation Information
Patent Citations
User manual providing terminal, server and method using augmented reality
KR101180278B1
Fishing reel
KR1020230141550A
System and method for automatically trading virtual currency
KR1020250010368A
Method for providing speech recognition based product guidance service using user manual
KR102585545B1
System and the controlling method thereof for providing interactive artificial intelligence-based automobile infotainment
KR102736282B1