Electronic device, method, and non-transitory computer-readable recording medium for providing content processed on basis of driving situation
The electronic device uses generative AI to process voice utterances and generate image content for display during vehicle travel, addressing the lack of interactive content generation in existing systems and enhancing user experience.
Patent Information
- Application Number
- PCT/KR2025/006209
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-22
- Filing Date
- 2025-05-08
- Publication Date
- 2025-12-26
AI Technical Summary
Existing systems fail to effectively generate and display image content based on voice utterances during vehicle travel, limiting the engagement and entertainment options for users.
An electronic device equipped with generative AI models processes audio data to identify voice utterances, generates corresponding image content, and displays it on an external device during vehicle travel, enhancing user experience.
Enables dynamic generation and display of image content relevant to audio data, improving user engagement and entertainment during vehicle journeys.
Smart Images

Figure KR2025006209_26122025_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transitory computer-readable recording medium for providing processed content based on driving conditions
[0001] The following descriptions relate to electronic devices, methods, and non-transitory computer-readable recording media that provide processed content based on driving conditions.
[0002] Artificial intelligence is a field of computer engineering and information technology that studies how to enable computers to think, learn, and develop themselves in ways that human intelligence can. It means enabling computers to imitate human intelligent behavior.
[0003] Artificial intelligence is intended to simulate human (or biological) neural activity, such as perception and / or inference, and can be implemented by hardware, software, or a combination of these designed to perform computations for simulating neural activity.
[0004] Artificial intelligence can generate new content (e.g., text, images, music, audio, video) based on instructions. Artificial intelligence that generates new content based on instructions can be referred to as generative artificial intelligence.
[0005] An electronic device is disclosed. The electronic device may include: at least one processor including communication circuitry, processing circuitry; and a memory storing instructions and including one or more storage media. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify original content for generating content to be played while a user of the electronic device is in a vehicle. The original content may include video content and / or audio content. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain audio data from the original content. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a voice utterance representing a subject of the audio data from among a plurality of voice utterances included in the audio data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate image content depicting the content of the spoken utterance by inputting a prompt based on the spoken utterance into a generative AI model. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit the image content to an external electronic device of the vehicle via the communication circuitry, such that the image content is displayed on a display included in the external electronic device of the vehicle during at least a portion of a time period during which the spoken utterance is played through a speaker included in the external electronic device of the vehicle.
[0006] An electronic device included in a vehicle is disclosed. The electronic device may include at least one processor including a speaker, a display, a communication circuit, and a processing circuit; and a memory storing instructions and including one or more storage media. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain, from an external electronic device via the communication circuit, information indicating original content to be played while a user of the electronic device is in the vehicle. The original content may include video content and / or audio content. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain audio data from the original content based on the information. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a voice utterance representing a subject of the audio data from among a plurality of voice utterances included in the audio data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate image content depicting the content of the spoken utterance by inputting a prompt based on the spoken utterance into a generative AI model. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the image content through the display during at least a portion of a time period during which the spoken utterance is played through the speaker.
[0007] A method is disclosed. The method can be performed in an electronic device including a communication circuit. The method can include an operation of identifying original content for generating content to be played while a user of the electronic device is in a vehicle. The original content can include video content and / or audio content. The method can include an operation of obtaining audio data from the original content. The method can include an operation of identifying a voice utterance representing a subject of the audio data among a plurality of voice utterances included in the audio data. The method can include an operation of generating image content depicting the content of the voice utterance by inputting a prompt based on the voice utterance into a generative AI model. The method can include an operation of transmitting the image content to an external electronic device of the vehicle through the communication circuit so that the image content is displayed on a display included in the external electronic device of the vehicle during at least a portion of a time period during which the voice utterance is played through a speaker included in the external electronic device of the vehicle.
[0008] A method is disclosed. The method can be performed in an electronic device included in a vehicle, the electronic device including a speaker, a display, and a communication circuit. The method can include an operation of obtaining, from an external electronic device via the communication circuit, information indicating original content to be played while a user of the electronic device is in the vehicle. The original content can include video content and / or audio content. The method can include an operation of obtaining audio data from the original content based on the information. The method can include an operation of identifying a voice utterance representing a subject of the audio data among a plurality of voice utterances included in the audio data. The method can include an operation of generating image content depicting the content of the voice utterance by inputting a prompt based on the voice utterance into a generative AI model. The method can include an operation of displaying the image content through the display during at least a portion of a time period during which the voice utterance is played through the speaker.
[0009] A non-transitory computer-readable storage medium is disclosed. The non-transitory computer-readable storage medium may store a program including instructions. The instructions, when individually or collectively executed by at least one processor of an electronic device including communication circuitry, may cause the electronic device to identify original content for generating content to be played while a user of the electronic device is in a vehicle. The original content may include video content and / or audio content. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain audio data from the original content. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a voice utterance representing a subject of the audio data from among a plurality of voice utterances included in the audio data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate image content depicting the content of the spoken utterance by inputting a prompt based on the spoken utterance into a generative AI model. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit the image content to an external electronic device of the vehicle via the communication circuitry, such that the image content is displayed on a display included in the external electronic device of the vehicle during at least a portion of a time period during which the spoken utterance is played through a speaker included in the external electronic device of the vehicle.
[0010] A non-transitory computer-readable recording medium is disclosed. The non-transitory computer-readable recording medium may store a program including instructions. The instructions, when individually or collectively executed by at least one processor of an electronic device included in a vehicle including a speaker, a display, and a communication circuit, may cause the electronic device to obtain, from an external electronic device via the communication circuit, information indicating original content to be played while a user of the electronic device is in the vehicle. The original content may include video content and / or audio content. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain audio data from the original content based on the information. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a voice utterance representing a subject of the audio data from among a plurality of voice utterances included in the audio data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate image content depicting the content of the spoken utterance by inputting a prompt based on the spoken utterance into a generative AI model. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the image content through the display during at least a portion of a time period during which the spoken utterance is played through the speaker.
[0011] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.
[0012] Figure 2 is a block diagram of an electronic device according to one embodiment.
[0013] FIG. 3 is a diagram illustrating an operation of an electronic device generating processed content according to one embodiment.
[0014] FIG. 4 is a diagram illustrating an operation of an electronic device generating processed content that summarizes content according to one embodiment.
[0015] FIG. 5 is a diagram illustrating an operation of an electronic device to determine content to be output based on a driving situation, according to one embodiment.
[0016] FIG. 6 is a diagram illustrating an operation of an electronic device updating a playlist according to a driving situation, according to one embodiment.
[0017] FIG. 7 illustrates an example of processed content output from a vehicle according to one embodiment.
[0018] FIGS. 8A and 8B are block diagrams of an electronic device and an external electronic device, according to one embodiment.
[0019] FIG. 9A is a flowchart illustrating the operation of an electronic device according to one embodiment.
[0020] FIG. 9b is a flowchart illustrating the operation of an electronic device according to one embodiment.
[0021] FIG. 10 is a diagram illustrating the structure of a generative AI model of an electronic device according to one embodiment.
[0022] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments.
[0023] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0024] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0025] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0026] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0027] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0028] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0029] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0030] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0031] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0032] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0033] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0034] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0035] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0036] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0037] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0038] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0039] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0040] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 664 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 6 ms or less for round trip) for URLLC realization.
[0041] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0042] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0043] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0044] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0045] Figure 2 is a block diagram of an electronic device according to one embodiment.
[0046] The electronic device (101) of FIG. 2 may correspond to the electronic device (101) of FIG. 1. The electronic device (101) of FIG. 2 may be described with reference to the electronic device (101) of FIG. 1.
[0047] Referring to FIG. 2, the electronic device (101) can communicate with an external electronic device (210). In one embodiment, the external electronic device (210) may be an electronic device mounted on a vehicle (200).
[0048] In one embodiment, the electronic device (101) may include a processor (120), a memory (130), a display (260), and a communication circuit (290). The processor (120) of FIG. 2 may correspond to the processor (120) of FIG. 1. The memory (130) of FIG. 2 may correspond to the memory (130) of FIG. 1. The display (260) of FIG. 2 may correspond to the display module (160) of FIG. 1. The communication circuit (290) of FIG. 2 may correspond to the communication module (190) of FIG. 1. In one embodiment, the program (140) of the memory (130) may include a path analysis unit (201), a content analysis unit (203), a prompt generation model (205), and a GEN AI (generative artificial intelligence) model (207). However, the present invention is not limited thereto. For example, the GEN AI model (207) may not be included in the electronic device (101). For example, the GEN AI model (207) may be included in a device external to the electronic device (101) (e.g., the electronic device (102), the electronic device (108), and / or the server (108) of FIG. 1).
[0049] In one embodiment, the external electronic device (210) may include a processor (225), a memory (235), a display (265), a communication circuit (295), and a speaker (255). The processor (225) of FIG. 2 may correspond to the processor (120) of FIG. 1. The memory (235) of FIG. 2 may correspond to the memory (130) of FIG. 1. The display (265) of FIG. 2 may correspond to the display module (160) of FIG. 1. The communication circuit (295) of FIG. 2 may correspond to the communication module (190) of FIG. 1. The speaker (255) of FIG. 2 may correspond to the audio output module (155) and / or the audio module (170) of FIG. 1.
[0050] In one embodiment, the path analysis unit (201), the content analysis unit (203), the prompt generation model (205), and the GEN AI model (207) may correspond to the application (146) of FIG. 1. For example, the path analysis unit (201), the content analysis unit (203), the prompt generation model (205), and the GEN AI model (207) may include instructions that may be executed by the processor (120).
[0051] In one embodiment, the path analysis unit (201) can identify a driving path. For example, the path analysis unit (201) can identify a driving path based on information obtained from the vehicle (200) (or an external electronic device (210)) and / or the electronic device (101). For example, the driving path can indicate a road on which the vehicle (200) is currently located and / or a road on which the vehicle (200) will drive in the future.
[0052] For example, information obtained from the automobile (200) (or the external electronic device (210)) and / or the electronic device (101) may include location information of the automobile (200) (or the external electronic device (210)) and / or the electronic device (101). For example, information obtained from the automobile (200) (or the external electronic device (210)) and / or the electronic device (101) may include current location, moving direction, moving speed, current driving time, and / or expected time required to move to a destination of the automobile (200) (or the external electronic device (210)) and / or the electronic device (101).
[0053] For example, information obtained from the vehicle (200) (or external electronic device (210)) and / or the electronic device (101) may include information about a destination. For example, information obtained from the vehicle (200) (or external electronic device (210)) and / or the electronic device (101) may include information about a driving environment along a route to a destination. For example, the driving environment may indicate whether there is a traffic jam (or a congested section), a speed limit, a driving speed, whether there is a danger zone (e.g., an accident-prone area), road width, the number of lanes, the direction of travel of a lane (e.g., going straight, turning left, turning right, and / or turning U-turn), a traffic signal situation, and / or a signal waiting time.
[0054] For example, information obtained from the vehicle (200) (or external electronic device (210)) and / or electronic device (101) may include the driving style of the vehicle (200) (e.g., highway priority, toll-free road priority, driving time priority, etc.).
[0055] For example, information obtained from the vehicle (200) (or external electronic device (210)) and / or the electronic device (101) may include autonomous driving performance of the vehicle (200) (e.g., autonomous driving level) and / or autonomous driving status (e.g., whether autonomous driving mode is on or off). Depending on the embodiment, the autonomous driving mode may be changed by a user's input (e.g., a command to turn autonomous driving on or off) or by the judgment of the vehicle (200).
[0056] In one embodiment, the path analysis unit (201) may obtain an input indicating a destination. In one embodiment, the path analysis unit (201) may obtain a driving route from the current location of the vehicle (200) to the destination based on obtaining the input indicating the destination. For example, a user may input information about the destination into the electronic device (101) while driving the vehicle (200) through the path analysis unit (201). In one embodiment, the driving route may include one or more roads that the vehicle (200) will pass through while moving from the current location to the destination.
[0057] In one embodiment, the path analysis unit (201) may classify the driving path into one or more segments. The one or more segments may include a first segment in which the degree of driving intervention of the vehicle (200) that the vehicle (200) requires from the user satisfies a first criterion. The one or more segments may include a second segment in which the degree of driving intervention of the vehicle (200) that the vehicle (200) requires from the user satisfies a second criterion that is lower than the first criterion. For example, when the user must grip the steering wheel of the vehicle (200) (or, when the user must steer the steering wheel of the vehicle (200)) (or, when the user must operate the (foot) brake of the vehicle (200)) while the vehicle (200) is driving and / or stopped, the degree of driving intervention of the vehicle (200) may be determined to meet the first criterion. For example, while the vehicle (200) is not in autonomous driving mode (or while in autonomous driving mode at level 2 or lower) (or, in autonomous driving mode at level 3, at the request of the vehicle (200)), the degree of driving intervention of the vehicle (200) may be determined to meet the first criterion. For example, if the degree of driving intervention of the vehicle (200) does not meet the first criterion, the degree of driving intervention of the vehicle (200) may be determined to meet the second criterion. For example, while the vehicle (200) is in autonomous driving mode (or, while in autonomous driving mode at level 4 or higher) (or, in autonomous driving mode at level 3, at the request of the vehicle (200)), the degree of driving intervention of the vehicle (200) may be determined to meet the second criterion.
[0058] In one embodiment, the path analysis unit (201) may identify an estimated time required to pass through each of one or more sections. For example, the path analysis unit (201) may identify an estimated time required to pass through each of one or more sections based on the driving environment and / or movement speed.
[0059] In one embodiment, the content analysis unit (203) can identify content to be output (or played) through the external electronic device (210) of the automobile (200). In one embodiment, the content analysis unit (203) can identify a list of contents to be output (or played) through the external electronic device (210) of the automobile (200). In one embodiment, the content being output through the external electronic device (210) may include the content being directly output through the external electronic device (210) or the content processed from the content being output through the external electronic device (210). In one embodiment, the content may be video content and / or audio content. For example, the video content may include video data and audio data. For example, the audio content may include audio data.
[0060] In one embodiment, the content analysis unit (203) may identify content to be output (or played) through the external electronic device (210) based on the communication connection between the electronic device (101) and the external electronic device (210). In one embodiment, the content analysis unit (203) may identify content to be output (or played) through the external electronic device (210) based on identifying a communication connection between the external electronic device (210) and a vehicle (200) registered as a vehicle of the user of the electronic device (101). In one embodiment, the content analysis unit (203) may identify video content for generating content to be played while the user of the electronic device (101) is riding in the vehicle (200). In one embodiment, the content analysis unit (203) may identify content being played on the electronic device (101) as content to be output (or played) through the external electronic device (210) based on the communication connection between the electronic device (101) and the external electronic device (210). In one embodiment, the content analysis unit (203) may identify content being played on the external electronic device (210) of the automobile (200) as content to be output (or played) through the external electronic device (210) based on the communication connection between the electronic device (101) and the external electronic device (210).
[0061] In one embodiment, the content analysis unit (203) can identify the purpose, type, content, presence or absence of video (or image), and / or playback time of the content identified through the external electronic device (210). For example, the purpose of the content may include viewing by the user of the electronic device (101) and / or viewing by passengers. The purpose of the content is not limited to these and may have various purposes. For example, the type of content may include movies, news, music videos, pop songs, and / or audio books. The type of content is not limited to these and may have various types.
[0062] In one embodiment, the content analysis unit (203) may extract (or analyze) the content of the content identified through the external electronic device (210). For example, the content analysis unit (203) may be an AI model (e.g., a stable diffusion model) for performing a task (e.g., captioning) of extracting words (or sentences) from the content of the content (e.g., text, image, or video) to describe the content. In one embodiment, the content analysis unit (203) may generate at least one sentence (or at least one word) that summarizes the content of the content from the content of the content.
[0063] In one embodiment, the content analysis unit (203) may generate at least one sentence (or at least one word) indicating the purpose, type, content, presence or absence of video (or image), and / or playback time of the content identified through the external electronic device (210).
[0064] In one embodiment, the content analysis unit (203) may obtain audio data from video content, which is original content. In one embodiment, the content analysis unit (203) may identify one or more voices (or voice utterances) representing a subject (or material) (or keyword) of the video content (or audio data) among a plurality of voices (or voice utterances) included in the audio data. For example, the content analysis unit (203) may be a program (or AI model) that performs an algorithm capable of processing natural language (e.g., morphological analysis, and / or tokenization for syntax analysis, and / or feature value extraction).
[0065] In one embodiment, the prompt generation model (205) may be an AI model (e.g., a stable diffusion model) for generating prompts (or task instructions) to be input to the GEN AI model (207) based on input.
[0066] In one embodiment, the prompt generation model (205) can obtain (or generate) a prompt. In one embodiment, the prompt generation model (205) can generate a prompt for obtaining (or generating) content to be output (or played) through an external electronic device (210). In one embodiment, the prompt can include data for guiding the generation of an image based on a message. For example, the prompt can include at least one sentence (or at least one word) for specifying content to be generated through the GEN AI model (207). For example, the prompt can include at least a portion of the content. Hereinafter, content generated through the GEN AI model (207) can be referred to as processed content. Content not generated through the GEN AI model (207) can be referred to as original content.
[0067] For example, the prompt generation model (205) can generate a prompt for generating processed audio content from original content through the GEN AI model (207). For example, the prompt generation model (205) can generate a prompt for generating processed image content from processed audio content through the GEN AI model (207). However, the present invention is not limited thereto. For example, the prompt generation model (205) can generate a prompt for generating processed image content from original content through the GEN AI model (207).
[0068] For example, at least one sentence (or at least one word) included in the prompt may describe data (e.g., original content, content of the original content, and / or a summary of the original content) that the GEN AI model (207) will reference when generating the processed content. For example, at least one sentence (or at least one word) included in the prompt may indicate data that the GEN AI model (207) will reference when generating the processed content.
[0069] In one embodiment, the prompt generation model (205) may generate a prompt for obtaining (or generating) processed content to be output (or played) through an external electronic device (210) based on information obtained from the content analysis unit (203). For example, the information obtained from the content analysis unit (203) may include the purpose, type, content, presence or absence of video (or image), and / or playing time of the original content.
[0070] For example, the prompt may include one or more words indicating the duration of the processed content generated from the original content. For example, the prompt may include one or more words indicating whether to generate processed image content from the original content (or processed audio content).
[0071] In one embodiment, the prompt generation model (205) may generate a prompt for obtaining (or generating) processed content to be output (or played) through an external electronic device (210) based on information obtained from the path analysis unit (201). For example, the information obtained from the path analysis unit (201) may include autonomous driving performance (e.g., autonomous driving level) of the vehicle (200) and / or whether autonomous driving is performed (e.g., whether the autonomous driving mode is on or off). For example, the information obtained from the path analysis unit (201) may include the current location, moving direction, moving speed, current driving time, and / or the expected time required to move to the destination of the vehicle (200) (or the external electronic device (210)) and / or the electronic device (101). For example, the information obtained from the path analysis unit (201) may include information on the driving environment along the path to the destination.
[0072] In one embodiment, the prompt generation model (205) can generate a prompt based on a plurality of voices (or voice utterances) included in original content (e.g., video content or audio content). For example, the prompt based on the plurality of voices (or voice utterances) can include one or more words representing a topic (or material) (or keyword) of the original content (e.g., video content or audio content). For example, the prompt based on the plurality of voices (or voice utterances) can include one or more words for instructing the GEN AI model (207) to generate image content describing the content of the voice utterance.
[0073] In one embodiment, the prompt generation model (205) may generate a prompt based on the remaining driving time. In one embodiment, the prompt generation model (205) may generate a prompt including one or more words indicating an instruction to make the playing time of the processed content (e.g., the processed audio content) shorter than the remaining driving time when the remaining driving time is determined to be shorter than the playing time of the original content (e.g., the video content or the audio content). However, the present invention is not limited thereto. In one embodiment, the prompt generation model (205) may generate a prompt including one or more words indicating an instruction to make the playing time of the processed content (e.g., the processed audio content) correspond to the remaining driving time (e.g., the playing time of the processed content is equal to or shorter than the remaining driving time).
[0074] In one embodiment, the prompt generation model (205) may generate prompts based on the vehicle's (200) occupancy status (e.g., occupants and / or number of occupants). Accordingly, the content generation style generated by the GEN AI model (207) may be changed based on the vehicle's (200) occupancy status (e.g., occupants and / or number of occupants).
[0075] In one embodiment, the GEN AI model (207) may include an AI model including multiple parameters related to a neural network having a structure based on an encoder and a decoder, such as a transformer. In one embodiment, the GEN AI model (207) may include a bi-directional model based on learning for an encoder (e.g., bidirectional encoder representations from transformers (BERT)) or an auto-encoding model (e.g., a diffusion model). In one embodiment, the GEN AI model (207) may include an auto-regressor model based on learning for a decoder (e.g., a generative pre-trained transformer (GPT)). In one embodiment, the GEN AI model (207) may include a sequence-to-sequence model (e.g., stable diffusion, DALL-E 2) based on learning for an encoder and a decoder. In one embodiment, the GEN AI model (207) may include, but is not limited to, a large language model (LLM) for processing natural language based on a massive number of parameters. The GEN AI model (207) may include parameters for driving a neural network such as a convolutional neural network (CNN), a recurrent neural network (RNN), a feedforward neural network (FNN), and / or a long short-term memory (LSTM).
[0076] In one embodiment, the GEN AI model (207) can generate processed content based on a prompt. For example, the GEN AI model (207) can generate processed audio content from original content based on the prompt. For example, the processed audio content can include audio content and audio signals included in the original content. For example, the processed audio content can include an audio signal generated after the prompt is generated.
[0077] In one embodiment, the GEN AI model (207) can generate processed image content from processed audio content based on a prompt. In one embodiment, the GEN AI model (207) can generate processed image content based on at least a portion of the original content (or processed audio content) based on information analyzed by the path analysis unit (201) included in the prompt and / or information analyzed by the content analysis unit (203).
[0078] For example, the processed image content may include a still image (or a plurality of still images) representing (or depicting) audio included in the original content (or the processed audio content). However, the present invention is not limited thereto. For example, the GEN AI model (207) may generate processed image content from the original content based on a prompt. For example, the processed image content may be a video representing (or depicting) audio included in the original content (or the processed audio content). For example, the processed image content may include an image that is distinct from video content included in the original content. For example, the processed image content may include an image that is generated after the prompt is generated.
[0079] In one embodiment, the processor (120) may output the generated processed content. For example, the processor (120) may transmit the generated processed content to an external electronic device (210) of the automobile (200) via the communication circuit (290). The external electronic device (210) may output the received generated processed content via the display (265) and / or the speaker (255). However, the present invention is not limited thereto. For example, the processor (120) may output the generated processed content via the electronic device (101). For example, the processor (120) may output the generated processed content via the display (260) and / or the audio output module (155).
[0080] In one embodiment, the processor (120) may transmit the processed image content to an external electronic device (210) of the automobile (200) via the communication circuit (290) such that the processed image content is displayed via a display (265) included in the automobile (200) during at least a portion of the time period during which the voice utterance is played via a speaker (255) included in the automobile (200).
[0081] In one embodiment, the processor (120) may cause an external electronic device (210) of the vehicle (200) to display one of the processed image content or the video content included in the original content through a display (265) included in the vehicle (200). For example, the processor (120) may transmit the processed image content to the external electronic device (210) of the vehicle (200) through the communication circuit (290) based on determining that a specified first condition is met. For example, the processor (120) may transmit the video content to the external electronic device (210) of the vehicle (200) through the communication circuit (290) based on determining that a specified second condition, distinct from the specified first condition, is met.
[0082] In one embodiment, the processor (120) may determine, based on location information of the vehicle (200), whether the vehicle (200) is located in a first section or a second section among sections on a driving route. In one embodiment, the processor (120) may determine, based on the driving route, that the vehicle (200) is located in a first section that satisfies a first criterion in which the degree of driving intervention required from the user is higher than the first criterion. In one embodiment, the processor (120) may determine, based on the driving route, that the vehicle (200) is located in a second section that satisfies a second criterion in which the degree of driving intervention required from the user is lower than the first criterion. For example, the first section may be a section where the autonomous driving mode is turned off (e.g., an alley). For example, the second section may be a section where the autonomous driving mode is turned on (e.g., a national road, a local road, a highway). However, the present invention is not limited thereto. For example, Section 1 may be a section that is a hazardous area (e.g., an accident-prone area). For example, Section 2 may be a section other than Section 1. For example, Section 1 may be a section with a speed limit and / or a speed limit higher than a designated speed. For example, Section 2 may be a section other than Section 1 (e.g., a congested area).
[0083] In one embodiment, the processor (120) may transmit processed image content to an external electronic device (210) of the vehicle (200) through the communication circuit (290) so that the processed image content is displayed through a display (265) included in the vehicle (200) when the vehicle (200) is determined to be located in the first section. In one embodiment, the processor (120) may transmit video content to an external electronic device (210) of the vehicle (200) through the communication circuit (290) so that video content included in the original content is displayed through the display (265) instead of the processed image content when the vehicle (200) is determined to be located in the second section.
[0084] In one embodiment, the processor (120) may identify the driving state of the vehicle (200). In one embodiment, when it is determined that the vehicle (200) is driving (or driving at a specified speed or higher), the processor (120) may transmit the processed image content to an external electronic device (210) of the vehicle (200) through the communication circuit (290) so that the processed image content is displayed through a display (265) included in the vehicle (200). In one embodiment, when it is determined that the vehicle (200) is stopped (or waiting for a signal), the processor (120) may transmit the video content to the external electronic device (210) of the vehicle (200) through the communication circuit (290) so that the video content included in the original content is displayed through the display (265) instead of the processed image content. However, the present invention is not limited thereto. In one embodiment, the processor (120) may transmit processed image content to an external electronic device (210) of the vehicle (200) through the communication circuit (290) when the distance between the vehicle (200) and the vehicle in front is determined to be less than a specified distance. In one embodiment, the processor (120) may transmit video content to an external electronic device (210) of the vehicle (200) through the communication circuit (290) when the distance between the vehicle (200) and the vehicle in front is determined to be greater than a specified distance.
[0085] In one embodiment, the processor (120) may shuffle a playlist of contents such that the processed content is output through the vehicle (200) while a first condition is met, and the original content is output through the vehicle (200) while a second condition is met. For example, the processor (120) may transmit the original content or the processed content to an external electronic device (210) of the vehicle (200) through the communication circuit (290) based on the shuffled playlist.
[0086] For example, the processor (120) may transmit the processed audio content and the processed image content to the external electronic device (210) of the vehicle (200) through the communication circuit (290) so that the processed audio content and the processed image content are output while the first condition is met. For example, the processor (120) may transmit the original audio content and the original image content to the external electronic device (210) of the vehicle (200) through the communication circuit (290) so that the original audio content and the original image content are output while the second condition is met.
[0087] For example, the processor (120) may shuffle a playlist so that processed content is transmitted to an external electronic device (210) of the vehicle (200) via the communication circuit (290) when the vehicle (200) is determined to be located in a first section. For example, the processor (120) may shuffle a playlist so that original content is transmitted to an external electronic device (210) of the vehicle (200) via the communication circuit (290) when the vehicle (200) is determined to be located in a second section.
[0088] According to an embodiment, the content analysis unit (203) may obtain audio data from video content, which is original content. In one embodiment, the content analysis unit (203) may identify music (or background music) included in the audio data. In one embodiment, the content analysis unit (203) may extract music (or background music) from the audio data.
[0089] In one embodiment, the content analysis unit (203) may request a prompt generation model (205) to generate information (or metadata) (e.g., title, year, artist, album, album art) of music (or background music) included in the audio data based on the fact that the audio data does not contain any speech utterances other than music (or background music).
[0090] In one embodiment, the prompt generation model (205) may generate a prompt (or work instruction) to be input to the GEN AI model (207) based on an input representing music (or background music) included in the audio data. In one embodiment, the prompt generation model (205) may generate a prompt for obtaining (or generating) processed content to be output (or played) through an external electronic device (210) based on music (or background music) included in the audio data obtained from the content analysis unit (203). For example, the prompt may generate a prompt for obtaining (or generating) processed content representing information of music (or background music) included in the audio data.
[0091] In one embodiment, the prompt generation model (205) may generate a prompt to generate information (or metadata) (e.g., title, year, artist, album, album art) of music (or background music) included in the audio data based on the fact that the audio data does not contain any speech utterances other than music (or background music).
[0092] In one embodiment, the GEN AI model (207) may generate processed image content representing information about the music (or background music) included in the audio data from the music (or background music) included in the audio data based on a prompt. For example, the processed image content may include a still image (or a plurality of still images) representing (or depicting) the music (or background music) included in the audio data. For example, the processed image content may represent information (or metadata) (e.g., title, year, artist, album, album art) of the music (or background music) included in the audio data.
[0093] As described above, the electronic device (101) can output images generated based on the content, or output videos included in the content, depending on the driving situation. Accordingly, in situations where the driver is required to concentrate, the electronic device (101) can output images instead of videos, thereby avoiding disruption to the driver's concentration. Furthermore, in situations where the driver is not required to concentrate (e.g., during autonomous driving), the electronic device (101) can output videos, thereby allowing the driver to enjoy the content richly.
[0094] FIG. 3 is a diagram illustrating an operation of an electronic device generating processed content according to one embodiment.
[0095] FIG. 3 may be described with reference to the electronic device (101) of FIG. 2 (or components of the electronic device (101)). The operations described with reference to FIG. 3 may be performed as instructions included in the path analysis unit (201), content analysis unit (203), prompt generation model (205), and / or GEN AI model (207) of FIG. 2 are executed by the processor (120) of the electronic device (101).
[0096] Referring to FIG. 3, the original content (310) may include video content (e.g., video content (420) of FIG. 4) and audio content (e.g., audio content (440) of FIG. 4). In one embodiment, the video content may include video frames representing one or more visual objects (311, 313). In one embodiment, the audio content may include audio segments including one or more utterances (315).
[0097] In one embodiment, the electronic device (101) may generate processed audio content (320) including one or more utterances (321) based on original content (310). For example, one or more utterances (321) included in the processed audio content (320) may correspond to one or more utterances (315) included in the original content (310). For example, one or more utterances (321) included in the processed audio content (320) may correspond to an utterance that summarizes one or more utterances (315) included in the original content (310).
[0098] In one embodiment, the electronic device (101) may generate processed image content (330) based on processed audio content (320). For example, the processed image content (330) may be an image (331) that includes a visual object (e.g., a graph, a building, a person) that depicts the content of the processed audio content (320). For example, the processed image content (330) may be a text-based image (335) that includes texts that summarize (or represent) the content of the processed audio content (320).
[0099] In FIG. 3, the processed image content (330) is illustrated as being generated based on the processed audio content (320), but this is merely an example. In some embodiments, the processed image content (330) may be generated based on the original content (310).
[0100] FIG. 4 is a diagram illustrating an operation of an electronic device generating processed content that summarizes content according to one embodiment.
[0101] FIG. 4 may be described with reference to the electronic device (101) of FIG. 2 (or components of the electronic device (101)). The operations described with reference to FIG. 4 may be performed as instructions included in the path analysis unit (201), content analysis unit (203), prompt generation model (205), and / or GEN AI model (207) of FIG. 2 are executed by the processor (120) of the electronic device (101).
[0102] Referring to FIG. 4, the original content (410) may include video content (420) and / or audio content (430). For example, the video content (420) may be a single video track. For example, the audio content (430) may be a single audio track. For example, the audio content (430) may be an audio track set to be played in synchronization with the display of the video content (420).
[0103] In one embodiment, the video content (420) may include one or more sets of video frames (421, 423, 425, 427, 429). In one embodiment, the audio content (430) may include one or more sets of audio segments (431, 433, 435, 437, 439) configured to be played at synchronized points in time with each of the one or more sets of video frames (421, 423, 425, 427, 429).
[0104] In one embodiment, the electronic device (101) can generate processed content (450) from original content (410). In one embodiment, the electronic device (101) can generate processed content (450) having a shorter playback time than the original content (410). In one embodiment, the electronic device (101) can generate processed audio content (460) based on the original content (410). In one embodiment, the electronic device (101) can generate processed image content (470) based on the processed audio content (460) (or the original content (410)). For example, the processed image content (470) can include an image that is set to be displayed in synchronization with the playback of the processed audio content (460).
[0105] In one embodiment, the processed audio content (460) may include one or more sets of processed audio segments (461, 463, 465, 467, 469). In one embodiment, the processed audio content (460) may include one or more sets of processed audio segments (461, 463, 465, 467, 469) generated by the GEN AI model (207) based on each of one or more playback sets. In one embodiment, a playback set may include a set of a pair of video frames and a set of audio segments that are played at the same point in time. For example, a set of video frames (421) and a set of audio segments (431) may be understood as one playback set. For example, a set of video frames (423) and a set of audio segments (433) may be understood as one playback set. However, the present invention is not limited thereto. In one embodiment, the processed audio content (460) may include one or more sets of processed audio segments (461, 463, 465, 467, 469) generated by the GEN AI model (207) based on each of one or more sets of video frames (421, 423, 425, 427, 429). In one embodiment, the processed audio content (460) may include one or more sets of processed audio segments (461, 463, 465, 467, 469) generated by the GEN AI model (207) based on each of one or more sets of audio segments (431, 433, 435, 437, 439).
[0106] In one embodiment, the processed image content (470) may include one or more sets of processed images (471, 473, 475, 477, 479). In one embodiment, the processed image content (470) may include one or more sets of processed images (471, 473, 475, 477, 479) generated by the GEN AI model (207) based on each of one or more sets of processed audio segments (461, 463, 465, 467, 469). However, the present invention is not limited thereto. In one embodiment, the processed image content (470) may include one or more sets of processed images (471, 473, 475, 477, 479) generated by the GEN AI model (207) based on each of one or more playback sets. In one embodiment, the processed image content (470) may include one or more sets of processed images (471, 473, 475, 477, 479) generated by the GEN AI model (207) based on each of one or more sets of video frames (421, 423, 425, 427, 429). In one embodiment, the processed image content (470) may include one or more sets of processed images (471, 473, 475, 477, 479) generated by the GEN AI model (207) based on each of one or more sets of audio segments (431, 433, 435, 437, 439).
[0107] In one embodiment, the electronic device (101) may be configured such that the playback time of each of the sets (461, 463, 465, 467, 469) of the processed audio segments may be less than or equal to the playback time of each of the playback sets of the original content (410). For example, the playback time of some of the sets (461, 463, 465, 467, 469) of the processed audio segments may be equal to the playback time of a playback set corresponding to some of the playback sets of the original content (410). For example, the playback time of other of the sets (461, 463, 465, 467, 469) of the processed audio segments may be shorter than the playback time of a playback set corresponding to other of the playback sets of the original content (410).
[0108] In one embodiment, the electronic device (101) may generate the processed content (450) based on a prompt to ensure that the playback time of the processed content (450) is shorter than the remaining driving time. However, the present invention is not limited thereto. In one embodiment, the electronic device (101) may generate the processed content (450) based on a prompt to ensure that the playback time of the processed content (450) is within the time allowed for playback of the original content (410). The time allowed for playback of the original content (410) may be less than or equal to an expected time required to pass through a first section where the degree of driving intervention of the vehicle (200) satisfies the first criterion. For example, since the vehicle (200) enters a different type of section (i.e., a section where the degree of driving intervention of the vehicle (200) does not meet the first criterion) after passing through the first section, the time allowed for playback of the original content (410) may be limited to a time less than the expected time required to pass through the first section where the degree of driving intervention of the vehicle (200) meets the first criterion.
[0109] According to an embodiment, the electronic device (101) may receive user input related to the content being output while outputting the content. For example, the user input related to the content being output may include a voice input to the electronic device (101) or a press input to a designated button, but is not limited thereto. The user input related to the content being output may include an input to an external electronic device (210).
[0110] In one embodiment, the electronic device (101) may receive a user input related to the output content, which requests an additional description (or additional material) for the output content. In one embodiment, the electronic device (101) may mark (or store) a playback time of the output content based on receiving the user input related to the output content. In one embodiment, the electronic device (101) may generate an additional description (or additional material) for the output content based on receiving the user input related to the output content. In one embodiment, the electronic device (101) may generate the additional description (or additional material) using a portion of the output content within a specified playback time range from a playback time point marked (or stored) for the output content. In one embodiment, the generation of additional description (or additional material) may be based on generating a prompt based on a portion of the content output within a specified playback time range from the playback point in time, and inputting the generated prompt into the GEN AI model (207). In one embodiment, the generated additional description (or additional material) may be processed audio content and / or processed image content.
[0111] In one embodiment, the electronic device (101) may output the generated additional description (or additional data) according to a predetermined condition. In one embodiment, the electronic device (101) may output the generated additional description (or additional data) when the driving situation of the vehicle (200) changes. In one embodiment, the electronic device (101) may output the generated additional description (or additional data) when the driving situation of the vehicle (200) is a first situation. In one embodiment, the electronic device (101) may output the generated additional description (or additional data) when the driving situation of the vehicle (200) is a second situation.
[0112] FIG. 5 is a diagram illustrating an operation of an electronic device to determine content to be output based on a driving situation, according to one embodiment.
[0113] FIG. 5 may be described with reference to the electronic device (101) of FIG. 2 (or components of the electronic device (101)). The operations described with reference to FIG. 5 may be performed as instructions included in the path analysis unit (201), content analysis unit (203), prompt generation model (205), and / or GEN AI model (207) of FIG. 2 are executed by the processor (120) of the electronic device (101).
[0114] Referring to FIG. 5, the electronic device (101) can identify whether the driving situation is a first situation or a second situation. For example, the first situation may be a situation in which the vehicle (200) is located in a first section where the degree of driving intervention of the vehicle (200) satisfies a first criterion. For example, the second situation may be a situation in which the vehicle (200) is located in a second section where the degree of driving intervention of the vehicle (200) required by the user is lower than the first criterion and satisfies a second criterion. For example, the first situation may include a manual driving mode or a semi-autonomous driving mode, which is a state in which user intervention is set to be greater than autonomous driving. For example, the second situation may include an autonomous driving mode, which is a state in which user intervention is set to be relatively less than autonomous driving.
[0115] For example, the first situation may be a situation in which the vehicle (200) is predicted (or identified) to be driving (or driving at a specified speed or higher). For example, the second situation may be a situation in which the vehicle (200) is predicted (or identified) to be stopped (or waiting for a signal). For example, the first situation may be a situation in which the user must directly drive. For example, the first situation may be a situation in which the vehicle (200) requires a low degree of driving intervention. For example, the first situation may include a situation in which the vehicle (200) is in a manual or semi-autonomous driving mode, requiring direct driving by the user. For example, the second situation may be a situation in which the vehicle (200) requires primarily driving rather than the user. For example, the second situation may be a situation in which the vehicle (200) requires a relatively high degree of driving intervention. For example, the second situation may include a situation in which the vehicle (200) is in an autonomous driving mode, requiring relatively little direct driving by the user.
[0116] For example, the first scenario may involve situations where the user must manually drive for safety reasons, such as narrow alleyways or roads requiring speed limits, where the risk of accidents during autonomous driving is relatively low. The second scenario may involve situations where the user must manually drive for safety reasons, such as national highways, local roads, or expressways, where the risk of accidents during autonomous driving is relatively low, and where relatively little driver intervention is required.
[0117] For example, the electronic device (101) can identify whether the driving situation is the first situation or the second situation, depending on the type of road on which the vehicle (200) is located. For example, the second situation may be a situation in which the vehicle (200) is driving on a national road, a local road, or an expressway. For example, the first situation may be a situation in which the vehicle (200) is driving on a general road or an alley.
[0118] For example, the electronic device (101) can predict (or identify) that the driving situation will be the first situation between time points (501, 503). For example, the electronic device (101) can predict (or identify) that the driving situation will be the second situation between time points (503, 505). For example, the electronic device (101) can predict (or identify) that the driving situation will be the first situation between time points (505, 507). For example, the electronic device (101) can predict (or identify) that the driving situation will be the second situation between time points (507, 509).
[0119] In one embodiment, the electronic device (101) may identify the content to be played based on the driving situation. For example, the electronic device (101) may identify the processed content (551, 555) as the content to be played, rather than the original content (511, 515), during a time interval (e.g., between time points (501, 503) and time points (505, 507)) in which the driving situation is predicted (or identified) to be the first situation. For example, the electronic device (101) may identify the original content (513, 517) as the content to be played, during a time interval (e.g., between time points (503, 505) and time points (507, 509)) in which the driving situation is predicted (or identified) to be the second situation.
[0120] In one embodiment, the electronic device (101) may transmit the identified playback content to an external electronic device (210) of the vehicle (200). For example, during a time interval between points (501, 503), the electronic device (101) may transmit the processed content (551) to the external electronic device (210) so that the processed content (551) is played through the external electronic device (210). For example, during a time interval between points (503, 505), the electronic device (101) may transmit the original content (513) to the external electronic device (210) so that the original content (513) is played through the external electronic device (210).
[0121] In FIG. 5, different original contents (511, 513, 515, 517) (or different processed contents (551, 555)) are played back in each situation, but this is only an example. For example, the original contents (511, 513, 515, 517) may be playback sets included in one single content. For example, the single content may be content including different chapters (or parts). For example, the original contents (511, 513, 515, 517) may form one single content. For example, when one video content includes different chapters (or scenes), the different chapters (or scenes) may be included in the original contents (511, 513, 515, 517). Corresponding, one video content may be one content including original contents (511, 513, 515, 517). For example, the electronic device (101) may generate processed content based on a portion of one single content to be played back during a first situation (e.g., a first chapter of one single content). For example, the electronic device (101) may output processed content corresponding to a portion of one single content (e.g., a first chapter of one single content) through an external electronic device (210) during the first situation. For example, the electronic device (101) may not generate processed content based on another portion of one single content to be played back during a second situation (e.g., a second chapter of one single content). For example, the electronic device (101) may output another portion of one single content through an external electronic device (210) during the second situation.
[0122] In one embodiment, the electronic device (101) can determine whether a situation has changed. In one embodiment, the electronic device (101) can adjust the length (or running time) of the content being played based on the determination that the situation has changed. In one embodiment, the determination of whether the situation has changed can be performed at a specified interval. In one embodiment, the determination of whether the situation has changed can be performed whenever an event occurs. Here, the event can include a path deviation (e.g., deviating from the path to the destination).
[0123] FIG. 6 is a diagram illustrating an operation of an electronic device updating a playlist according to a driving situation, according to one embodiment.
[0124] FIG. 6 may be described with reference to the electronic device (101) of FIG. 2 (or components of the electronic device (101)). The operations described with reference to FIG. 6 may be performed as instructions included in the path analysis unit (201), content analysis unit (203), prompt generation model (205), and / or GEN AI model (207) of FIG. 2 are executed by the processor (120) of the electronic device (101).
[0125] FIG. 6 illustrates, compared to FIG. 5, that the electronic device (101) can shuffle a playlist depending on the driving situation. For example, the processor (120) can transmit original content or processed content to an external electronic device (210) of the automobile (200) via the communication circuit (290) based on the shuffled playlist.
[0126] Referring to FIG. 6, the electronic device (101) can identify whether the driving situation is a first situation or a second situation. For example, the first situation may be a situation in which the vehicle (200) is located in a first section where the degree of driving intervention of the vehicle (200) satisfies a first criterion. For example, the second situation may be a situation in which the vehicle (200) is located in a second section where the degree of driving intervention of the vehicle (200) requested by the user is lower than the first criterion. For example, the first situation may be a situation in which the vehicle (200) is predicted (or identified) to be driving (or driving at a specified speed or higher). For example, the second situation may be a situation in which the vehicle (200) is predicted (or identified) to be stopped (or waiting for a signal).
[0127] For example, the electronic device (101) can predict (or identify) that the driving situation will be the first situation between time points (601, 603). For example, the electronic device (101) can predict (or identify) that the driving situation will be the second situation between time points (603, 605). For example, the electronic device (101) can predict (or identify) that the driving situation will be the first situation between time points (605, 607). For example, the electronic device (101) can predict (or identify) that the driving situation will be the second situation between time points (607, 609).
[0128] In one embodiment, the electronic device (101) can shuffle a playlist of contents (611, 613, 615, 617, 619, 621, 623) based on a driving situation. For example, the electronic device (101) may shuffle a playlist of contents (611, 613, 615, 617, 619, 621, 623) such that audio contents (613, 617, 619) or processed contents (631, 633) among the original contents (611, 613, 615, 617, 619, 621, 623) are played during a time interval (e.g., between time points (601, 603) and between time points (605, 607)) in which the driving situation is predicted (or identified) to be the first situation. For example, the electronic device (101) may shuffle the playlist of contents (611, 613, 615, 617, 619, 621, 623) so that the video contents (611, 615) among the original contents (611, 613, 615, 617, 619, 621, 623) are played first during a time interval (e.g., between time points (601, 603) and between time points (605, 607)) in which the driving situation is predicted (or identified) to be a second situation.
[0129] In one embodiment, the electronic device (101) may first assign video contents (611, 615, 621, 623) among the contents (611, 613, 615, 617, 619, 621, 623) included in the playlist to be played in the second situation. In one embodiment, if the electronic device (101) cannot assign some video contents (621, 623) among the video contents (611, 615, 621, 623) to be played in the second situation, the electronic device (101) may not assign some video contents (621, 623) to be played in the first situation. In one embodiment, the electronic device (101) can shuffle the playlist of contents (611, 613, 615, 617, 619, 621, 623) based on the inclusion order of the contents (611, 613, 615, 617, 619, 621, 623) included in the playlist. In one embodiment, the electronic device (101) can shuffle the playlist of contents (611, 613, 615, 617, 619, 621, 623) such that the playback time of the contents to be played in the first situation and / or the second situation does not exceed the duration of the corresponding situation. In one embodiment, if the playback time of the content to be played in the second situation exceeds the duration of the second situation, the electronic device (101) may convert the portion of the content to be played that exceeds the duration of the second situation into processed content and then output it.
[0130] For example, the electronic device (101) can shuffle the playlist so that the video content (611) included first among the contents (611, 613, 615, 617, 619, 621, 623) included in the playlist is played in the second situation between the times (603, 605). For example, the electronic device (101) can shuffle the playlist so that the audio content (613) included next among the contents (611, 613, 615, 617, 619, 621, 623) included in the playlist is played in the first situation between the times (601, 603). For example, the electronic device (101) can shuffle the playlist so that the video content (613) included next among the contents (611, 613, 615, 617, 619, 621, 623) included in the playlist is played in a second situation between the times (607, 609). For example, the electronic device (101) can shuffle the playlist so that the audio content (617) included next among the contents (611, 613, 615, 617, 619, 621, 623) included in the playlist is played in a first situation between the times (601, 603).
[0131] In one embodiment, the electronic device (101) can transmit the identified playable content from the shuffled playlist to an external electronic device (210) of the automobile (200).
[0132] For example, during the time interval between points (601, 603), the electronic device (101) can transmit the contents (613, 617, 631) to the external electronic device (210) so that the contents (613, 617, 631) are played back through the external electronic device (210). For example, the external electronic device (210) can display the processed contents (631) generated based on the video contents (621) during the time interval between points (601, 603).
[0133] For example, during the time interval between points (603, 605), the electronic device (101) can transmit the content (611) to the external electronic device (210) so that the content (611) is played through the external electronic device (210).
[0134] For example, during the time interval between points (605, 607), the electronic device (101) can transmit the contents (619, 633) to the external electronic device (210) so that the contents (619, 633) are played back through the external electronic device (210). For example, the external electronic device (210) can display the processed contents (633) generated based on the video contents (623) during the time interval between points (605, 607).
[0135] For example, during the time interval between points (607, 609), the electronic device (101) can transmit the content (615) to the external electronic device (210) so that the content (615) is played through the external electronic device (210).
[0136] In FIG. 6, a single video content is shown being output in a second situation, but this is merely an example. For example, the electronic device (101) can shuffle the playlist so that one or more video contents are output in a second situation.
[0137] FIG. 7 illustrates an example of processed content output from a vehicle according to one embodiment.
[0138] FIG. 7 may be described with reference to the electronic device (101) of FIG. 2 (or components of the electronic device (101)). The operations described with reference to FIG. 7 may be performed as instructions included in the path analysis unit (201), content analysis unit (203), prompt generation model (205), and / or GEN AI model (207) of FIG. 2 are executed by the processor (120) of the electronic device (101).
[0139] Referring to FIG. 7, a vehicle (700) may include a windshield (701). In one embodiment, the vehicle (700) may include displays (710, 720) on the windshield (701). In one embodiment, the vehicle (700) may include at least one display (730) in a center console portion. In one embodiment, the displays (710, 720, 730) may correspond to the display (265) of the external electronic device (210) of the vehicle (200) of FIG. 2. In one embodiment, the vehicle (700) may include at least one speaker (not shown). In one embodiment, the at least one speaker (not shown) of the vehicle (700) may correspond to the speaker (255) of the external electronic device (210) of the vehicle (200) of FIG. 2.
[0140] In one embodiment, the electronic device (101) can transmit content through a communication connection with an external electronic device (210) included in the automobile (700). The electronic device (101) can transmit content to the external electronic device (210) so that the external electronic device (210) outputs the content using at least one of the displays (710, 720, 730).
[0141] For example, the external electronic device (210) can output content through at least one of the displays (710, 720, 730). For example, the external electronic device (210) can output content through at least one of the displays (710, 720, 730) while outputting audio content (740).
[0142] For example, referring to FIG. 7, the external electronic device (210) can output content (711, 721) through displays (710, 720) of the windshield (701). In one embodiment, the displays (710, 720) of the windshield (701) can be a head-up display (HUD) that projects content (711, 721) onto the windshield (701) and provides the content (711, 721) reflected through the windshield (701) to the user. However, the present invention is not limited thereto. In one embodiment, the displays (710, 720) can be a head-up display (HUD) that is positioned in a space between the user and the windshield (701) and directly provides the content (711, 721) to the user. According to an embodiment, the display (710) may be a reflective area that reflects light from the head-up display. For example, the reflective area may reflect light from a head-up display at a location other than the reflective area toward the user. According to an embodiment, at least one portion of the windshield (701) may include a (transparent) display. For example, the displays (710, 720) of the windshield (701) may be transparent displays.
[0143] In one embodiment, the electronic device (101) may transmit content to the external electronic device (210) so that a visual object of a specified type among the displays (710, 720) is output from a display (710) that is closer to the driver. For example, the visual object of the specified type may be a processed image content generated from the original content. In one embodiment, the electronic device (101) may transmit content to the external electronic device (210) so that a visual object of another specified type among the displays (710, 720) is output from a display (720) that is further away from the driver. For example, the visual object of another specified type may be an image content included in the original content.
[0144] In one embodiment, the electronic device (101) can identify content that cannot be shown to the user depending on the driving conditions. For example, while displaying content on the display (720), the electronic device (101) can identify content that cannot be shown to the user depending on the driving conditions. For example, when the electronic device (101) can display content on the display (720), the electronic device (101) can identify content that cannot be shown to the user depending on the driving conditions. For example, while displaying content (721) to the user through the display (720), the electronic device (101) can determine whether there is content that cannot be shown to the user depending on the driving conditions.
[0145] In one embodiment, the electronic device (101) may display contents that cannot be shown to the user depending on the driving situation. For example, referring to FIG. 4, in a situation where one or more sets of processed images (471, 473) are not displayed, and a set of processed images (475) is displayed through the display (720), the electronic device (101) may display one or more sets of processed images (471, 473) through the display (720) instead of the set of processed images (475). For example, referring to FIG. 4, in a situation where one or more sets of processed images (471, 473) are not displayed, in a situation where a set of processed images (475) is displayed through a display (720), the electronic device (101) can display one or more sets of processed images (471, 473) together with the set of processed images (475) through the display (720).
[0146] According to an embodiment, the electronic device (101) may display content (721) when the driving situation of the vehicle (200) satisfies a preset condition (e.g., a first condition or a second condition). According to an embodiment, the electronic device (101) may display content (721) at a preset location (e.g., a display (710) or a display (720)) or a location based on the driving situation when the driving situation of the vehicle (200) satisfies the preset condition. For example, the location based on the driving situation may be the location of the display (720) when the first condition is met. For example, the location based on the driving situation may be the location of the display (710) when the second condition is met.
[0147] According to an embodiment, the electronic device (101) may determine the location of a location to display content (e.g., a display (710, 720)) based on the passenger status of the vehicle (200) (e.g., passengers and / or the number of passengers). For example, if passengers are also in the first row passenger seat, the electronic device (101) may determine the location to display content as the display (720). For example, if passengers are only in the first row driver's seat, the electronic device (101) may determine the location to display content as the display (710).
[0148] FIG. 8A is a block diagram of an electronic device and an external electronic device according to one embodiment.
[0149] FIG. 8A may illustrate a situation in which at least some of the functions of the electronic device (101) are included in an external electronic device (210) of the automobile (200), as compared to FIG. 2. The functions of the path analysis unit (801), the content analysis unit (803), the prompt generation model (805), and / or the GEN AI model (807) of FIG. 8A may correspond to the functions of the path analysis unit (201), the content analysis unit (203), the prompt generation model (205), and / or the GEN AI model (207) described in FIG. 2. However, the present invention is not limited thereto. For example, the GEN AI model (807) may not be included in the external electronic device (210). For example, the GEN AI model (807) may be included in a device external to the external electronic device (210) (e.g., the server (108) of FIG. 1).
[0150] In one embodiment, referring to FIG. 8A, the electronic device (101) may transmit original content (810) output through the external electronic device (210) to the external electronic device (210). In one embodiment, the electronic device (101) may transmit the original content (810) to the external electronic device (210) as content to be played according to a playlist.
[0151] In one embodiment, the external electronic device (210) can identify a driving path through the path analysis unit (801). For example, the driving path can indicate a road on which the vehicle (200) is currently located and / or a road on which the vehicle (200) will drive in the future.
[0152] In one embodiment, the external electronic device (210) may identify, through the content analysis unit (803), the purpose, type, content, presence or absence of video (or image), and / or playback time of content to be output (or played) through the external electronic device (210) of the automobile (200). In one embodiment, the external electronic device (210) may extract (or analyze) the content of the identified content through the content analysis unit (803). In one embodiment, the external electronic device (210) may generate, through the content analysis unit (803), at least one sentence (or at least one word) indicating the purpose, type, content, presence or absence of video (or image), and / or playback time of the identified content.
[0153] In one embodiment, the external electronic device (210) can obtain (or generate) a prompt through a prompt generation model (805). In one embodiment, the external electronic device (210) can generate a prompt for obtaining (or generating) content to be output (or played) through the external electronic device (210) through the prompt generation model (805). In one embodiment, the external electronic device (210) can generate processed content based on the prompt through a GEN AI model (807).
[0154] In one embodiment, the processor (225) of the external electronic device (210) can output the generated processed content. The external electronic device (210) can output the generated processed content through a display (265) and / or a speaker (255).
[0155] In one embodiment, the processor (225) may display one of the processed image content or the video content included in the original content through the display (265) included in the vehicle (200). For example, the processor (225) may display the processed image content based on determining that a specified first condition is satisfied. For example, the processor (225) may display the video content based on determining that a specified second condition, distinct from the specified first condition, is satisfied. For example, the processor (225) may shuffle the playlist so that the processed content is output when the vehicle (200) is determined to be located in the first section. For example, the processor (225) may shuffle the playlist so that the original content is output when the vehicle (200) is determined to be located in the second section.
[0156] FIG. 8b is a block diagram of an electronic device and an external electronic device according to one embodiment.
[0157] FIG. 8B may illustrate a situation in which at least some of the functions of the electronic device (101) are included in an external electronic device (210) of the automobile (200), as compared to FIG. 2. The functions of the path analysis unit (801), the content analysis unit (803), the prompt generation model (805), and / or the GEN AI model (807) of FIG. 8B may correspond to the functions of the path analysis unit (201), the content analysis unit (203), the prompt generation model (205), and / or the GEN AI model (207) described in FIG. 2.
[0158] In one embodiment, referring to FIG. 8B, the electronic device (101) may transmit processed content (820) generated based on original content output through the external electronic device (210) to the external electronic device (210). In one embodiment, the electronic device (101) may transmit the processed content (820) to the external electronic device (210) as content to be played according to a playlist.
[0159] In one embodiment, the external electronic device (210) may detect driving conditions, such as traffic conditions, weather conditions, and / or lighting, through one or more sensors (e.g., a camera) of the automobile (200). In one embodiment, the external electronic device (210) may determine whether to turn the autonomous driving mode on or off based on the driving conditions.
[0160] In one embodiment, the external electronic device (210) may output processed content (820) based on a determination that the autonomous driving mode is turned on. In one embodiment, the external electronic device (210) may reprocess the processed content (820) based on a determination that the autonomous driving mode is turned off. For example, the reprocessing of the processed content (820) may include changing the content so as not to distract the driver from driving. For example, the external electronic device (210) may reprocess the processed content (820) using the path analysis unit (801), the content analysis unit (803), the prompt generation model (805), and / or the GEN AI model (807). However, the present invention is not limited thereto. In one embodiment, the external electronic device (210) may request original content from the electronic device (101) in addition to the processed content (820) based on a determination that the autonomous driving mode is turned on. For example, an external electronic device (210) can output original content obtained from an electronic device (101).
[0161] FIG. 9A is a flowchart illustrating the operation of an electronic device according to one embodiment.
[0162] FIG. 9A may be described with reference to the electronic device (101) of FIG. 2 (or components of the electronic device (101)). The operations described with reference to FIG. 9A may be performed as instructions included in the path analysis unit (201), content analysis unit (203), prompt generation model (205), and / or GEN AI model (207) of FIG. 2 are executed by the processor (120) of the electronic device (101).
[0163] Referring to FIG. 9A, in operation 910, the electronic device (101) may identify content. For example, the electronic device (101) may identify content to be output (or played) through an external electronic device (210) of the automobile (200). In one embodiment, the electronic device (101) may identify a list of content to be output (or played) through the external electronic device (210) of the automobile (200). In one embodiment, the content may be video content and / or audio content. For example, the video content may include video data and audio data. For example, the audio content may include audio data.
[0164] In operation 920, the electronic device (101) may obtain audio data. For example, the electronic device (101) may obtain audio data by inputting a prompt including one or more words indicating the purpose, type, content, presence or absence of video (or image), and / or playback time of the content into the GEN AI model (207). For example, the obtained audio data may be at least partially different from the audio data included in the original content.
[0165] In operation 930, the electronic device (101) may generate visual content. In one embodiment, the electronic device (101) may identify one or more voices (or voice utterances) representing a topic (or material) (or keyword) of the original content among a plurality of voices (or voice utterances) included in audio data. In one embodiment, the electronic device (101) may generate a prompt based on the plurality of voices (or voice utterances) included in the original content (e.g., video content or audio content). For example, the prompt based on the plurality of voices (or voice utterances) may include one or more words representing a topic (or material) (or keyword) of the original content (e.g., video content or audio content). For example, a prompt based on multiple voices (or voice utterances) may include one or more words to instruct the GEN AI model (207) to generate visual content that describes the content of the voice utterances.
[0166] In one embodiment, the electronic device (101) can generate visual content by inputting a prompt to the GEN AI model (207). For example, the visual content can include a still image (or multiple still images) representing (or depicting) a utterance included in the original content (or audio data).
[0167] In operation 940, the electronic device (101) may output visual content together with audio data. For example, the processor (120) may transmit the visual content together with the audio data generated through the communication circuit (290) to an external electronic device (210) of the automobile (200). The external electronic device (210) may output the visual content together with the received audio data through the display (265) and / or the speaker (255).
[0168] FIG. 9b is a flowchart illustrating the operation of an electronic device according to one embodiment.
[0169] FIG. 9B may be described with reference to the electronic device (101) of FIG. 2 (or components of the electronic device (101)). The operations described with reference to FIG. 9B may be performed as instructions included in the path analysis unit (201), content analysis unit (203), prompt generation model (205), and / or GEN AI model (207) of FIG. 2 are executed by the processor (120) of the electronic device (101).
[0170] Among the operations of FIG. 9b, operations 910, 920, 930, and 940 may correspond to operations 910, 920, 930, and 940 of FIG. 9a.
[0171] Referring to FIG. 9b, at operation 950, the electronic device (101) can identify a destination path.
[0172] In one embodiment, the electronic device (101) may obtain an input indicating a destination. In one embodiment, the electronic device (101) may identify a destination path from the current location of the vehicle (200) to the destination based on obtaining the input indicating the destination.
[0173] In one embodiment, the electronic device (101) may classify a destination route into one or more segments. The one or more segments may include a first segment where the degree of driving intervention required by the vehicle (200) from the user satisfies a first criterion. The one or more segments may include a second segment where the degree of driving intervention required by the vehicle (200) from the user satisfies a second criterion that is lower than the first criterion.
[0174] In one embodiment, the electronic device (101) may identify an estimated time required to pass through each of one or more segments. For example, the electronic device (101) may identify an estimated time required to pass through each of one or more segments based on the driving environment and / or moving speed.
[0175] In operation 910, the electronic device (101) can identify content.
[0176] In operation 920, the electronic device (101) may obtain audio data. For example, the electronic device (101) may obtain audio data by inputting a prompt including one or more words indicating the purpose, type, content, presence or absence of video (or image), and / or playback time of the content into the GEN AI model (207). For example, the obtained audio data may be at least partially different from the audio data included in the original content.
[0177] In one embodiment, the electronic device (101) may generate a prompt to obtain (or generate) processed content to be output (or played) through the external electronic device (210) based on information related to the destination path identified in operation 950. For example, the information related to the destination path identified in operation 950 may include autonomous driving performance (e.g., autonomous driving level) of the vehicle (200) and / or whether autonomous driving is on or off. For example, the information related to the destination path identified in operation 950 may include a current location, a moving direction, a moving speed, a current driving time, and / or an expected time required to travel to the destination of the vehicle (200) (or the external electronic device (210)) and / or the electronic device (101). For example, the information related to the destination path identified in operation 950 may include information about a driving environment along the path to the destination.
[0178] For example, the electronic device (101) may obtain audio data by inputting a prompt including one or more words to the GEN AI model (207) to obtain (or generate) processed content to be output (or played) through the external electronic device (210) based on one or more words indicating the purpose, type, content, presence or absence of video (or image), and / or playing time of the content and information related to the destination path identified in operation 950. For example, the obtained audio data may be at least partially different from audio data included in the original content.
[0179] At operation 930, the electronic device (101) can generate visual content. At operation 940, the electronic device (101) can output the visual content together with audio data.
[0180] FIG. 10 is a diagram illustrating the structure of a generative AI model of an electronic device according to one embodiment.
[0181] Figure 10 can be explained with reference to Figures 1 and 2.
[0182] In one embodiment, a user interface (UI) (1010) may be an element for interaction between an electronic device (101) and a user. For example, the UI (1010) may be an element for receiving (or acquiring) a user's input. For example, the UI (1010) may be an element for providing (or outputting) a result of the user's input.
[0183] In one embodiment, the UI (1010) may be a graphical UI (GUI) for user interaction (e.g., acquisition of touch input and / or display of results) via the display (260). In one embodiment, the UI (1010) may be a voice UI (VUI) (or an auditory UI (AUI)) for user interaction (e.g., acquisition of voice signals and / or output of audio signals) via the audio output module (155) (or the audio module (170)). In one embodiment, the UI (1010) may be a natural UI (NUI) for user interaction (e.g., acquisition of user gestures or gaze) via the camera module (180). In one embodiment, the UI (1010) may be a physical UI (PUI) (or a tangible UI (TUI)) for user interaction (e.g., input to a physical button) via the input module (150).
[0184] In one embodiment, the UI (1010) may be configured for interaction between the electronic device (101) and the user for the purpose of querying the user and / or responding to the query. In one embodiment, the UI (1010) may receive user input. For example, the user input may include natural language input (e.g., voice signal input, and / or text input), and / or input for selecting (or indicating) content (e.g., images and / or videos). The user input may also be in a mixed form of the above-described natural language, images, sounds, and context information. The user input may also be in a non-natural language form, such as selecting a menu.
[0185] In one embodiment, the UI (1010) may identify (or obtain) information about the context related to the time (or situation) at which the user's input is obtained, for the purpose of the user's inquiry and / or response to the inquiry. In one embodiment, the information about the context may include additional information at the time of the user input. For example, the additional information may include information about the application currently being used by the user or location information of the user (or the electronic device (101)).
[0186] In one embodiment, the electronic device (101) may output the results of a generative AI model (1050) to the user via a UI (1010). In one embodiment, the output may be in the form of natural language or specific content. In one embodiment, the output may be provided in a form requested by the user (e.g., an action).
[0187] In one embodiment, the UI (1010) may transmit information about the user's input and / or context to an artificial intelligence (AI) framework (1020). In one embodiment, the AI framework (1020) may obtain information about the user's input and / or context from the UI (1010).
[0188] In one embodiment, the UI (1010) may include one or more analyzers (1011, 1013, 1015). In one embodiment, the content analyzer (1011) may correspond to the content analysis unit (203) of FIG. 2. In one embodiment, the user analyzer (1013) may be included as part of the path analysis unit (201) and / or the content analysis unit (203) of FIG. 2. In one embodiment, the environment analyzer (1013) may be included as part of the path analysis unit (201) of FIG. 2.
[0189] In one embodiment, the AI framework (1020) may coordinate and / or control each of the components (e.g., interface management element (1021), prompt design element (1023), or output modification element (1025)) necessary to perform a task according to the user's intent identified based on the user's query (or query included in the user input).
[0190] In one embodiment, the AI framework (1020) may transmit user input obtained from the UI (1010) to a prompt design element (1023).
[0191] In one embodiment, the prompt design element (1023) may generate a prompt suitable for inputting user input into a generative AI model (1050) (e.g., a large language model (LLM) and / or a larger multimodal model (LMM)).
[0192] In one embodiment, the prompt design element (1023) may be an AI component that utilizes a machine learning algorithm or neural network. Accordingly, the prompt design element (1023) may change the prompt generated through learning.
[0193] In one embodiment, the prompt design element (1023) may access a knowledge element (or knowledge repository (1030)) containing user preference data, a prompt library, and prompt examples based on user input to generate a prompt, and pass the generated prompt to a generative AI model (1050) (e.g., an LLM and / or LMM).
[0194] In one embodiment, the prompt design element (1023) may correspond to the prompt generation model (205) of FIG. 2.
[0195] In one embodiment, the interface management element (1021) can communicate with external elements. In one embodiment, the interface management element (1021) can obtain additional information from external elements when there is a request for additional information when passing user input (or a prompt based on user input) as input to the generative AI model (1050). In one embodiment, the interface management element (1021) can establish a channel for communicating with the outside of the AI framework (1020) through an application programming interface (API), and can enable the AI framework (1020) to access various data sources (e.g., a knowledge repository (1030)) through the established channel. In addition, the interface management element (1021) can request the service element (1040) through the API to perform an action that performs the user input as a final result rather than an intermediate result when the application and / or service needs to perform the action. Information obtained from external elements may be used to generate prompts in the prompt design element (1023) along with user input or may be passed as input to the generative AI model (1050).
[0196] In one embodiment, the output modification component (1025) can fine-tune the output from the generative AI model (1050). For example, the output modification component (1025) can verify (e.g., verify for relevance, bias, and / or harmfulness) the content generated by the generative AI model (1050) (e.g., LLM and / or LMM). In addition, the output modification component (1025) can determine to what extent the output result matches the result desired by the user and, if additional adjustment is needed, can perform additional tasks for additional adjustment. Additionally, the output modification component (1025) can provide the user with hints to prevent (or reduce) undesired output.
[0197] In one embodiment, a generative AI model (1050) may generally refer to an artificial intelligence neural network that generates new types of data based on user input information. The generative AI model (1050) may include a model that generates images and / or a model that generates language. The model that generates images may include a generative adversarial network (GAN), a variational autoencoder (VAE), and / or a diffusion-based generative model that uses a VAE and a transformer architecture. The model that generates language may be a model trained to output statistically most appropriate output values based on input values. For example, the model that generates language may include a model such as CHAT-GPT 3 or CHAT-GPT 4. In addition, the generative AI model (1050) may be an LMM that can recognize various types of data input, such as text, images, and voice, and generate new data corresponding thereto.
[0198] In one embodiment, the generative AI model (1050) may correspond to the GEN AI model (207) of FIG. 2.
[0199] As described above, the electronic device (101) may include a communication circuit (290), at least one processor (120) including a processing circuit; and a memory (130) storing instructions and including one or more storage media. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify original content (410) for generating content to be played while a user of the electronic device (101) is in a vehicle (200). The original content (410) may include video content (420) and / or audio content (430). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain audio data (460) from the original content (410). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify a voice utterance representing a subject of the audio data (460) from among a plurality of voice utterances included in the audio data (460). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate image content (470) describing the content of the voice utterance by inputting a prompt based on the voice utterance into a generative AI model (207).The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to transmit the image content (470) to the external electronic device (210) of the automobile (200) via the communication circuit (290) such that the image content (470) is displayed via the display (265) included in the external electronic device (210) of the automobile (200) during at least a portion of the time period during which the voice utterance is played via the speaker (255) included in the external electronic device (210) of the automobile (200).
[0200] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify a driving path of the vehicle (200) based on location information. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to transmit the image content (470) to the external electronic device (210) via the communication circuit (290) such that the image content (470) is displayed via the display (265) when the electronic device (101) determines that the vehicle (200) is located in a first section where a degree of driving intervention required from the user satisfies a first criterion based on the driving path. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to transmit the image content (420) to the external electronic device (210) through the communication circuit (290) such that the image content (420) is displayed through the display (265), instead of the image content (470), when the electronic device (101) determines that the vehicle (200) is located in a second section of the driving route that satisfies a second criterion in which the degree of driving intervention is lower than the first criterion.
[0201] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain an input indicating a destination. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain the driving route from a current location to the destination based on obtaining the input. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to determine, based on the location information, whether the automobile (200) is located in the first section or the second section among the sections on the driving route.
[0202] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify a driving state of the vehicle (200). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to transmit the image content (470) to the external electronic device (210) via the communication circuit (290) so that the image content (470) is displayed via the display (265) when the vehicle (200) is determined to be driving based on the driving state. The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to transmit the image content (420) to the external electronic device (210) via the communication circuit (290) so that the image content (420) is displayed via the display (265), instead of the image content (470), when the vehicle (200) is determined to be stopped based on the driving state.
[0203] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify data indicating whether the automobile (200) is autonomously driven. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to transmit the image content (470) to the external electronic device (210) via the communication circuit (290) so that the image content (470) is displayed via the display (265) when the degree of driving intervention required from the user is determined to meet a first criterion based on the data indicating whether the automobile is autonomously driven. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to transmit the image content (420) to the external electronic device (210) through the communication circuit (290) so that the image content (420) is displayed through the display (265) instead of the image content (470) when the degree of driving intervention required from the user is determined to be lower than the first criterion based on the data indicating whether the autonomous driving is performed.
[0204] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify data indicative of a remaining driving time. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain the audio data (460) having a shorter playback time than the playback time of the original content (410), based on the data indicative of the remaining driving time, when the playback time of the original content (410) is determined to be longer than the remaining driving time.
[0205] The playback time of the above audio data (460) can correspond to the remaining driving time.
[0206] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify a playlist of a plurality of contents. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify a driving route of the vehicle (200) based on location information. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to classify a plurality of sections included in the driving route into a first section or a second section based on the driving route. The first section may be a section in which the degree of driving intervention required from the user of the vehicle (200) satisfies a first criterion. The second section may be a section in which the vehicle (200) satisfies a second criterion in which the degree of driving intervention is lower than the first criterion. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to shuffle the playlist so as to play audio content (430) among the plurality of contents in the first section and video content (420) among the plurality of contents in the second section.
[0207] As described above, the electronic device (210) included in the automobile (200) may include at least one processor (225) including a speaker (255), a display (265), a communication circuit (295), a processing circuit; and a memory (235) storing instructions and including one or more storage media. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to obtain, from an external electronic device (101) via the communication circuit (295), information indicating original content (410) to be played while the user of the electronic device (210) is in the automobile (200). The original content (410) may include video content (420) and / or audio content (430). The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to obtain audio data (460) from the original content (410) based on the information. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to identify a voice utterance representing a subject of the audio data (460) from among a plurality of voice utterances included in the audio data (460). The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to generate image content (470) describing the content of the voice utterance by inputting a prompt based on the voice utterance into a generative AI model (807).The above instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to display the image content (470) through the display (265) during at least a portion of the time period during which the voice utterance is reproduced through the speaker (255).
[0208] The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to identify a driving path of the vehicle (200) based on location information. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to display the image content (470) through the display (265) when it is determined, based on the driving path, that the vehicle (200) is located in a first section where a degree of driving intervention required from the user satisfies a first criterion. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to display the video content (420) through the display (265) instead of the image content (470) when the vehicle (200) is determined to be located in a second section of the driving route that satisfies a second criterion in which the degree of driving intervention is lower than the first criterion.
[0209] The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to obtain an input indicating a destination. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to obtain a driving route from a current location to the destination based on obtaining the input. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to determine, based on the location information, whether the automobile (200) is located in the first section or the second section among the sections on the driving route.
[0210] The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to identify a driving state of the vehicle (200). The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to display the image content (470) through the display (265) when the vehicle (200) is determined to be driving based on the driving state. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to display the video content (420) through the display (265) instead of the image content (470) when the vehicle (200) is determined to be stopped based on the driving state.
[0211] The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to identify data indicating whether the vehicle (200) is autonomously driven. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to display the image content (470) through the display (265) when the degree of driving intervention required from the user is determined to meet a first criterion based on the data indicating whether the vehicle is autonomously driven. The above instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to display the video content (420) through the display (265) instead of the image content (470) when the degree of driving intervention required from the user is determined to be lower than the first criterion based on the data indicating whether the autonomous driving is performed.
[0212] The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to identify data indicative of a remaining driving time. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to obtain the audio data (460) having a playback time shorter than the playback time of the original content (410), based on the data indicative of the remaining driving time, when the playback time of the original content (410) is determined to be longer than the remaining driving time.
[0213] The playback time of the above audio data (460) can correspond to the remaining driving time.
[0214] The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to identify a playlist of a plurality of contents. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to identify a driving route of the vehicle (200) based on location information. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to classify a plurality of sections included in the driving route into a first section and a second section based on the driving route. The first section may be a section in which the degree of driving intervention required from the user of the vehicle (200) satisfies a first criterion. The second section may be a section in which the vehicle (200) satisfies a second criterion in which the degree of driving intervention is lower than the first criterion. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to shuffle the playlist, such that audio content (430) among the plurality of contents is played in the first section, and video content (420) among the plurality of contents is played in the second section.
[0215] As described above, the method may be performed in an electronic device (101) including a communication circuit (290). The method may include an operation of identifying original content (410) for generating content to be played while a user of the electronic device (101) is in a vehicle (200). The original content (410) may include video content (420) and / or audio content (430). The method may include an operation of obtaining audio data (460) from the original content (410). The method may include an operation of identifying a voice utterance representing a subject of the audio data (460) among a plurality of voice utterances included in the audio data (460). The method may include an operation of generating image content (470) depicting the content of the voice utterance by inputting a prompt based on the voice utterance into a generative AI model (207). The method may include an operation of transmitting the image content (470) to the external electronic device (210) of the automobile (200) via the communication circuit (290) such that the image content (470) is displayed via the display (265) included in the external electronic device (210) of the automobile (200) during at least a portion of the time period during which the voice utterance is played via the speaker (255) included in the external electronic device (210) of the automobile (200).
[0216] The method may include an operation of identifying a driving path of the vehicle (200) based on location information. The method may include an operation of transmitting the image content (470) to the external electronic device (210) through the communication circuit (290) so that the image content (470) is displayed through the display (265) when it is determined that the vehicle (200) is located in a first section where a degree of driving intervention required from the user satisfies a first criterion based on the driving path. The method may include an operation of transmitting the image content (420) to the external electronic device (210) through the communication circuit (290) so that the image content (420) is displayed through the display (265) instead of the image content (470), when it is determined that the vehicle (200) is located in a second section that satisfies a second criterion in which the degree of driving intervention is lower than the first criterion based on the driving path.
[0217] The method may include an operation of identifying a driving state of the vehicle (200). The method may include an operation of transmitting the image content (470) to the external electronic device (210) through the communication circuit (290) so that the image content (470) is displayed through the display (265) when the vehicle (200) is determined to be driving based on the driving state. The method may include an operation of transmitting the image content (420) to the external electronic device (210) through the communication circuit (290) so that the image content (420) is displayed through the display (265) instead of the image content (470) when the vehicle (200) is determined to be stopped based on the driving state.
[0218] The method may include an operation of identifying data indicating whether the automobile (200) is autonomously driven. The method may include an operation of transmitting the image content (470) to the external electronic device (210) through the communication circuit (290) so that the image content (470) is displayed on the display (265) when the degree of driving intervention required from the user is determined to satisfy a first criterion based on the data indicating whether the automobile is autonomously driven. The method may include an operation of transmitting the image content (420) to the external electronic device (210) through the communication circuit (290) so that the image content (420) is displayed on the display (265) instead of the image content (470) when the degree of driving intervention required from the user is determined to satisfy a second criterion that is lower than the first criterion based on the data indicating whether the automobile is autonomously driven.
[0219] As described above, the method may be performed in an electronic device (210) included in a vehicle (200) including a speaker (255), a display (265), and a communication circuit (295). The method may include an operation of obtaining, from an external electronic device (101) via the communication circuit (295), information indicating original content (410) to be played while a user of the electronic device (210) is in the vehicle (200). The original content (410) may include video content (420) and / or audio content (430). The method may include an operation of obtaining audio data (460) from the original content (410) based on the information. The method may include an operation of identifying a voice utterance indicating a subject of the audio data (460) from among a plurality of voice utterances included in the audio data (460). The method may include an operation of generating image content (470) depicting the content of the voice utterance by inputting a prompt based on the voice utterance into a generative AI model (807). The method may include an operation of displaying the image content (470) through the display (265) during at least a portion of a time period during which the voice utterance is played through the speaker (255).
[0220] As described above, a non-transitory computer readable storage medium can store a program including instructions. The instructions, when individually or collectively executed by at least one processor (120) of an electronic device (101) including a communication circuit (290), can cause the electronic device (101) to identify original content (410) for generating content to be played while a user of the electronic device (101) is in a vehicle (200). The original content (410) can include video content (420) and / or audio content (430). The instructions, when individually or collectively executed by the at least one processor (120), can cause the electronic device (101) to obtain audio data (460) from the original content (410). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify a voice utterance representing a subject of the audio data (460) from among a plurality of voice utterances included in the audio data (460). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate image content (470) describing the content of the voice utterance by inputting a prompt based on the voice utterance into a generative AI model (207).The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to transmit the image content (470) to the external electronic device (210) of the automobile (200) via the communication circuit (290) such that the image content (470) is displayed via the display (265) included in the external electronic device (210) of the automobile (200) during at least a portion of the time period during which the voice utterance is played via the speaker (255) included in the external electronic device (210) of the automobile (200).
[0221] As described above, a non-transitory computer readable storage medium can store a program including instructions. The instructions, when individually or collectively executed by at least one processor (225) of an electronic device (210) included in a vehicle (200) including a speaker (255), a display (265), and a communication circuit (295), can cause the electronic device (210) to obtain, from an external electronic device (101) via the communication circuit (295), information representing original content (410) to be played while a user of the electronic device (210) is in the vehicle (200). The original content (410) can include video content (420) and / or audio content (430). The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to obtain audio data (460) from the original content (410) based on the information. The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to identify a voice utterance representing a subject of the audio data (460) from among a plurality of voice utterances included in the audio data (460). The instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to generate image content (470) describing the content of the voice utterance by inputting a prompt based on the voice utterance into a generative AI model (807).The above instructions, when individually or collectively executed by the at least one processor (225), may cause the electronic device (210) to display the image content (470) through the display (265) during at least a portion of the time period during which the voice utterance is reproduced through the speaker (255).
[0222] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0223] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0224] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0225] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0226] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0227] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device (101), Communication circuit (290), At least one processor (120) comprising a processing circuit; and A memory (130) storing instructions and including one or more storage media, wherein the instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: The user of the electronic device (101) identifies original content (410) for generating content to be played while riding in a car (200), and the original content (410) includes video content (420) and / or audio content (430). Obtain audio data (460) from the above original content (410), Identifying a voice utterance representing the subject of the audio data (460) among a plurality of voice utterances included in the audio data (460), By inputting a prompt based on the above voice utterance into a generative AI model (207), image content (470) describing the content of the above voice utterance is generated, Causing the image content (470) to be transmitted to the external electronic device (210) of the automobile (200) through the communication circuit (290) so that the image content (470) is displayed through the display (265) included in the external electronic device (210) of the automobile (200). Electronic devices.
2. In claim 1, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Identifying the driving path of the vehicle (200) based on location information, Based on the driving route, when it is determined that the vehicle (200) is located in a first section where the degree of driving intervention required from the user satisfies a first criterion, the image content (470) is transmitted to the external electronic device (210) through the communication circuit (290) so that the image content (470) is displayed through the display (265), and Based on the driving path, when it is determined that the vehicle (200) is located in a second section that satisfies a second criterion in which the degree of driving intervention is lower than the first criterion, the image content (420) is transmitted to the external electronic device (210) through the communication circuit (290) so that the image content (420) is displayed through the display (265) instead of the image content (470). Electronic devices.
3. In claim 2, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Obtain input indicating the destination, Based on obtaining the above input, obtaining the driving route from the current location to the destination, Based on the above location information, causing the vehicle (200) to determine whether it is located in the first section or the second section among the sections on the driving route. Electronic devices.
4. In any one of claims 1 to 3, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Identify the driving status of the above vehicle (200), Based on the driving status, when it is determined that the vehicle (200) is driving, the image content (470) is transmitted to the external electronic device (210) through the communication circuit (290) so that the image content (470) is displayed through the display (265), and Based on the driving state, when it is determined that the vehicle (200) is stopped, the image content (420) is transmitted to the external electronic device (210) through the communication circuit (290) so that the image content (420) is displayed through the display (265) instead of the image content (470). Electronic devices.
5. In any one of claims 1 to 4, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Identify data indicating whether the above vehicle (200) is autonomously driven, Based on the data indicating whether the autonomous driving is performed, when the degree of driving intervention required from the user is determined to meet the first criterion, the image content (470) is transmitted to the external electronic device (210) through the communication circuit (290) so that the image content (470) is displayed through the display (265), and Based on the data indicating whether the autonomous driving is performed, when it is determined that the degree of driving intervention required from the user satisfies a second criterion lower than the first criterion, the image content (420) is transmitted to the external electronic device (210) through the communication circuit (290) so that the image content (420) is displayed through the display (265) instead of the image content (470). Electronic devices.
6. In any one of claims 1 to 5, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Identify data indicating remaining driving time, Based on the data indicating the remaining driving time, the playback time of the original content (410) is determined to be longer than the remaining driving time, thereby causing the audio data (460) having a playback time shorter than the playback time of the original content (410) to be acquired. Electronic devices.
7. In claim 6, The playback time of the above audio data (460) corresponds to the remaining driving time. Electronic devices.
8. In any one of claims 1 to 6, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Identify a playlist of multiple contents, Identifying the driving path of the vehicle (200) based on location information, Based on the driving route, a plurality of sections included in the driving route are classified into a first section or a second section, and the first section is a section in which the degree of driving intervention required of the vehicle (200) from the user satisfies a first standard, and the second section is a section in which the degree of driving intervention required of the vehicle (200) from the user satisfies a second standard lower than the first standard. Causing the playlist to be shuffled so that audio content (430) among the plurality of contents is played in the first section and video content (420) among the plurality of contents is played in the second section. Electronic devices.
9. In an electronic device (210) included in a vehicle (200), Speaker (255), Display (265), Communication circuit (295), At least one processor (225) comprising a processing circuit; and A memory (235) storing instructions and including one or more storage media, wherein the instructions, when individually or collectively executed by the at least one processor (225), cause the electronic device (210) to: Through the above communication circuit (295), information indicating original content (410) to be played while the user of the electronic device (210) is riding in the automobile (200) is obtained from an external electronic device (101), and the original content (410) includes video content (420) and / or audio content (430). Based on the above information, audio data (460) is obtained from the original content (410), Identifying a voice utterance representing the subject of the audio data (460) among a plurality of voice utterances included in the audio data (460), By inputting a prompt based on the above voice utterance into a generative AI model (807), image content (470) describing the content of the above voice utterance is generated, Causing the above image content (470) to be displayed through the display (265), Electronic devices.
10. In claim 9, The above instructions, when individually or collectively executed by the at least one processor (225), cause the electronic device (210) to: Identifying the driving path of the vehicle (200) based on location information, Based on the driving path, when it is determined that the vehicle (200) is located in a first section where the degree of driving intervention required from the user satisfies a first criterion, the image content (470) is displayed through the display (265), and Based on the driving path, when it is determined that the vehicle (200) is located in a second section that satisfies a second criterion in which the degree of driving intervention is lower than the first criterion, the video content (420) is caused to be displayed through the display (265) instead of the image content (470). Electronic devices.
11. In claim 10, The above instructions, when individually or collectively executed by the at least one processor (225), cause the electronic device (210) to: Obtain input indicating the destination, Based on obtaining the above input, obtaining a driving route from the current location to the destination, Based on the above location information, causing the vehicle (200) to determine whether it is located in the first section or the second section among the sections on the driving route. Electronic devices.
12. In any one of claims 9 to 11, The above instructions, when individually or collectively executed by the at least one processor (225), cause the electronic device (210) to: Identify the driving status of the above vehicle (200), Based on the above driving status, when it is determined that the vehicle (200) is driving, the image content (470) is displayed through the display (265), and Based on the driving state, when it is determined that the vehicle (200) is stopped, the video content (420) is displayed through the display (265) instead of the image content (470). Electronic devices.
13. In any one of claims 9 to 12, The above instructions, when individually or collectively executed by the at least one processor (225), cause the electronic device (210) to: Identify data indicating whether the above vehicle (200) is autonomously driven, Based on the data indicating whether the autonomous driving is performed, if the degree of driving intervention required from the user is determined to meet the first criterion, the image content (470) is displayed through the display (265), and Based on the data indicating whether the autonomous driving is performed, the degree of driving intervention required from the user is determined to meet a second criterion that is lower than the first criterion, thereby causing the video content (420) to be displayed through the display (265) instead of the image content (470). Electronic devices.
14. In claim 9, The above instructions, when individually or collectively executed by the at least one processor (225), cause the electronic device (210) to: Identify data indicating remaining driving time, Based on the data indicating the remaining driving time, the playback time of the original content (410) is determined to be longer than the remaining driving time, thereby causing the audio data (460) having a playback time shorter than the playback time of the original content (410) to be acquired. Electronic devices.
15. In claim 14, The playback time of the above audio data (460) corresponds to the remaining driving time. Electronic devices.
Citation Information
Patent Citations
Vehicle wallpaper generation method and device, electronic equipment and readable storage medium
CN116501432A
In-vehicle information interaction method, in-vehicle infotainment system and vehicle
CN117275488A
On-vehicle content reproduction device
JP2006133006A
Electronic device including indicator label
KR1020240014404A
Multimedia automatic generation system for automatically generating multimedia suitable for user's voice data by using artificial intelligence
KR102213618B1