Signal processing device and vehicle control device including same

WO2026169100A1PCT designated stage Publication Date: 2026-08-13LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-10
Publication Date
2026-08-13

Smart Images

  • Figure KR2026002421_13082026_PF_FP_ABST
    Figure KR2026002421_13082026_PF_FP_ABST
Patent Text Reader

Abstract

A signal processing device according to an embodiment of the present disclosure, and a vehicle control device including same, comprises: a memory; and a processor that receives outside vehicle camera data and inside vehicle camera data, wherein the processor classifies emotion data of a passenger on the basis of the inside vehicle camera data, generates a highlight video including an external video from the outside vehicle camera data in response to the classified emotion data, and controls the generated highlight video to be stored in the memory. Accordingly, it is possible to provide an emotion-based highlight video.
Need to check novelty before this filing date? Find Prior Art

Description

Signal processing device, and vehicle control device having the same

[0001] The present disclosure relates to a signal processing device and a vehicle control device equipped with the same, and more specifically, to a signal processing device capable of providing emotion-based highlight images and a vehicle control device equipped with the same.

[0002] A vehicle is a device that moves the user in the desired direction. A typical example is an automobile.

[0003] Meanwhile, for the convenience of users, a vehicle control device is installed inside the vehicle.

[0004] The vehicle control unit includes a signal processing unit and can perform signal processing based on sensor data from various internal sensor devices.

[0005] Meanwhile, the signal processing device can provide images related to vehicle driving upon completion of vehicle driving based on signal processing.

[0006] However, when providing such videos related to vehicle driving, there is a disadvantage in that it is difficult to provide videos that the passengers are interested in.

[0007] The problem that the present disclosure aims to solve is to provide a signal processing device capable of providing emotion-based highlight images, and a vehicle control device equipped with the same.

[0008] Another problem that the present disclosure aims to solve is to provide a signal processing device capable of providing a highlight image including a point of interest based on emotion, and a vehicle control device equipped with the same.

[0009] To solve the above technical problem, a signal processing device according to one embodiment of the present disclosure and a vehicle control device equipped therewith include a memory and a processor that receives camera data from outside the vehicle and camera data from inside the vehicle. The processor classifies emotion data of a passenger based on camera data inside the vehicle, generates a highlight image including an external image from camera data from outside the vehicle corresponding to the classified emotion data, and controls the storage of the generated highlight image in the memory.

[0010] Meanwhile, if the passenger's emotional data is the first emotional data, the processor can extract a first external image at a time corresponding to the first emotional data from camera data outside the vehicle and generate a first highlight image including the first external image.

[0011] Meanwhile, when the passenger's emotional data is the first emotional data and the second emotional data, the processor may extract a first external image at a time corresponding to the first emotional data from camera data outside the vehicle and generate a first highlight image including the first external image, extract a second external image at a time corresponding to the second emotional data from camera data outside the vehicle, and generate a second highlight image including the second external image.

[0012] Meanwhile, the processor can generate a separate highlight video for each of the multiple emotional data when the passenger's emotional data consists of multiple emotional data.

[0013] Meanwhile, if the passenger's emotion data consists of multiple emotion data, the processor can generate a single highlight image including an external image corresponding to the multiple emotion data.

[0014] Meanwhile, the processor can generate a highlight image including an external image and camera information captured from the external image.

[0015] Meanwhile, the processor can generate a map image containing vehicle movement location information and a highlight image containing an external image.

[0016] Meanwhile, the processor can generate location information corresponding to the external image and a highlight image including the external image.

[0017] Meanwhile, the processor can generate a highlight image including an external image from camera data outside the vehicle and an internal image from camera data inside the vehicle, in response to the classified emotion data.

[0018] Meanwhile, the processor can receive additional voice signals from the occupant, extract facial expression information of the occupant from camera data inside the vehicle, and classify the occupant's emotion data based on the facial expression information and the occupant's voice signals.

[0019] Meanwhile, the processor may, when driving a first driving path at a first time point, extract a first external image at a time point corresponding to the first emotional data and generate a first highlight image including the first external image, and when driving a first driving path at a second time point different from the first time point, extract a second external image at a time point corresponding to the second emotional data and generate a second highlight image including the second external image.

[0020] Meanwhile, the processor can play a highlight video based on an event or transmit it to a mobile terminal or another vehicle.

[0021] Meanwhile, the processor can generate multimodal context data based on camera data outside the vehicle, camera data inside the vehicle, and voice signals of the occupant, generate a prompt for classifying the occupant's emotion data based on the multimodal context data, and extract the occupant's emotion data based on an inference result or response result based on the generated prompt.

[0022] Meanwhile, the processor can detect the occupant's gaze position based on camera data inside the vehicle, extract an external image including a point of interest corresponding to the occupant's gaze position, and generate a highlight image including the extracted external image.

[0023] Meanwhile, the processor detects the gaze position of the occupant based on camera data inside the vehicle, performs voice recognition based on the occupant's voice signal, and if the performed voice recognition content includes a confirmation request, extracts an external image including a point of interest corresponding to the gaze position of the occupant at the time of the confirmation request utterance, and generates a highlight image including the extracted external image.

[0024] A signal processing device and a vehicle control device equipped with the same according to one embodiment of the present disclosure further include a neural processor, wherein the processor generates multimodal context data based on camera data outside the vehicle, camera data inside the vehicle, and a voice signal of a passenger, generates a prompt for classifying the passenger's emotion data based on the multimodal context data, receives an inference result or response result based on the prompt generated by the neural processor, and can extract the passenger's emotion data based on the inference result or response result based on the generated prompt.

[0025] Meanwhile, the processor executes a hypervisor and, on the hypervisor, executes a display virtualization machine for a display, and the display virtualization machine extracts emotion data of the occupant based on inference results or response results based on generated prompts, and can generate a highlight image including external images from camera data outside the vehicle based on the extracted emotion data.

[0026] Meanwhile, the processor further runs a gateway virtualization machine for gateway operation on the hypervisor, and the gateway virtualization machine receives camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of the occupant, and can transmit the camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of the occupant to a display virtualization machine through shared memory within the hypervisor.

[0027] A signal processing device and a vehicle control device equipped with the same according to another embodiment of the present disclosure include a memory and a processor that receives camera data outside the vehicle, camera data inside the vehicle, and a voice signal of a passenger. The processor classifies the passenger's emotion data based on the passenger's voice signal and the camera data inside the vehicle, generates a highlight image including an external image from the camera data outside the vehicle corresponding to the classified emotion data, and controls the storage of the generated highlight image in the memory.

[0028] A signal processing device and a vehicle control device equipped with the same according to one embodiment of the present disclosure include a memory and a processor that receives camera data from outside the vehicle and camera data from inside the vehicle. The processor classifies emotion data of a passenger based on camera data inside the vehicle, generates a highlight image including an external image from camera data from outside the vehicle corresponding to the classified emotion data, and controls the storage of the generated highlight image in the memory. Accordingly, it is possible to provide an emotion-based highlight image.

[0029] Meanwhile, if the passenger's emotion data is the first emotion data, the processor can extract a first external image at a time corresponding to the first emotion data from camera data outside the vehicle and generate a first highlight image including the first external image. Accordingly, an emotion-based highlight image can be provided.

[0030] Meanwhile, when the passenger's emotion data is the first emotion data and the second emotion data, the processor may extract a first external image at a time point corresponding to the first emotion data from camera data outside the vehicle and generate a first highlight image including the first external image, and extract a second external image at a time point corresponding to the second emotion data from camera data outside the vehicle and generate a second highlight image including the second external image. Accordingly, it becomes possible to provide an emotion-based highlight image.

[0031] Meanwhile, the processor can generate separate highlight videos for each of the multiple emotion data when the passenger's emotion data consists of multiple emotion data. Accordingly, it becomes possible to provide emotion-based highlight videos.

[0032] Meanwhile, if the passenger's emotion data consists of multiple emotion data, the processor can generate a single highlight video including external images corresponding to the multiple emotion data. Accordingly, it becomes possible to provide an emotion-based highlight video.

[0033] Meanwhile, the processor can generate a highlight video that includes external video and camera information captured from the external video. Accordingly, it becomes possible to provide emotion-based highlight videos.

[0034] Meanwhile, the processor can generate a map image containing vehicle movement location information and a highlight image containing external footage. Accordingly, it becomes possible to provide emotion-based highlight images.

[0035] Meanwhile, the processor can generate location information corresponding to the external image and a highlight image containing the external image. Accordingly, it becomes possible to provide an emotion-based highlight image.

[0036] Meanwhile, the processor can generate a highlight image including an external image from camera data outside the vehicle and an internal image from camera data inside the vehicle, in response to the classified emotion data. Accordingly, it becomes possible to provide an emotion-based highlight image.

[0037] Meanwhile, the processor can receive additional voice signals from the occupant, extract facial expression information of the occupant from camera data inside the vehicle, and classify the occupant's emotion data based on the facial expression information and the occupant's voice signals. Accordingly, it becomes possible to provide emotion-based highlight videos.

[0038] Meanwhile, the processor may extract a first external image at a time corresponding to the first emotional data and generate a first highlight image including the first external image when driving a first driving path at a first time point, and extract a second external image at a time corresponding to the second emotional data and generate a second highlight image including the second external image when driving a first driving path at a second time point different from the first time point. Accordingly, it is possible to provide an emotion-based highlight image.

[0039] Meanwhile, the processor can play highlight videos based on events or transmit them to mobile terminals or other vehicles. Accordingly, it becomes possible to provide emotion-based highlight videos.

[0040] Meanwhile, the processor can generate multimodal context data based on camera data from outside the vehicle, camera data from inside the vehicle, and the passenger's voice signal, generate a prompt for classifying the passenger's emotion data based on the multimodal context data, and extract the passenger's emotion data based on the inference result or response result based on the generated prompt. Accordingly, it becomes possible to provide emotion-based highlight videos.

[0041] Meanwhile, the processor can detect the occupant's gaze position based on camera data inside the vehicle, extract an external image containing a point of interest corresponding to the occupant's gaze position, and generate a highlight image containing the extracted external image. Accordingly, it becomes possible to provide a highlight image containing a point of interest based on emotion.

[0042] Meanwhile, the processor detects the occupant's gaze position based on camera data inside the vehicle, performs speech recognition based on the occupant's voice signal, and if the performed speech recognition content includes a confirmation request, extracts an external image containing a point of interest corresponding to the occupant's gaze position at the time of the confirmation request utterance, and generates a highlight image containing the extracted external image. Accordingly, it becomes possible to provide a highlight image containing a point of interest based on emotion.

[0043] A signal processing device and a vehicle control device equipped with the same according to one embodiment of the present disclosure further include a neural processor, wherein the processor generates multimodal context data based on camera data outside the vehicle, camera data inside the vehicle, and a voice signal of a passenger, generates a prompt for classifying the passenger's emotion data based on the multimodal context data, receives an inference result or response result based on the prompt generated by the neural processor, and extracts the passenger's emotion data based on the inference result or response result based on the generated prompt. Accordingly, it is possible to provide an emotion-based highlight video.

[0044] Meanwhile, the processor executes a hypervisor and, on the hypervisor, executes a display virtualization machine for a display, and the display virtualization machine extracts emotion data of the occupant based on inference results or response results based on generated prompts, and can generate a highlight video including external images from camera data outside the vehicle based on the extracted emotion data. Accordingly, it is possible to provide an emotion-based highlight video.

[0045] Meanwhile, the processor further runs a gateway virtualization machine for gateway operation on the hypervisor, and the gateway virtualization machine receives camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of the occupants, and transmits the camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of the occupants to a display virtualization machine through shared memory within the hypervisor. Accordingly, it becomes possible to provide emotion-based highlight videos.

[0046] A signal processing device and a vehicle control device equipped with the same according to another embodiment of the present disclosure include a memory and a processor that receives camera data from outside the vehicle, camera data from inside the vehicle, and a voice signal of a passenger. The processor classifies the passenger's emotion data based on the passenger's voice signal and the camera data inside the vehicle, generates a highlight image including an external image from the camera data from outside the vehicle corresponding to the classified emotion data, and controls the storage of the generated highlight image in the memory. Accordingly, it is possible to provide an emotion-based highlight image.

[0047] Figure 1 is a drawing illustrating an example of the exterior and interior of a vehicle.

[0048] Figure 2 is a diagram illustrating an example of the architecture of a vehicle control device.

[0049] FIG. 3a is a drawing illustrating an example of the arrangement of displays inside a vehicle.

[0050] Figure 3b is a drawing illustrating another example of the arrangement of displays inside a vehicle.

[0051] FIG. 4 is an example of an internal block diagram of a vehicle control device according to an embodiment of the present disclosure.

[0052] FIG. 5 is an example of a configuration diagram of a signal processing device according to an embodiment of the present disclosure.

[0053] FIG. 6 is an example of a block diagram of a vehicle control device according to an embodiment of the present disclosure.

[0054] FIG. 7 is an example of a block diagram of a signal processing device according to an embodiment of the present disclosure.

[0055] FIGS. 8a to 8c are drawings referenced in the description of FIG. 7.

[0056] FIG. 9 is another example of a block diagram of a signal processing system according to an embodiment of the present disclosure.

[0057] FIG. 10 is an example of a flowchart illustrating a method of operation of a signal processing device according to an embodiment of the present disclosure.

[0058] FIG. 11 is an example of an internal block diagram of a signal processing device according to an embodiment of the present disclosure.

[0059] FIGS. 12a to 26 are drawings referenced in the description of FIG. 10 or FIG. 11.

[0060] The present disclosure will be described in more detail below with reference to the drawings.

[0061] The suffixes "module" and "part" for components used in the following description are assigned solely for the ease of drafting this specification and do not inherently confer any particularly significant meaning or role. Accordingly, the terms "module" and "part" may be used interchangeably.

[0062] Figure 1 is a drawing illustrating an example of the exterior and interior of a vehicle.

[0063] Referring to the drawing, the vehicle (200) is operated by a plurality of wheels (103FR, 103FL, 103RL,...) that rotate by a power source, and a steering wheel (150) for controlling the direction of travel of the vehicle (200).

[0064] Meanwhile, the vehicle (200) may further be equipped with a camera (195), etc., for acquiring an image of the front of the vehicle.

[0065] Meanwhile, the vehicle (200) may be equipped with a plurality of displays (180a, 180b) for displaying images, information, etc. inside.

[0066] In FIG. 1, a cluster display (180a) and an IVI (In-Vehicle Infotainment) display (180b) are exemplified as multiple displays (180a, 180b). In addition, a HUD (Head Up Display) and the like are also possible.

[0067] Meanwhile, the IVI (In-Vehicle Infotainment) display (180b) may also be named a Center Information Display or an AVN (Audio Video Navigation) display.

[0068] Meanwhile, the vehicle (200) described in this specification may be a concept that includes all of the following: a vehicle equipped with an engine as a power source, a hybrid vehicle equipped with an engine and an electric motor as a power source, an electric vehicle equipped with an electric motor as a power source, etc.

[0069] Figure 2 is a diagram illustrating an example of the architecture of a vehicle control device.

[0070] Referring to the drawing, the architecture (300a) of the vehicle control device can correspond to a zone-based architecture.

[0071] Accordingly, sensor devices and processors inside the vehicle may be placed in each of the multiple zones (Z1 to Z4), and a signal processing device (170a) including a gateway (GWDa) may be placed in the central area of ​​the multiple zones (Z1 to Z4).

[0072] Meanwhile, the signal processing device (170a) may additionally include an autonomous driving control module (ACC), a cockpit control module (CPG), etc., in addition to the gateway (GWDa).

[0073] The gateway (GWDa) within the signal processing device (170a) may be a High Performance Computing (HPC) gateway.

[0074] That is, the signal processing device (170a) of FIG. 2 is an integrated HPC and can exchange data with an external communication module (not shown) or a processor (not shown) in a plurality of zones (Z1 to Z4).

[0075] FIG. 3a is a drawing illustrating an example of the arrangement of displays inside a vehicle.

[0076] Referring to the drawing, the vehicle interior may be equipped with a cluster display (180a), an IVI (In-Vehicle Infotainment) display (180b), a rear seat entertainment display (180c, 180d), a rearview mirror display (not shown), etc.

[0077] Meanwhile, in addition to the display, an interior camera (195i) may be installed inside the vehicle.

[0078] Figure 3b is a drawing illustrating another example of the arrangement of displays inside a vehicle.

[0079] A vehicle control device (100) according to an embodiment of the present disclosure may include a plurality of displays (180a to 180b) and a signal processing device (170) that performs signal processing for displaying images, information, etc. on the plurality of displays (180a to 180b) and outputs an image signal to at least one display (180a to 180b).

[0080] Among the plurality of displays (180a to 180b), the first display (180a) is a cluster display (180a) for displaying driving status, operation information, etc., and the second display (180b) may be an IVI (In-Vehicle Infotainment) display (180b) for displaying vehicle operation information, navigation map, various entertainment information or video.

[0081] The signal processing device (170) has a processor (175) inside and can execute a first virtualization machine to a third virtualization machine (not shown) on a hypervisor (not shown) within the processor (175).

[0082] A second virtualization machine (not shown) operates for the first display (180a), and a third virtualization machine (not shown) can operate for the second display (180b).

[0083] Meanwhile, the first virtualization machine (not shown) within the processor (175) can be controlled to set up a shared memory (508) based on a hypervisor (505) for the same data transmission to the second virtualization machine (not shown) and the third virtualization machine (not shown). Accordingly, the same information or the same image can be synchronized and displayed on the first display (180a) and the second display (180b) within the vehicle.

[0084] Meanwhile, the first virtualization machine (not shown) within the processor (175) shares at least a portion of the data with the second virtualization machine (not shown) and the third virtualization machine (not shown) for data sharing processing. Accordingly, data can be shared and processed by multiple virtualization machines for multiple displays within the vehicle.

[0085] Meanwhile, the first virtualization machine (not shown) within the processor (175) can receive and process wheel speed sensor data of the vehicle and transmit the processed wheel speed sensor data to at least one of the second virtualization machine (not shown) or the third virtualization machine (not shown). Accordingly, the wheel speed sensor data of the vehicle can be shared with at least one virtualization machine, etc.

[0086] Meanwhile, the vehicle control device (100) according to the embodiment of the present disclosure may further include a rear seat entertainment display (180c) for displaying driving status information, simple navigation information, various entertainment information or images.

[0087] The signal processing device (170) can control the RSE display (180c) by running a fourth virtualization machine (not shown) in addition to the first to third virtualization machines (not shown) on a hypervisor (not shown) within the processor (175).

[0088] Accordingly, various displays (180a to 180c) can be controlled using a single signal processing device (170).

[0089] Meanwhile, some of the multiple displays (180a to 180c) operate under a Linux OS, and others can operate under a Web OS.

[0090] A signal processing device (170) according to an embodiment of the present disclosure can control displays (180a to 180c) operating under various operating systems (OS) to synchronize and display the same information or the same image.

[0091] Meanwhile, FIG. 3b illustrates that a vehicle speed indicator (212a) and a vehicle interior temperature indicator (213a) are displayed on a first display (180a), a home screen (222) including a plurality of applications, a vehicle speed indicator (212b), and a vehicle interior temperature indicator (213b) is displayed on a second display (180b), and a second home screen (222b) including a plurality of applications and a vehicle interior temperature indicator (213c) is displayed on a third display (180c).

[0092] FIG. 4 is an example of an internal block diagram of a vehicle control device according to an embodiment of the present disclosure.

[0093] Referring to the drawings, a vehicle control device (100) according to an embodiment of the present disclosure may include an input unit (110), a communication unit (120) for communication with an external device, a plurality of communication modules (EMa~EMd) for internal communication, a memory (140), a signal processing unit (170), a plurality of displays (180a~180c), an audio output unit (185), and a power supply unit (190).

[0094] Multiple communication modules (EMa~EMd) can be placed in each of the multiple zones (Z1~Z4) of FIG. 2, for example.

[0095] Meanwhile, the signal processing device (170) may have a communication switch (736b) inside for data communication with each communication module (EM1~EM4).

[0096] Each communication module (EM1~EM4) can perform data communication with a plurality of sensor devices (SN), ECU (770), or area signal processing device (170Z).

[0097] Meanwhile, a plurality of sensor devices (SN) may include a camera (195), lidar (196), radar (197), or position sensor (198).

[0098] The input unit (110) may be equipped with physical buttons, pads, etc. for button input, touch input, etc.

[0099] Meanwhile, the input unit (110) may be equipped with a microphone (not shown) for user voice input.

[0100] The communication unit (120) can exchange data wirelessly with a mobile terminal (600) or a server (400).

[0101] In particular, the communication unit (120) can wirelessly exchange data with the vehicle driver's mobile terminal. Various data communication methods are possible as wireless data communication methods, such as Bluetooth, WiFi, WiFi Direct, and APiX.

[0102] The communication unit (120) can receive weather information, road traffic condition information, for example, TPEG (Transport Protocol Expert Group) information from a mobile terminal (600) or a server (400). To this end, the communication unit (120) may be equipped with a mobile communication module (not shown).

[0103] A plurality of communication modules (EM1~EM4) can receive sensor data, etc. from an ECU (770), a sensor device (SN), or a region signal processing device (170Z), and transmit the received sensor data to the signal processing device (170).

[0104] Here, the sensor data may include at least one of vehicle direction data, vehicle location data (GPS data), vehicle angle data, vehicle speed data, vehicle acceleration data, vehicle tilt data, vehicle forward / reverse data, battery data, fuel data, tire data, vehicle lamp data, vehicle interior temperature data, and vehicle interior humidity data.

[0105] Such sensor data can be obtained from a heading sensor, a yaw sensor, a gyro sensor, a position module, a vehicle forward / reverse sensor, a wheel sensor, a vehicle speed sensor, a vehicle body inclination sensor, a battery sensor, a fuel sensor, a tire sensor, a steering sensor based on steering wheel rotation, a vehicle interior temperature sensor, a vehicle interior humidity sensor, etc.

[0106] Meanwhile, the position module may include a GPS module or a position sensor (198) for receiving GPS information.

[0107] Meanwhile, at least one of the multiple communication modules (EM1 to EM4) can transmit location information data sensed from a GPS module or a location sensor (198) to a signal processing device (170).

[0108] Meanwhile, at least one of the plurality of communication modules (EM1 to EM4) can receive vehicle front image data, vehicle side image data, vehicle rear image data, and obstacle distance information around the vehicle from a camera (195), lidar (196), radar (197), etc., and transmit the received information to a signal processing device (170).

[0109] The memory (140) can store various data for the overall operation of the vehicle control device (100), such as a program for processing or controlling the signal processing device (170).

[0110] For example, memory (140) can store data regarding a hypervisor, a first virtualization machine to a third virtualization machine, for execution within a processor (175).

[0111] The audio output unit (185) converts an electrical signal from the signal processing device (170) into an audio signal and outputs it. To do this, a speaker or the like may be provided.

[0112] The power supply unit (190) can supply power necessary for the operation of each component under the control of the signal processing unit (170). In particular, the power supply unit (190) can receive power from a battery inside the vehicle, etc.

[0113] The signal processing unit (170) controls the overall operation of each unit within the vehicle control unit (100).

[0114] For example, the signal processing device (170) may include a processor (175) that performs signal processing for a vehicle display (180a, 180b).

[0115] The processor (175) can run a first virtualization machine to a third virtualization machine (not shown) on a hypervisor (505 in FIG. 5) within the processor (175).

[0116] Among the first to third virtual machines (not shown), the first virtual machine (not shown) may be named a Server Virtual Machine, and the second to third virtual machines (not shown) may be named a Guest Virtual Machine.

[0117] For example, a first virtualization machine (not shown) within a processor (175) can receive sensor data from a plurality of sensor devices, such as vehicle sensor data, location information data, camera image data, audio data, or touch input data, and process or modify it to output it.

[0118] In this way, by performing most of the data processing in the first virtualization machine (not shown), 1:N data sharing becomes possible.

[0119] As another example, the first virtualization machine (not shown) can directly receive and process CAN data, Ethernet data, audio data, radio data, USB data, and wireless communication data for the second virtualization machine to the third virtualization machine (not shown).

[0120] And, the first virtualization machine (not shown) can transmit the processed data to the second virtualization machine to the third virtualization machine (not shown).

[0121] Accordingly, among the first to third virtualization machines (not shown), only the first virtualization machine (not shown) receives sensor data, communication data, or external input data from a plurality of sensor devices and performs signal processing, thereby reducing the signal processing burden on other virtualization machines and enabling 1:N data communication, which enables synchronization when sharing data.

[0122] Meanwhile, the first virtualization machine (not shown) can control the sharing of the same data with the second virtualization machine (not shown) and the third virtualization machine (not shown) by writing data to the shared memory (508 in FIG. 5).

[0123] For example, the first virtualization machine (not shown) can record vehicle sensor data, the location information data, the camera image data, or the touch input data in a shared memory (508) and control the sharing of the same data with the second virtualization machine (not shown) and the third virtualization machine (not shown). Accordingly, data sharing in a 1:N manner becomes possible.

[0124] Ultimately, by performing most of the data processing on the first virtualization machine (not shown), 1:N data sharing becomes possible.

[0125] Meanwhile, the first virtualization machine (not shown) within the processor (175) can control the second virtualization machine (not shown) and the third virtualization machine (not shown) to set up a shared memory (508) based on the hypervisor (505) for the same data transmission.

[0126] Meanwhile, the signal processing device (170) can process various signals such as audio signals, video signals, and data signals. To this end, the signal processing device (170) can be implemented in the form of a System On Chip (SOC).

[0127] Meanwhile, the signal processing device (170) in the display device (100) of FIG. 4 may be the same as the signal processing device (170) of the vehicle control device of FIG. 5 and below.

[0128] FIG. 5 is an example of a configuration diagram of a signal processing device according to an embodiment of the present disclosure.

[0129] Referring to the drawings, a signal processing device (170) according to an embodiment of the present disclosure includes a processor (175).

[0130] Meanwhile, the signal processing device (170) may be named as an HPC (High Performance Computing) signal processing device.

[0131] Meanwhile, the signal processing device (170) according to an embodiment of the present disclosure may further include a neural processor (179).

[0132] Meanwhile, the neural processor (179) may also be referred to as an on-device-based learning processor.

[0133] Meanwhile, the processor (175) in the signal processing device (170) can execute the hypervisor (505) and execute the first to third virtualization machines (510 to 530) on the hypervisor (505).

[0134] The first virtualization machine (510) may be a gateway virtualization machine corresponding to the gateway (GWDa) of FIG. 2.

[0135] The second virtualization machine (520) may be a driving control virtualization machine corresponding to the autonomous driving control module (ACC) of FIG. 2.

[0136] The driving control virtualization machine (520) at this time can control the vehicle driving assistance (ADAS) or the autonomous driving (AD).

[0137] The third virtualization machine (530) may be a display virtualization machine corresponding to the cockpit control module (CPG) of FIG. 2 or the display (180a, 180b, 180c) of FIG. 3.

[0138] For example, the third virtualization machine (530) may include a cluster virtualization machine (530b) for a cluster display (180a) and an IVI virtualization machine (530b) for an IVI display (180b).

[0139] Meanwhile, the third virtualization machine (530) may further include a HUD virtualization machine (not shown) for a HUD display (180c).

[0140] Meanwhile, the processor (175) can share data between each virtualization machine (510~530) through shared memory (505) within the hypervisor (505).

[0141] Meanwhile, the processor (175) can share data with a plurality of sensor devices (SN) or ECUs (770) or area signal processing devices (170Z) through shared memory (505) within the hypervisor (505).

[0142] Meanwhile, the processor (175) can exchange data with an external server (400) or mobile terminal (600) through the communication device (120) of FIG. 4.

[0143] Meanwhile, the data received by the processor (175) within the signal processing device (170) may include camera data or sensor data.

[0144] For example, sensor data within the vehicle may include at least one of vehicle wheel speed data, vehicle direction data, vehicle location data (GPS data), vehicle angle data, vehicle speed data, vehicle acceleration data, vehicle tilt data, vehicle forward / reverse data, battery data, fuel data, tire data, vehicle lamp data, vehicle interior temperature data, vehicle interior humidity data, vehicle external radar data, and vehicle external lidar data.

[0145] Meanwhile, camera data may include external vehicle camera data and internal vehicle camera data.

[0146] Meanwhile, the processor (175) within the signal processing device (170) can execute multiple virtualization machines (510 to 530) based on safety standards.

[0147] Meanwhile, the processor (175) in the signal processing unit (170a) can execute the hypervisor (505) and, on the hypervisor (505), execute the first to third virtualization machines (510 to 530) according to the automotive safety integrity level (Automotive SIL and ASIL).

[0148] For example, the first virtualization machine (510) may be a virtualization machine corresponding to ASIL C or ASIL D, in which the sum of Severity, Exposure, and Controllability in the Automotive Safety Integrity Level (ASIL) is 9 or 10.

[0149] Meanwhile, ASIL D can correspond to the grade requiring the highest safety level.

[0150] The first virtualization machine (510) can run a safety operating system (not shown) and an application (not shown) on the safety operating system.

[0151] Meanwhile, the first virtualization machine (510) may run a container runtime (not shown) and a container runtime (not shown) on a safety operating system.

[0152] Meanwhile, unlike the drawing, the first virtualization machine (510) may also be executed through a separate processor core instead of the processor (175).

[0153] Meanwhile, the second virtualization machine (520) may be a virtualization machine corresponding to ASIL A or ASIL B, in which the sum of Severity, Exposure, and Controllability in the Automotive Safety Integrity Level (ASIL) is 7 or 8.

[0154] Meanwhile, the second virtualization machine (520) can run an operating system (not shown), a container runtime (not shown) on the operating system (not shown), and a container (not shown) on the container runtime.

[0155] Alternatively, the second virtualization machine (520) can run an operating system (not shown) and an application on the operating system (not shown).

[0156] Meanwhile, the third virtualization machine (530) may be a virtualization machine corresponding to Quality Management (QM), which is the lowest safety level and non-mandatory grade in the Automotive Safety Integrity Level (ASIL).

[0157] Meanwhile, the third virtualization machine (530) can run an operating system (not shown), a container runtime (not shown) on the operating system (not shown), and a container (not shown) on the container runtime.

[0158] Alternatively, the third virtualization machine (530) can run an operating system (not shown) and an application on the operating system (not shown).

[0159] FIG. 6 is an example of a block diagram of a vehicle control device according to an embodiment of the present disclosure.

[0160] Referring to the drawings, a vehicle control device (900) according to an embodiment of the present disclosure comprises a signal processing device (170).

[0161] A vehicle control device (900) according to an embodiment of the present disclosure may further include at least one display.

[0162] In the drawing, at least one display is exemplified as a cluster display (180a) and an IVI display (180b).

[0163] Meanwhile, the vehicle control device (900) may further include a plurality of area signal processing devices (170Z1 to 170Z4).

[0164] The signal processing device (170) at this time is a high-performance centralized signal processing and control device having a plurality of CPUs (175), GPUs (178), NPUs (179), etc., and can be named as a High Performance Computing (HPC) signal processing device or a central signal processing device.

[0165] Multiple area signal processing devices (170Z1~170Z4) and signal processing device (170) are connected by wired cables (CB1~CB4).

[0166] Meanwhile, multiple area signal processing devices (170Z1~170Z4) can be connected to each other by wired cables (CBa~CBd).

[0167] The wired cable (CBa~CBd) at this time may include a CAN communication cable, an Ethernet communication cable, or a PCI Express cable.

[0168] Meanwhile, the signal processing device (170) according to the embodiment of the present disclosure may have at least one processor (175, 178, 177) and a large-capacity storage device (925).

[0169] For example, a signal processing device (170) according to an embodiment of the present disclosure may include a central processor (175, 177), a graphics processor (178), and a neural processor (179).

[0170] Meanwhile, sensor data can be transmitted from at least one of the multiple area signal processing devices (170Z1 to 170Z4) to the signal processing device (170). In particular, the sensor data can be stored in a storage device (925) within the signal processing device (170).

[0171] The sensor data at this time may include at least one of camera data, lidar data, radar data, vehicle direction data, vehicle position data (GPS data), vehicle angle data, vehicle speed data, vehicle acceleration data, vehicle tilt data, vehicle forward / reverse data, battery data, fuel data, tire data, vehicle lamp data, vehicle interior temperature data, and vehicle interior humidity data.

[0172] In the drawing, camera data from a camera (195a) and lidar data from a lidar sensor (196) are input to a first area signal processing device (170Z1), and the camera data and lidar data are transmitted to a signal processing device (170) via a second area signal processing device (170Z2) and a third area signal processing device (170Z3), etc.

[0173] Meanwhile, since the data reading or writing speed to the storage device (925) is faster than the network speed when sensor data is transmitted from at least one of the multiple area signal processing devices (170Z1~170Z4) to the signal processing device (170), it is desirable to perform multipath routing so that network bottlenecks do not occur.

[0174] To this end, the signal processing device (170) according to an embodiment of the present disclosure can perform multipath routing based on a Software Defined Network (SDN). Accordingly, a stable network environment can be secured when reading or writing data of the storage device (925). Furthermore, since data can be transmitted to the storage device (925) using multiple paths, data can be transmitted by dynamically changing the network configuration.

[0175] Data communication between a plurality of area signal processing devices (170Z1~170Z4) and a signal processing device (170) within a vehicle control device (900) according to an embodiment of the present disclosure is preferably Peripheral Component Interconnect Express communication or Ethernet communication for high-bandwidth, low-latency communication.

[0176] FIG. 7 is an example of a block diagram of a signal processing device according to an embodiment of the present disclosure.

[0177] Referring to the drawings, the signal processing system (700) according to an embodiment of the present disclosure may be referred to as an IVEX (In Vehicle Experience) system.

[0178] Meanwhile, the signal processing system (700) according to an embodiment of the present disclosure may include an edge sensor group (SN), a signal processing device (170) in a vehicle, and a server (400).

[0179] Meanwhile, the signal processing device (170) according to an embodiment of the present disclosure includes a recognizer (720), an insightor (750), and an illustrator (780).

[0180] Meanwhile, the edge sensor group (SN) can correspond to a plurality of sensor devices (SN) of FIG. 2.

[0181] Meanwhile, the recognizer (720), the insightor (750), and the illustrator (780) may be included in the signal processing device (170) of the vehicle display device (100) of FIG. 4.

[0182] In particular, the recognizer (720), the insightor (750), and the illustrator (780) may be included in the processor (175) of the signal processing device (170).

[0183] Meanwhile, the edge sensor group (SN) may include vehicle interior sensors (710) and vehicle exterior sensors (715).

[0184] Meanwhile, the edge sensor group (SN) may further include a data receiving unit (718) for receiving vehicle internal data or vehicle external data. The data receiving unit (718) may correspond to the communication device (120) of FIG. 2.

[0185] Meanwhile, the vehicle interior sensors (710) are sensors placed inside the vehicle (200) and may include a front camera (195), lidar (196), radar (197), interior camera (195i), vehicle interior temperature sensor, or vehicle interior humidity sensor.

[0186] The vehicle external sensors (715) are sensors positioned on the exterior of the vehicle (200) and may include an external camera, lidar (196), radar (197), heading sensor, yaw sensor, gyro sensor, position module, vehicle forward / reverse sensor, wheel sensor, vehicle speed sensor, vehicle body inclination sensor, battery sensor, fuel sensor, tire sensor, or steering sensor based on steering wheel rotation. The position module may include a GPS module or a position sensor (198) for receiving GPS information.

[0187] Sensing data may include vehicle interior sensing data and vehicle exterior sensing data.

[0188] The vehicle interior sensing data may be data sensed by the vehicle interior sensors (710).

[0189] Vehicle interior sensing data may include at least one of vehicle interior temperature data, vehicle interior humidity data, battery data, fuel data, vehicle lamp data, tire data, or vehicle interior camera data, or audio data received through a microphone.

[0190] The vehicle external sensing data may be data sensed by the vehicle external sensors (715).

[0191] Vehicle external sensing data may include at least one of vehicle location data (GPS), vehicle direction data, vehicle angle data, vehicle speed data, vehicle acceleration data, vehicle tilt data, whether the vehicle is moving forward or backward, front camera data, or rear camera data.

[0192] Meanwhile, the recognizer (720) can recognize the vehicle situation based on sensing data received from the edge sensor group (SN). In this regard, the recognizer (720) may be referred to as a situation recognition unit.

[0193] The vehicle situation at this time may include external vehicle conditions and internal vehicle conditions.

[0194] Meanwhile, the recognizer (720) can recognize the vehicle situation based on the vehicle internal sensing data or the vehicle internal sensing data, and can generate vehicle situation information regarding the recognized vehicle situation.

[0195] For example, the vehicle situation information (740) may include at least one of the external situation information of the vehicle and the situation information of the vehicle's occupants.

[0196] As another example, the vehicle situation information (740) may include at least one of the vehicle's external situation information, the vehicle's internal situation information, and the vehicle's occupant situation information.

[0197] Meanwhile, the recognizer (720) can recognize external situation information of the vehicle based on external sensing data from the edge sensor group (SN), recognize internal situation information of the vehicle based on internal sensing data of the vehicle, or recognize situation information of the vehicle's occupant based on internal sensing data of the vehicle.

[0198] Meanwhile, the recognizer (720) can output external situation information of the vehicle based on internal vehicle sensing data or external vehicle sensing data, output internal situation information of the vehicle based on internal vehicle sensing data or external vehicle sensing data, or output situation information of the vehicle's occupant based on internal vehicle sensing data or external vehicle sensing data.

[0199] Meanwhile, the recognizer (720) can perform preprocessing and calibration of the vehicle interior sensing data or the vehicle interior sensing data, and can obtain vehicle situation information based on the preprocessed and calibrated sensing data.

[0200] For example, the internal recognition unit (725) within the recognition unit (720) can perform preprocessing and calibration (721) on the vehicle internal sensing data, and can obtain internal situation information of the vehicle based on the preprocessed and calibrated vehicle internal sensing data.

[0201] Specifically, the internal recognizer (725) can perform preprocessing and calibration (721) on the vehicle interior sensing data, perform primitive detection (722) or sensor fusion (723), and based on this, obtain information about the vehicle's interior situation.

[0202] For example, an external recognizer (730) within the recognizer (720) can perform preprocessing and calibration (721) on the external vehicle sensing data, and can obtain external situation information of the vehicle based on the preprocessed and calibrated external vehicle sensing data.

[0203] Specifically, the external recognizer (730) can perform preprocessing and calibration (731) on external vehicle sensing data, execute a perception engine (732), perform sensor fusion (733), or perform segmentation (734), and based on this, obtain external situation information of the vehicle.

[0204] Meanwhile, the passenger recognition device (730) within the recognition device (720) can perform facial recognition (741) of the passenger based on the vehicle interior sensing data, recognize distraction (742), recognize gaze, position, gesture (743), recognize emotion (744), recognize whether the passenger is drinking or taking medication (746), recognize drowsiness (747), or recognize information (748) regarding other actions, and obtain situational information of the vehicle's passenger based thereon.

[0205] Meanwhile, the situational information of the occupant may include at least one of the following: a face identifier (Face ID) of the occupant inside the vehicle (200), distraction information, gaze information, position information, gesture information, emotion information, information on whether alcohol or drugs have been taken, drowsiness information, or other information regarding actions.

[0206] Meanwhile, the recognizer (720) can transmit the generated vehicle situation information to the insightor (750).

[0207] Meanwhile, the insightor (750) can generate an inference result or a response result based on the vehicle situation information received from the recognizer (720).

[0208] Meanwhile, the insightor (750) can generate an inference result or a response result based on the vehicle situation information received from the recognizer (720), and can generate or control a service to be executed based on the inference result or the response result.

[0209] Accordingly, the insightor (750) may be named a response result generation unit or a service generation unit.

[0210] Meanwhile, the insightor (750) may include a multimodal context engine (755), a context controller (754), a safety assistance engine (757), a user characteristic engine (758), an AI orchestrator (760), and an AI model set (765).

[0211] Meanwhile, the multimodal context engine (755) can generate multimodal context data based on the vehicle situation information received from the recognizer (720).

[0212] For example, multimodal context data may include at least one of text data or image data describing a vehicle situation generated based on vehicle situation information.

[0213] Meanwhile, the context controller (754) may be a component included in the multimodal context engine (755) or provided separately from the multimodal context engine (755).

[0214] Meanwhile, the context controller (754) can execute the process of generating multimodal context data when it receives a passenger query from the AI ​​orchestrator (760).

[0215] For example, the passenger curie may be a voice recognition result corresponding to a voice command spoken by the passenger.

[0216] Meanwhile, the safety assist engine (757) can determine whether the current situation is a safe situation or a dangerous situation based on the vehicle situation information received from the recognizer (720).

[0217] Meanwhile, the safety assist engine (757) can transmit a driver assistance control command or a warning notification output command to the safety application (790) of the illustrator (780) if the current situation is determined to be a dangerous situation.

[0218] Meanwhile, the user characteristic engine (758) can generate user context data or passenger context data based on user information or passenger information.

[0219] Meanwhile, user information or passenger information may be referred to as user persona or passenger persona.

[0220] Meanwhile, user information or passenger information may include at least one of the user or passenger's nationality, age, gender, occupation, personality, or psychological type (Myers-Briggs Type Indicator, MBTI).

[0221] Meanwhile, the AI ​​orchestrator (760) can generate a prompt based on at least one of multimodal context data or passenger context data.

[0222] Meanwhile, the AI ​​orchestrator (760) can send the generated prompt to the AI ​​model set (765).

[0223] Meanwhile, the AI ​​model set (765) may include at least one AI model.

[0224] For example, the AI ​​model set (765) may include an on-device AI model (767).

[0225] Meanwhile, the on-device AI model (767) may include at least one AI model. For example, the on-device AI model (767) may include a Small Language Model (LLM).

[0226] Meanwhile, the AI ​​model set (765) may include an interface (766) for data exchange with the AI ​​model (769) in the server (400).

[0227] Meanwhile, the AI ​​model (769) within the server (400) may include at least one AI model. For example, the AI ​​model (769) within the server (400) may include a Large Language Model (LLM).

[0228] That is, the on-device AI model (767) may have a smaller capacity or size than the AI ​​model (769) in the server (400).

[0229] Meanwhile, the AI ​​model set (765) can output an inference result in response to a prompt received from the AI ​​orchestrator (760) and can transmit the inference result to the AI ​​orchestrator (760).

[0230] Meanwhile, the AI ​​orchestrator (760) can generate additional prompts based on the inference results and send the additional prompts to the AI ​​model set (765).

[0231] Meanwhile, the AI ​​model set (765) that receives the additional prompt can output an additional inference result in response to the additional prompt and can transmit the additional inference result to the AI ​​orchestrator (760).

[0232] The AI ​​orchestrator (760) can obtain an inference result or additional inference result received from the AI ​​model set (765) as a response result, and can output the obtained response result to the illustrator (780).

[0233] Meanwhile, the illustrator (780) can execute a service, output service information, execute an application, or output application information based on the response result output from the insightor (750).

[0234] For example, the illustrator (780) can execute a navigation service, execute a vehicle driving assistance control service, execute an autonomous driving service, or execute a display-related service based on the response result output from the insightor (750).

[0235] As another example, the illustrator (780) can run a navigation application, run a vehicle driving assistance control application, run an autonomous driving application, or run a display-related application based on the response result output from the insightor (750).

[0236] Meanwhile, the illustrator (780) may include a multimodal output encoder (781), a visual interface (782), an audio interface (785), and a safety application (790).

[0237] Meanwhile, the multimodal output encoder (781) can encode the response result output from the insightor (750) and output the encoded response result data to the visual interface (782) or audio interface (785).

[0238] Meanwhile, the visual interface (782) or audio interface (785) may be referred to as an output interface.

[0239] Meanwhile, the visual interface (782) can output response result data output from the multimodal output encoder (781).

[0240] Accordingly, at least one of the plurality of displays (180a to 180c) of FIG. 4 can display an image based on response result data from the visual interface (782).

[0241] Meanwhile, the visual interface (782) can output augmented reality (AR) video based on response result data or mixed reality (MR) video data based on response result data.

[0242] Meanwhile, the audio interface (785) can output response result data in the form of audio.

[0243] Accordingly, the audio output unit (185) of FIG. 4 can output a sound corresponding to the response result data from the audio interface (785).

[0244] Meanwhile, the safety application (790) can perform Advanced Driver Assistance System (ADAS) control based on the driver assistance control command or the response result received from the insightor (750).

[0245] For example, the safety application (790) can output warning notification data according to a warning notification output command.

[0246] In response to this, at least one of the electronic control unit (770) of FIG. 4 or a plurality of displays (180a to 180c) can output warning notification data.

[0247] FIGS. 8a to 8c are drawings referenced in the description of FIG. 7.

[0248] First, FIG. 8a is a drawing referenced in the description of the insight of FIG. 7.

[0249] Referring to the drawing, the insightor (750) may include a multimodal context engine (755), an AI orchestrator (760), and a multimodal LLM (767).

[0250] Meanwhile, the multimodal context engine (755) may include a multimodal signal adapter (751), a multimodal indexer (812), a multimodal context buffer (752), a multimodal context retriever (814), a multimodal context descriptor (815), a multimodal event monitor (811), and a context controller (754).

[0251] Unlike Fig. 7, the multimodal context engine (755) may include a context controller (754).

[0252] Meanwhile, the multimodal signal adapter (751) can generate pre-processed multimodal data by filtering, cleaning, synchronizing, and reformulating data received from various sensors or vehicle situation information received from the recognizer (720).

[0253] Meanwhile, the multimodal signal adapter (751) can receive at least one of audio data received through a microphone, front image data captured through a front camera, ADAS information obtained from the front image data or vehicle sensor, location data, IVI (In-Vehicle Infotainment) system data, IVI display information, DMS (Driver Monitoring System) information or IMS (Interior Monitoring System) information based on image data captured through an interior camera (195i), and biometric data obtained from a biometric sensor.

[0254] Meanwhile, the multimodal indexer (812) can generate multimodal processing data by dividing the preprocessed multimodal data into chunks and can index the multimodal processing data.

[0255] Meanwhile, the multimodal indexer (812) can obtain an encoding vector or keyword representing the attribute (or meaning) of the multimodal processed data divided into chunks as an index.

[0256] Meanwhile, the multimodal context buffer (752) can store multimodal processing data and an index corresponding to the multimodal processing data.

[0257] Meanwhile, the multimodal context retriever (814) can search for multimodal processing data most related to the passenger query through an index from the multimodal context buffer (752), and can select the searched multimodal processing data as a context candidate for the creation of multimodal context data.

[0258] Meanwhile, the multimodal context descriptor (815) can reconfigure multimodal processing data selected as a context candidate into multimodal context data having a prompt form that the multimodal LLM (767) can interpret.

[0259] Meanwhile, the multimodal event monitor (811) can monitor whether an index matching the trigger condition is entered when a trigger condition for multimodal data is registered. If an index matching the trigger condition is entered, the multimodal event monitor (811) can generate a trigger event to generate a proactive service query.

[0260] Meanwhile, the context controller (754) can control the overall operation of the multimodal context engine (755).

[0261] Meanwhile, the context controller (754) can execute the process of generating multimodal context data when it receives a passenger query from the AI ​​orchestrator (760).

[0262] Meanwhile, the context controller (754) can generate a preemptive service query when it receives a trigger event from the multimodal event monitor (811).

[0263] Meanwhile, the AI ​​orchestrator (760) can generate a prompt by combining the passenger query and the multimodal context, and can send the generated prompt to the multimodal LLM (767).

[0264] Meanwhile, the passenger curry may be a query in the form of recognized text based on a voice command spoken by the passenger. The voice command may be received through a microphone, and the voice command may be converted into text through an Automatic Speech Recognition (ASR) process.

[0265] Meanwhile, passenger curry may be text converted through an Automatic Speech Recognition (ASR) process.

[0266] Meanwhile, multimodal LLM (767) may be an example of an AI model included in the set of AI models (765) of FIG. 6a.

[0267] Meanwhile, the multimodal LLM (767) can output an inference result from a prompt received from the AI ​​orchestrator (760) and can transmit the inference result to the AI ​​orchestrator (760).

[0268] Meanwhile, the inference result may include a result representing a response to the prompt.

[0269] For example, the inference result may include an API call result regarding whether an API call corresponding to a driving assistance function was successfully performed, and a feedback generation result regarding whether feedback corresponding to the API call was successfully generated.

[0270] Next, Fig. 8b is a drawing referenced in the description of the AI ​​orchestrator of Fig. 7.

[0271] Referring to the drawing, the AI ​​orchestrator (760) may include a task arbitrator (771), a prompt manager (763), a sub-agent set (761), a knowledge database (762), a workflow controller (772), and a tool set / adapter set (764).

[0272] Meanwhile, the task arbitrator (771) can determine one of the multiple sub-agents based on the voice recognition result and the multimodal context data output from the multimodal context engine (755).

[0273] Meanwhile, the prompt manager (763) can generate a prompt for the operation of the determined sub-agent. The prompt manager (763) can generate a prompt based on multimodal context data and passenger queries.

[0274] Meanwhile, the prompt manager (763) can generate a prompt based on information about the functions that the determined sub-agent can perform, multimodal context data, the results of previously performed functions, conversation history and passenger query.

[0275] Meanwhile, the sub-agent set (761) may include multiple sub-agents.

[0276] For example, the sub-agent set (761) may include a navigation agent for navigation services, a vehicle function agent for providing vehicle functions, a telephony agent for automated telephone answering services, and a Q&A agent for providing response services to queries.

[0277] Meanwhile, the sub-agent determined by the task arbitrator (771) can call the cloud AI model (769) or on-device AI model (767) assigned to it.

[0278] Meanwhile, the cloud AI model (769) or on-device AI model (767) may be a Large Language Model (LLM).

[0279] Meanwhile, a sub-agent within a sub-agent set (761) can call a cloud AI model (769) or an on-device AI model (767) assigned to it to obtain an inference result corresponding to a prompt from the model.

[0280] Meanwhile, a sub-agent in the sub-agent set (761) can call the knowledge database (762) to provide additional information based on the obtained inference result, obtain additional information from the knowledge database (762), specify the name of the function to be executed and the parameter value of the function, and determine the feedback phrase to be provided to the user.

[0281] Meanwhile, a sub-agent in the sub-agent set (761) can transmit to the workflow controller (772) a parsing result including additional information called from the knowledge database (762), the name of the function to be executed, the parameter value of the function, and a feedback phrase, which is generated by parsing the inference result received from the model.

[0282] Meanwhile, the knowledge database (762) can store additional information and information about functions. The information about functions may include the name of the function and the parameter values ​​of the function.

[0283] Meanwhile, the workflow controller (772) can generate control commands to perform the corresponding function based on the parsing results and can store the results of the conversation and function performed by the AI ​​model and sub-agent.

[0284] Meanwhile, the tool set / adapter set (764) can call an API corresponding to a control command generated by the workflow controller (772). The tool set / adapter set (764) can convert the control command into an execution command of an actual function or an API call command that the IVI system can understand and execute, and can execute the converted command.

[0285] Fig. 8c is an example of an internal block diagram of the server of Fig. 7.

[0286] Referring to the drawing, the server (400) may represent a device that trains an artificial neural network using a machine learning algorithm or uses a trained artificial neural network.

[0287] Here, the server (400) may be composed of multiple servers to perform distributed processing. Alternatively, the server (400) may be defined as a 5G network.

[0288] The server (400) may be included as part of the vehicle (200) and may perform at least some of the AI ​​processing together.

[0289] The server (400) may include a communication interface (410), memory (430), a learning processor (440), and a processor (470).

[0290] The communication interface (410) can transmit and receive data with the vehicle (200) or an external device.

[0291] The memory (430) may include a model storage unit (431). The model storage unit (431) may store a model (or artificial neural network, 531a) that is being learned or has been learned through the learning processor (440).

[0292] The learning processor (440) can train the artificial neural network (431a) using the training data.

[0293] For example, the learning processor (440) can execute a learning model. The learning model may include a cloud AI model (769) such as FIG. 7.

[0294] The learning model may be used while mounted on the server (400) of the artificial neural network, or may be used while mounted on an external device such as a vehicle (200).

[0295] The learning model may be implemented in hardware, software, or a combination of hardware and software. If part or all of the learning model is implemented in software, one or more instructions constituting the learning model may be stored in memory (430).

[0296] The processor (470) can infer a result value for new input data using a learning model and generate a response or control command based on the inferred result value.

[0297] Meanwhile, the processor (470) can control the learning processor (440) to execute the cloud AI model (769) based on requests from the signal processing device (170) in the vehicle (200), etc.

[0298] FIG. 9 is another example of a block diagram of a signal processing system according to an embodiment of the present disclosure.

[0299] Referring to the drawings, the signal processing system (900) according to an embodiment of the present disclosure may be referred to as an IVEX (In Vehicle Experience) system or a vehicle control device.

[0300] Meanwhile, the signal processing system (900) according to an embodiment of the present disclosure includes an edge sensor group (SN) and a signal processing device (170) in a vehicle.

[0301] Meanwhile, the signal processing system (900) according to an embodiment of the present disclosure may further include a server (400).

[0302] Meanwhile, the signal processing device (170) according to an embodiment of the present disclosure includes a recognizer (720), an insightor (750), and an illustrator (780).

[0303] Meanwhile, the edge sensor group (SN) can correspond to a plurality of sensor devices (SN) of FIG. 2.

[0304] Meanwhile, the edge sensor group (SN) may include a microphone (112) which is an example of an input device (110), a vehicle interior camera (195i), a vehicle front camera (195), and a position sensor (198) that receives GPS data, etc.

[0305] Among these, the microphone (112) and the vehicle interior camera (195i) may be included in the vehicle interior sensors (710) of FIG. 7.

[0306] Meanwhile, the vehicle front camera (195) or position sensor (198) may be included in the vehicle external sensors (715).

[0307] Meanwhile, the recognizer (720) can recognize the vehicle situation based on sensing data received from the edge sensor group (SN). In this regard, the recognizer (720) may be referred to as a situation recognition unit.

[0308] The vehicle situation at this time may include external vehicle conditions and internal vehicle conditions.

[0309] Meanwhile, the recognizer (720) can recognize the vehicle situation based on the vehicle internal sensing data or the vehicle internal sensing data, and can generate vehicle situation information regarding the recognized vehicle situation.

[0310] For example, the vehicle situation information (740) may include at least one of the external situation information of the vehicle and the situation information of the vehicle's occupants.

[0311] As another example, the vehicle situation information (740) may include at least one of the vehicle's external situation information, the vehicle's internal situation information, and the vehicle's occupant situation information.

[0312] Meanwhile, the recognizer (720) can recognize external situation information of the vehicle based on external sensing data from the edge sensor group (SN), recognize internal situation information of the vehicle based on internal sensing data of the vehicle, or recognize situation information of the vehicle's occupant based on internal sensing data of the vehicle.

[0313] Meanwhile, the recognizer (720) can output external situation information of the vehicle based on internal vehicle sensing data or external vehicle sensing data, output internal situation information of the vehicle based on internal vehicle sensing data or external vehicle sensing data, or output situation information of the vehicle's occupant based on internal vehicle sensing data or external vehicle sensing data.

[0314] Meanwhile, the recognizer (720) can perform preprocessing and calibration of the vehicle interior sensing data or the vehicle interior sensing data, and can obtain vehicle situation information based on the preprocessed and calibrated sensing data.

[0315] For example, the recognizer (720) can perform preprocessing and calibration on the vehicle interior sensing data, and can obtain information on the vehicle's interior situation based on the preprocessed and calibrated vehicle interior sensing data.

[0316] For example, the recognizer (720) can perform preprocessing and calibration on the vehicle external sensing data, and can obtain external situation information of the vehicle based on the preprocessed and calibrated vehicle external sensing data.

[0317] Meanwhile, the recognition device (720) can perform facial identification of the occupant based on the vehicle interior sensing data, recognize distraction, recognize gaze, position, gesture, recognize emotion, recognize whether alcohol or drugs have been taken, recognize drowsiness, or recognize information (748) about other actions, and obtain situational information of the vehicle occupant based thereon.

[0318] Meanwhile, the situational information of the occupant may include at least one of the following: a face identifier (Face ID) of the occupant inside the vehicle (200), distraction information, gaze information, position information, gesture information, emotion information, information on whether alcohol or drugs have been taken, drowsiness information, or other information regarding actions.

[0319] Meanwhile, the recognizer (720) can transmit the generated vehicle situation information to the insightor (750).

[0320] Meanwhile, the recognizer (720) performs voice recognition based on the passenger's voice signal from the microphone (112), and if a confirmation request is included in the performed voice recognition content, it extracts the confirmation request (724).

[0321] For example, the confirmation request (724) may include phrases such as "What's that?"

[0322] And, the recognizer (720) can send a confirmation request to the AI ​​orchestrator (760).

[0323] Meanwhile, the recognizer (720) receives camera data from outside the vehicle from the vehicle's external camera (195) and performs object detection based on the camera data from outside the vehicle.

[0324] For example, the recognizer (720) can receive vehicle front image data from the vehicle's external camera (195), perform segmentation on the vehicle front image data, and detect multiple objects (735).

[0325] Meanwhile, the recognizer (720) receives location information (737) from the location sensor (178) and can transmit the location information (737) to the insightor (750).

[0326] Meanwhile, the recognizer (720) receives camera data inside the vehicle from the vehicle's internal camera (195i) and detects the occupant's gaze position (743) based on the camera data inside the vehicle.

[0327] For example, the recognizer (720) can receive image data of the interior of the vehicle from the vehicle's interior camera (195i) and detect the gaze position (743) of the occupant, particularly the driver, among the image data of the interior of the vehicle.

[0328] And, the recognizer (720) can select an object (743b) corresponding to the occupant's gaze position based on the extracted gaze position (743) and a plurality of objects (735).

[0329] Meanwhile, the recognizer (720) can crop or extract an area corresponding to the passenger's gaze position (743) from the vehicle's front image data from the vehicle's external camera (195) based on the passenger's gaze position (743).

[0330] For example, the recognizer (720) can crop or extract an area containing an object corresponding to the passenger's gaze position (743) among a plurality of objects (735).

[0331] Meanwhile, the recognizer (720) can transmit an object (743b) corresponding to the passenger's gaze position to the insightor (750).

[0332] Additionally, the recognizer (720) can further transmit multiple objects (735), a gaze position (743), and location information (737) to the insightor (750), respectively.

[0333] That is, the recognizer (720) can transmit an object (743b) corresponding to the passenger's gaze position, a plurality of objects (735), a gaze position (743), and location information (737) to the insightor (750).

[0334] Meanwhile, the insightor (750) can generate an inference result or a response result based on the vehicle situation information received from the recognizer (720).

[0335] Meanwhile, the insightor (750) can generate an inference result or a response result based on the vehicle situation information received from the recognizer (720), and can generate or control a service to be executed based on the inference result or the response result.

[0336] Accordingly, the insightor (750) may be named a response result generation unit or a service generation unit.

[0337] Meanwhile, the insightor (750) may include a multimodal context engine (755), an AI orchestrator (760), and an AI model set (765).

[0338] Meanwhile, the multimodal context engine (755) can generate multimodal context data based on the vehicle situation information received from the recognizer (720).

[0339] For example, multimodal context data may include at least one of text data, image data, or audio data describing a vehicle situation generated based on vehicle situation information.

[0340] Meanwhile, the multimodal context engine (755) may include a multimodal signal adapter (751), a multimodal context buffer (752), a multimodal context retriever (814), etc.

[0341] The multimodal signal adapter (751) can receive an object (743b) corresponding to the passenger's gaze position, a plurality of objects (735), a gaze position (743), and location information (737).

[0342] Additionally, the multimodal signal adapter (751) can control the synchronization of received data, such as an object (743b) corresponding to the passenger's gaze position, a plurality of objects (735), a gaze position (743), and location information (737), and convert it into synchronized vector data to be stored in a database within the multimodal context buffer (752).

[0343] In particular, the multimodal signal adapter (751) can be controlled to store synchronized vector data for a predetermined period (e.g., 10 seconds) in a database within the multimodal context buffer (752).

[0344] Accordingly, context analysis before and after the receipt of a voice recognition-based confirmation request becomes possible.

[0345] The multimodal context buffer (752) can store synchronized vector data.

[0346] The multimodal context retriever (814) can receive a voice recognition-based confirmation request from the AI ​​orchestrator (760).

[0347] When a voice recognition-based confirmation request is received, the multimodal context retriever (814) can search for multimodal processing data related to the voice recognition-based confirmation request through the index of the multimodal context buffer (752) based on the confirmation request, and can select the searched multimodal processing data as a context candidate for the creation of multimodal context data.

[0348] Meanwhile, the multimodal context retriever (814) can receive voice recognition-based confirmation requests and map data from the AI ​​orchestrator (760).

[0349] Meanwhile, when a voice recognition-based confirmation request and map data are received, the multimodal context retriever (814) can search for multimodal processing data related to the voice recognition-based confirmation request through the index of the multimodal context buffer (752) based on the confirmation request or map data, and can select the searched multimodal processing data as a context candidate for the creation of multimodal context data.

[0350] Meanwhile, the multimodal context retriever (814) can reconfigure multimodal processing data selected as a context candidate into multimodal context data having a prompt form that the multimodal LLM (767) can interpret.

[0351] For example, the multimodal context retriever (814) can extract context information corresponding to a voice recognition-based confirmation request from a vector database within the multimodal context buffer (752), and extract attribute information for multiple extracted objects and the context corresponding to the confirmation request.

[0352] And, the multimodal context retriever (814) can transmit, based on the extracted context, an object (743b) corresponding to the passenger's gaze position, multiple objects (735), gaze position (743), location information (737), and map data to the prompt manager (763) in the AI ​​orchestrator (760).

[0353] Meanwhile, the AI ​​orchestrator (760) can receive a voice recognition-based confirmation request from the recognizer (720) and control the generation of multimodal context data based on the confirmation request.

[0354] To this end, the AI ​​orchestrator (760) may include a confirmation request agent (761w), a prompt manager (763), a user attribute database (794), a record agent (797), etc.

[0355] The AI ​​orchestrator (760) can transmit the voice recognition-based confirmation request to the multimodal context retriever (814) when the voice recognition-based confirmation request is received from the recognizer (720).

[0356] In particular, when the AI ​​orchestrator (760) receives a voice recognition-based confirmation request from the recognizer (720), it can transmit the voice recognition-based confirmation request and map data to the multimodal context retriever (814).

[0357] Meanwhile, the prompt manager (763) within the AI ​​orchestrator (760) can receive from the multimodal context retriever (814) an object (743b) corresponding to the rider's gaze position, multiple objects (735), a gaze position (743), location information (737), and map data to the prompt manager (763) within the AI ​​orchestrator (760).

[0358] Meanwhile, the prompt manager (763) within the AI ​​orchestrator (760) can generate a confirmation request and a prompt corresponding to the user characteristics by referring to the vector database within the user characteristics database (794), and transmit the generated prompt to the confirmation request agent (761w).

[0359] The confirmation request agent (761w) can control the operation of an on-device AI model (767) in the AI ​​model set (765) or an AI model (769) in the server (400) based on the received prompt.

[0360] At this time, the AI ​​model (769) in the server (400) can share data with the AI ​​model set (765) through the interface (766) in the AI ​​model set (765).

[0361] For example, the confirmation request agent (761w) can distribute the load of the on-device AI model (767) in the AI ​​model set (765) or the AI ​​model (769) in the server (400) based on the received prompt, and control the on-device AI model (767) or the AI ​​model (769) in the server (400) to operate respectively based on the load distribution.

[0362] Meanwhile, the confirmation request agent (761w) can receive an inference result or response result corresponding to the prompt from an on-device AI model (767) or an AI model (769) within the server (400).

[0363] And, the confirmation request agent (761w) can transmit the inference result or response result to the illustrator (780).

[0364] Meanwhile, the illustrator (780) can execute a service, output service information, execute an application, or output application information based on the inference result or response result output from the insightor (750).

[0365] For example, the illustrator (780) can execute a navigation service or a display-related service based on the response result output from the insightor (750).

[0366] Meanwhile, the illustrator (780) may include a multimodal output encoder (781) including a modality manager (795), a visual interface (782), an audio interface (785), etc.

[0367] Meanwhile, the multimodal output encoder (781) can encode the inference result or response result output from the insightor (750) and output the encoded inference result data or response result data to the visual interface (782) or audio interface (785).

[0368] Meanwhile, the visual interface (782) or audio interface (785) may be referred to as an output interface.

[0369] Meanwhile, the visual interface (782) can output response result data output from the multimodal output encoder (781).

[0370] Accordingly, at least one of the plurality of displays (180a to 180c) of FIG. 4 can display an image based on response result data from the visual interface (782).

[0371] Meanwhile, the visual interface (782) can output augmented reality (AR) video data (783) based on response result data or mixed reality (MR) video data (784) based on response result data.

[0372] Meanwhile, the audio interface (785) can output response result data in the form of audio.

[0373] For example, the audio interface (785) can convert text-based response result data into audio data and output the converted audio data (798).

[0374] Accordingly, the audio output unit (185) of FIG. 4 can output a sound corresponding to the response result data from the audio interface (785).

[0375] In summary, FIG. 9, the processor (175) in the signal processing device (170) receives camera data from outside the vehicle, camera data from inside the vehicle, and a voice signal of the passenger, performs object detection based on the camera data from outside the vehicle, detects the passenger's gaze position based on the camera data inside the vehicle, and performs voice recognition based on the passenger's voice signal.

[0376] Meanwhile, if a confirmation request is included in the voice recognition content performed, the processor (175) extracts multimodal context data corresponding to the confirmation request and selects a point of interest corresponding to the passenger's gaze position based on an inference result or response result obtained based on the extracted multimodal context data. Accordingly, it is possible to accurately select a point of interest in response to the confirmation request. Furthermore, it is possible to select a point of interest that matches the passenger's characteristics in response to the confirmation request and provide information.

[0377] Meanwhile, the processor (175) can collect information about the selected point of interest and output the collected information. Accordingly, the processor (175) can collect information about the selected point of interest and output the collected information. Accordingly, the point of interest can be accurately selected in response to a confirmation request.

[0378] Meanwhile, the processor (175) within the signal processing device (170) can select a fixed object-based point of interest or a moving object-based point of interest when selecting at least one point of interest. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.

[0379] Meanwhile, fixed objects may include buildings, signboards, roads, etc.

[0380] Meanwhile, moving objects may include vehicles, motorcycles, people, animals, robots, etc.

[0381] Meanwhile, the processor (175) in the signal processing device (170) can vary the point of interest based on the vehicle's driving path or the vehicle's situation information.

[0382] For example, the processor (175) within the signal processing device (170) can control the vehicle's driving path based on road traffic during driving and output point of interest information in response to the driving path of the vehicle being changed.

[0383] As another example, the processor (175) within the signal processing unit (170) can vary the driving path of the vehicle or vary the selected point of interest information based on situational information of the occupant, such as urgent business of the occupant inside the vehicle during driving. Accordingly, the point of interest can be accurately selected based on the situational information of the occupant.

[0384] Meanwhile, the processor (175) in the signal processing device (170) can select a group of points of interest including a plurality of points of interest for vehicle progress.

[0385] Meanwhile, the interest point group may include fixed object-based interest points and moving object-based interest points.

[0386] Specifically, the processor (175) in the signal processing device (170) can combine a first point of interest corresponding to a fixed object and a second point of interest corresponding to a moving object to set a first point of interest group.

[0387] For example, the processor (175) within the signal processing unit (170) can be controlled to output composite point of interest information, such as "There is a cafe in the yellow building where the woman walking with a black umbrella is located," based on first point of interest information corresponding to a yellow building in front of the vehicle and second point of interest information corresponding to a woman walking with a black umbrella on the right side in front of the vehicle. Accordingly, information regarding the points of interest can be accurately output.

[0388] Meanwhile, the interest point group may include road-based interest points and object-based interest points.

[0389] Alternatively, the interest point group may include map data-based interest points and object-based interest points.

[0390] Alternatively, the interest point group may include road-based interest points and object-based interest points.

[0391] Meanwhile, the processor (175) can select a point of interest corresponding to the position of the passenger's gaze at the time when the voice related to the confirmation request is spoken, based on the inference result or response result obtained based on the extracted multimodal context data, and output information about the selected point of interest. Accordingly, it is possible to accurately select a point of interest in response to the confirmation request.

[0392] Specifically, if the confirmation request includes a phrase such as "What's that?", the processor (175) can control to select an external point of interest corresponding to the passenger's gaze position, collect information about the selected point of interest, and output the collected information.

[0393] For example, based on an external object corresponding to a confirmation request and the passenger's gaze position, a first building on the front right of the vehicle is selected, and if the passenger's characteristic data is focused on restaurant data, a first restaurant within the first building can be selected as a point of interest, and information about the selected first restaurant can be collected and provided.

[0394] As another example, based on an external object corresponding to a confirmation request and the passenger's gaze position, a first building on the front right of the vehicle is selected, and if the passenger's characteristic data is focused on hospital data, a first hospital within the first building can be selected as a point of interest, and information regarding the selected first hospital can be collected and provided.

[0395] As another example, based on an external object corresponding to a confirmation request and the passenger's gaze position, a first vehicle on the front left of the vehicle is selected, and the processor (175) selects the first vehicle as a point of interest when the passenger's characteristic data is focused on cost data, collects purchase information regarding price, displacement, etc. related to cost data among the information about the first vehicle, and can provide the collected purchase information.

[0396] That is, the processor (175) can select a moving point of interest, such as the first vehicle, based on a confirmation request and the position of the passenger's gaze, and can control to output information about the selected moving point of interest.

[0397] Meanwhile, when a moving object is selected as a point of interest, the processor (175) may collect information about the moving object through additional analysis or detailed analysis within the acquired image data in addition to collecting information from an external server (400).

[0398] Meanwhile, the processor (175) can select a fixed object as a point of interest, such as the first building, based on a confirmation request and the occupant's line of sight position, and can control to output information about the selected fixed object.

[0399] Meanwhile, the processor (175) can collect information about the selected point of interest through an external server (400) and output the collected information about the selected point of interest as image data or audio data. Accordingly, the point of interest can be accurately selected in response to a confirmation request.

[0400] For example, the processor (175) can control the information about the collected points of interest to be displayed on the display (180).

[0401] As another example, the processor (175) can control the information about the collected point of interest to be output through the audio output unit (185).

[0402] Meanwhile, the processor (175) can detect multiple objects based on camera data from outside the vehicle and recognize the class and attributes of each object. Accordingly, it is possible to output information with improved accuracy in response to a verification request.

[0403] Meanwhile, the processor (175) can recognize the gaze vector information of the occupant based on camera data inside the vehicle. Accordingly, it is possible to output information with improved accuracy in response to a verification request.

[0404] Meanwhile, the processor (175) detects multiple objects based on camera data outside the vehicle, recognizes the gaze vector information of the occupant based on camera data inside the vehicle, and can select an object among the multiple objects that corresponds to the gaze vector information as a point of interest. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.

[0405] Meanwhile, the processor (175) receives location information from the location sensor (198) and can control the synchronization of the location information with the object detected based on camera data outside the vehicle and the occupant's gaze vector information recognized based on camera data inside the vehicle, so that the location information is stored in a database within the multimodal context buffer (752). Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.

[0406] Meanwhile, the processor (175) performs voice recognition based on the passenger's voice signal, extracts context information corresponding to a confirmation request from the database within the multimodal context buffer (752) among the performed voice recognition contents, and can extract context data corresponding to a confirmation request based on the context information. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.

[0407] Meanwhile, the processor (175) can generate a prompt corresponding to the occupant characteristics based on context data including context information, a portion of image data, and map data.

[0408] And, the processor (175) can select an object corresponding to the occupant's gaze position based on the inference result or response result obtained based on the prompt, and output information about the selected object. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.

[0409] Meanwhile, the processor (175) can generate a prompt corresponding to the occupant characteristics based on context data including context information, a portion of image data, and map data.

[0410] And, the processor (175) controls the transmission of the generated prompt to an external server (400), receives an inference result or a response result from the external server (400), selects a point of interest corresponding to the passenger's gaze position based on the received inference result or response result, and outputs information about the selected point of interest. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.

[0411] Meanwhile, the signal processing device (170) according to an embodiment of the present disclosure may further include a neural processor (179).

[0412] Meanwhile, the processor (175) can generate a prompt corresponding to the occupant characteristics based on context data including context information, part of image data, and map data, and transmit the prompt to the neural processor (179).

[0413] The neural processor (179) can output an inference result or a response result based on a prompt from the processor (175).

[0414] Meanwhile, the processor (175) receives an inference result or response result based on a prompt generated by the neural processor (179), selects a point of interest corresponding to the passenger's gaze position based on the received inference result or response result, and outputs information about the selected point of interest. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.

[0415] Meanwhile, the processor (175) can run a hypervisor (505) as in FIG. 5 and run a display virtualization machine (530) for a display on the hypervisor (505).

[0416] Meanwhile, the display virtualization machine (530) can select a point of interest corresponding to the passenger's gaze position based on an inference result or response result obtained based on the extracted context data, and output information about the selected point of interest. Accordingly, it is possible to accurately select a point of interest.

[0417] Meanwhile, the processor (175) may further run a gateway virtualization machine (510) for gateway operation on the hypervisor (505) as shown in FIG. 5.

[0418] Meanwhile, the gateway virtualization machine (510) receives camera data from outside the vehicle, camera data from inside the vehicle, and a voice signal of the passenger, and can transmit the camera data from outside the vehicle, camera data from inside the vehicle, and the voice signal of the passenger to the display virtualization machine (530) through the shared memory (508) within the hypervisor (505).

[0419] Meanwhile, the display virtualization machine (530) can receive camera data outside the vehicle, camera data inside the vehicle, and a voice signal of a passenger through a shared memory (508) within the hypervisor (505), and perform signal processing of the camera data outside the vehicle, camera data inside the vehicle, or the voice signal of a passenger.

[0420] For example, the IVI virtualization machine (530b) within the display virtualization machine (530) can perform signal processing of the passenger's voice signal.

[0421] Specifically, the IVI virtual machine (530b) within the display virtual machine (530) can perform voice recognition based on the passenger's voice signal.

[0422] Meanwhile, the processor (175) can further run a driving control virtualization machine (520) on the hypervisor (505) as shown in FIG. 5.

[0423] Meanwhile, the driving control virtualization machine (520) receives camera data outside the vehicle, camera data inside the vehicle, and a passenger's voice signal through a shared memory (508) within the hypervisor (505), and can execute a driving control service or a driving control application based on the camera data outside the vehicle, camera data inside the vehicle, and the passenger's voice signal. Accordingly, it is possible to provide emotion-based highlight videos.

[0424] FIG. 10 is an example of a flowchart illustrating a method of operation of a signal processing device according to an embodiment of the present disclosure.

[0425] Referring to the drawings, a processor (175) in a signal processing device (170) according to one embodiment of the present disclosure can collect data (S1015).

[0426] Meanwhile, the processor (175) receives camera data from outside the vehicle and camera data from inside the vehicle.

[0427] Camera data from outside the vehicle may include front camera data from the vehicle's front camera, rear camera data from the vehicle's rear camera, left front camera data from the left front camera, or right front camera data from the right front camera.

[0428] Meanwhile, the processor (175) receives vehicle interior sensing data from vehicle interior sensors (710) and can receive vehicle exterior sensing data from vehicle exterior sensors (715). The vehicle interior sensing data may include vehicle interior camera data, audio data received through a microphone, steering wheel, acceleration / deceleration pedal input, and other occupant body signal data.

[0429] External vehicle sensing data may include vehicle location data (GPS), etc.

[0430] Next, the processor (175) can set the mode (S1015).

[0431] The processor (175) can set the mode based on multimodal input data at the start of driving.

[0432] Meanwhile, the processor (175) can generate multimodal context data based on external vehicle sensing data and internal vehicle sensing data.

[0433] And, the processor (175) can set the mode based on multimodal context data.

[0434] Meanwhile, the processor (175) can set a mode based on multimodal context data based on external vehicle sensing data and internal vehicle sensing data, and data stored in a database during vehicle driving.

[0435] For example, the processor (175) can set any one of a destination-based recording mode, a psychological analysis mode, or a health mode.

[0436] Meanwhile, the processor (175) can control the execution of a destination-based recording mode when, during vehicle driving, it is determined to be a destination-based driving such as a trip based on the driver's schedule information, GPS information, navigation information, and conversation information of the driver or passenger.

[0437] Meanwhile, the processor (175) can control the psychological analysis mode to be performed when, during vehicle operation, the psychological state of the driver or passenger continues to be maintained in a specific psychological state based on the driver's face analysis or voice analysis, or when the change in the psychological state exceeds an allowable limit.

[0438] Meanwhile, the processor (175) can analyze the health condition based on the body signal data of the driver or passenger while the vehicle is in motion, and control the health mode to be performed based on the analyzed health condition.

[0439] Next, the processor (175) can evaluate the similarity (S1020).

[0440] Meanwhile, the processor (175) can output a similarity or evaluation index by combining preference-related parameter updates and multi-input data according to mode setting information and personalization information.

[0441] Meanwhile, the processor (175) can analyze the passenger's facial expressions, voice tone, language expressions, etc. from the vehicle interior camera (195i) and microphone (112).

[0442] For example, the processor (175) can perform facial expression analysis, tone, speed, volume, melody pattern analysis, etc. based on a neural network (Convolutional Neural Network; CNN).

[0443] For example, the processor (175) can perform speech recognition based on natural language processing (NLP), etc.

[0444] Meanwhile, the processor (175) can extract emotion keywords based on facial expression analysis, tone, speed, volume, melody pattern analysis, or speech recognition.

[0445] For example, the processor (175) can extract emotion keywords such as happy, sad, anger, fear, surprise, anxiety, joy, melancholy, and consolation, or classify the passenger's emotion data based on facial expression analysis, tone, speed, volume, melody pattern analysis, or speech recognition.

[0446] Meanwhile, the processor (175) can perform landscape recognition from camera data from an external camera based on a Vision Language Model (VLM), etc.

[0447] Meanwhile, the processor (175) can output an evaluation index for the recognized landscape.

[0448] In particular, the processor (175) can output an evaluation index for the perceived scenery based on personalized information based on the passenger's preferences, etc.

[0449] Meanwhile, the processor (175) can return the result class (cause, symptom) to be estimated in each mode and the corresponding similarity through a multimodal artificial intelligence model learned according to the input source required for each mode, such as a destination-based recording mode, a psychological analysis mode, or a health mode.

[0450] Meanwhile, the processor (175) can add information from external data of the user characteristic engine (758) or the recognizer (720) if the system receives or requires information on points of interest or related diseases.

[0451] Next, the processor (175) can detect an event (S1025).

[0452] For example, the processor (175) can filter input data at each unit time.

[0453] Specifically, the processor (175) can filter input data such as duplicate databases, sentiment indices, information on areas of interest, voice, and scenery.

[0454] Meanwhile, the processor (175) can determine whether the input data is meaningful at the current time by comprehensively judging the multimodal input data based on the results inferred according to the determined mode.

[0455] Meanwhile, the processor (175) can record changes in the index for sections where a specific emotional index is high or the range of change is large, sections where the driver missed but thought it meaningful, sections where a specific psychological state is indicated or where the psychological state changes, etc.

[0456] Meanwhile, the processor (175) can actively determine data at a desired time based on a determined mode and record changes in the index for the interval before and after that time.

[0457] Next, the processor (175) can store images and data (S1030).

[0458] Meanwhile, the processor (175) determines the length of the recording intervals before and after the video clip based on the transmitted data, and can set and store the optimal video clip interval length according to the data storage space.

[0459] Next, the processor (175) determines whether a content creation event occurs (S1035), and if applicable, can create content (S1040).

[0460] Meanwhile, the processor (175) can collect necessary or similar clips through current driving and past database retrievers and generate content based on the collected clipped images.

[0461] That is, the processor (175) generates a highlight image including an external image from camera data outside the vehicle in response to classified emotion data, and controls the generated highlight image to be stored in memory (925). Accordingly, it becomes possible to provide an emotion-based highlight image.

[0462] FIG. 11 is an example of an internal block diagram of a signal processing device according to an embodiment of the present disclosure.

[0463] Referring to the drawings, a signal processing system (1100) according to an embodiment of the present disclosure includes an edge sensor group (SN) and a signal processing device (170) in a vehicle.

[0464] The edge sensor group (SN) may include vehicle interior sensors (710) and vehicle exterior sensors (715).

[0465] Meanwhile, the vehicle interior sensors (710) may include an interior camera (195i), a microphone (112), a GPS sensor, a vehicle interior temperature sensor, or a vehicle interior humidity sensor.

[0466] The vehicle external sensors (715) may include an external camera, lidar (196), radar (197), etc.

[0467] Meanwhile, the edge sensor group (SN) may further include a data receiving unit (718) for communication with a mobile terminal (600) or a server (400).

[0468] Meanwhile, the signal processing device (170) according to an embodiment of the present disclosure includes a recognizer (720), an insightor (750), and an illustrator (780).

[0469] Meanwhile, the recognizer (720) can recognize the vehicle situation based on the sensing data received from the edge sensor group (SN).

[0470] The vehicle situation at this time may include external vehicle conditions and internal vehicle conditions.

[0471] Meanwhile, the recognizer (720) can recognize the vehicle situation based on vehicle internal sensing data, vehicle external sensing data, or communication data, and can generate vehicle situation information regarding the recognized vehicle situation.

[0472] To this end, the recognizer (720) may include a detector (1105) and a situation processing unit (740).

[0473] Meanwhile, the internal recognition unit (1105) within the detector (1105) can recognize the situation inside the vehicle based on the vehicle interior sensing data.

[0474] For example, the internal recognizer (1105) can recognize driver gaze information, facial essential voice, emotion information, heart rate variability (HRV) information, driving path information, etc. based on vehicle interior sensing data.

[0475] Meanwhile, the external recognizer (1107) within the detector (1105) can recognize the external situation of the vehicle based on external sensing data.

[0476] Meanwhile, the data recognizer (1108) within the detector (1105) can recognize schedule information, personal characteristic information, health history information, etc. based on communication data.

[0477] The situation government processor (740) can perform comprehensive scene understanding recognition based on vehicle internal situation information from the internal recognizer (1105), vehicle external situation information from the external recognizer (1107), and communication data from the data recognizer (1108).

[0478] To this end, the situation government processor (740) may include a comprehensive scene understanding recognizer (1111) that performs comprehensive scene understanding recognition, and a data collector (1120) that collects or clips data.

[0479] The data collector (1120) may include a data buffer (1121), a content data clip database (1123), a data manager (1124), etc.

[0480] The data manager (1124) can perform duplicate database processing, management of old data, and management of memory (925).

[0481] The Insightor (750) may include a multimodal context engine (755), an AI orchestrator (760), a highlight video module (1140), and a user characteristic database (794).

[0482] The multimodal context engine (755) can generate multimodal context data based on the vehicle situation information received from the recognizer (720).

[0483] Meanwhile, the multimodal context engine (755) can generate multimodal context data based on camera data outside the vehicle, camera data inside the vehicle, and voice signals of the occupant.

[0484] Meanwhile, the multimodal context engine (755) can extract passenger emotion data based on the inference result or response result based on the generated prompt.

[0485] Meanwhile, the multimodal context engine (755) classifies the passenger's emotional data based on camera data inside the vehicle.

[0486] The multimodal context engine (755) may include a recording engine (799) for performing recording modes, etc. Meanwhile, the recording engine (799) may be named Sketch of Memory Engine, as shown in the drawing.

[0487] The AI ​​orchestrator (760) can generate a prompt for classifying the passenger's emotion data based on multimodal context data.

[0488] Meanwhile, the AI ​​orchestrator (760) can obtain inference results or response results based on prompts by using an artificial intelligence model.

[0489] Meanwhile, the AI ​​orchestrator (760) may include a recording agent (797) for calling the recording engine (799) within the multimodal context engine (755). The recording agent (797) may be named a ketch of memory agent, as shown in the drawing.

[0490] Meanwhile, a recording agent (797) within the AI ​​orchestrator (760) can call a recording engine (799) within the multimodal context engine (755) and generate a prompt corresponding to data from the recording engine (799).

[0491] Meanwhile, the recording agent (797) in the AI ​​orchestrator (760) can control the operation of the on-device AI model (767) in the AI ​​model set (765) or the AI ​​model (769) in the server (400) based on the generated prompt.

[0492] For example, the record agent (797) can distribute the load of the on-device AI model (767) in the AI ​​model set (765) or the AI ​​model (769) in the server (400) based on the prompt, and control the on-device AI model (767) or the AI ​​model (769) in the server (400) to operate respectively based on the load distribution.

[0493] Meanwhile, the recording agent (797) can receive an inference result or response result corresponding to a prompt from an on-device AI model (767) or an AI model (769) within the server (400).

[0494] Meanwhile, the recording agent (797) can transmit the inference result or response result based on the prompt to the recording engine (799).

[0495] Meanwhile, the recording agent (797) within the AI ​​orchestrator (760) can send a request to produce a highlight video to the highlight video module (1140) after sending the inference result or response result based on the prompt to the recording engine (799).

[0496] Additionally, the recording agent (797) within the AI ​​orchestrator (760) can transmit driving mode information, judgment information, etc. to the highlight video module (1140).

[0497] Meanwhile, the recording engine (799) can receive inference results or response results based on prompts from the recording agent (797) within the AI ​​orchestrator (760).

[0498] Meanwhile, the recording engine (799) can receive driver point of interest (POI), personal body signal data, etc. from the user characteristic database (794).

[0499] The recording engine (799) within the multimodal context engine (755) may include a mode setting unit (1132), a similarity evaluation unit (1134), an event detection engine (1136), etc.

[0500] The mode setting unit (1132) can set the mode.

[0501] The mode setting unit (1132) can set the mode based on multimodal input data when driving starts.

[0502] Meanwhile, the mode setting unit (1132) can generate multimodal context data based on external vehicle sensing data and internal vehicle sensing data.

[0503] And, the mode setting unit (1132) can set the mode based on multimodal context data.

[0504] Meanwhile, the mode setting unit (1132) can set the mode based on multimodal context data based on external vehicle sensing data and internal vehicle sensing data, and data stored in a database during vehicle driving.

[0505] For example, the processor (175) can set any one of a destination-based recording mode, a psychological analysis mode, or a health mode.

[0506] Meanwhile, the mode setting unit (1132) can control the execution of a destination-based recording mode when, during vehicle driving, it is determined to be a destination-based driving such as travel based on the driver's schedule information, GPS information, navigation information, and conversation information of the driver or passenger.

[0507] Meanwhile, the mode setting unit (1132) can control the psychological analysis mode to be performed when, during vehicle driving, the psychological state of the driver or passenger continues to be maintained in a specific psychological state based on the driver's face analysis or voice analysis, or when the change in the psychological state exceeds an allowable limit.

[0508] Meanwhile, the mode setting unit (1132) can analyze the health condition based on the physical signal data of the driver or passenger while the vehicle is in motion, and control the health mode to be performed based on the analyzed health condition.

[0509] Next, the similarity evaluation unit (1134) can evaluate similarity.

[0510] Meanwhile, the similarity evaluation unit (1134) can output a similarity or evaluation index by combining preference-related parameter updates and multi-input data according to mode setting information and personalization information.

[0511] Meanwhile, the similarity evaluation unit (1134) can analyze the passenger's facial expressions, voice tone, language expressions, etc. from the vehicle interior camera (195i) and microphone (112).

[0512] For example, the similarity evaluation unit (1134) can perform facial expression analysis, tone, speed, volume, melody pattern analysis, etc. based on a neural network (Convolutional Neural Network; CNN).

[0513] For example, the similarity evaluation unit (1134) can perform speech recognition based on natural language processing (NLP), etc.

[0514] Meanwhile, the similarity evaluation unit (1134) can extract emotion keywords based on facial expression analysis, tone, speed, volume, melody pattern analysis, or speech recognition.

[0515] For example, the similarity evaluation unit (1134) can extract emotion keywords such as happy, sad, anger, fear, surprise, anxiety, joy, melancholy, and consolation, or classify the passenger's emotion data based on facial expression analysis, tone, speed, volume, melody pattern analysis, or speech recognition.

[0516] Meanwhile, the similarity evaluation unit (1134) can perform landscape recognition from camera data from an external camera based on a Vision Language Model (VLM), etc.

[0517] Meanwhile, the similarity evaluation unit (1134) can output an evaluation index for the recognized landscape.

[0518] In particular, the similarity evaluation unit (1134) can output an evaluation index for the recognized scenery based on personalized information based on the passenger's preferences, etc.

[0519] Meanwhile, the similarity evaluation unit (1134) can return the result class (cause, symptom) to be estimated in each mode and the corresponding similarity through a multimodal artificial intelligence model learned according to the input source required for each mode, such as a destination-based recording mode, a psychological analysis mode, or a health mode.

[0520] Meanwhile, the similarity evaluation unit (1134) can add information from external data of the user characteristic engine (758) or the recognizer (720) when receiving or requiring information on points of interest or related diseases.

[0521] Next, the event detection engine (1136) can detect events.

[0522] For example, the event detection engine (1136) can filter input data at each unit time.

[0523] Specifically, the event detection engine (1136) can filter input data such as duplicate databases, sentiment indices, information on areas of interest, voice, and landscape.

[0524] Meanwhile, the event detection engine (1136) can determine whether the input data is meaningful at the current time by comprehensively judging the multimodal input data based on the results inferred according to the determined mode.

[0525] Meanwhile, the event detection engine (1136) may send a request to a data collector (1120 in FIG. 12d) to record data of the corresponding section in order to record changes in the index for sections where the specific emotional index is high or the range of change is large, sections where the driver missed but considers meaningful, sections where a specific psychological state is indicated or where the psychological state changes.

[0526] Meanwhile, the event detection engine (1136) can actively determine data at a desired time based on a determined mode and record changes in the index for the interval before and after that time.

[0527] Meanwhile, the highlight video module (1140) may include a content creation judgment module (1142) and a content generator (1144).

[0528] The content production decision module (1142) can make a decision for content production when a request for highlight video production is received from the AI ​​orchestrator (760) or the multimodal context engine (755).

[0529] And, when the content creation judgment module (1142) determines to create content, it can transmit a content creation signal to the content generator (1144).

[0530] Meanwhile, the content generator (1144) generates a highlight image including an external image from camera data outside the vehicle in response to classified emotion data based on a content creation signal.

[0531] For example, the content generator (1144) can generate a highlight video including an external image from camera data outside the vehicle based on an internal artificial intelligence model or an external artificial intelligence model.

[0532] Meanwhile, the memory (925) can store the highlight video generated by the content generator (1144).

[0533] Meanwhile, the content generator (1144) can extract a first external image at a time corresponding to the first emotional data from camera data outside the vehicle when the passenger's emotional data is the first emotional data, and generate a first highlight image including the first external image. Accordingly, an emotion-based highlight image can be provided.

[0534] Meanwhile, the content generator (1144), when the passenger's emotional data is the first emotional data and the second emotional data, can extract a first external image at a time corresponding to the first emotional data from the camera data outside the vehicle and generate a first highlight image including the first external image, and extract a second external image at a time corresponding to the second emotional data from the camera data outside the vehicle and generate a second highlight image including the second external image. Accordingly, it is possible to provide an emotion-based highlight image.

[0535] Meanwhile, the content generator (1144) can generate a separate highlight video for each of the multiple emotional data when the passenger's emotional data consists of multiple emotional data. Accordingly, it is possible to provide an emotion-based highlight video.

[0536] Meanwhile, the content generator (1144) can generate a single highlight video including an external video corresponding to the multiple emotional data when the passenger's emotional data is multiple emotional data. Accordingly, it is possible to provide an emotion-based highlight video.

[0537] Meanwhile, the illustrator (780) receives the highlight video output from the insightor (750) and can add additional information, etc. to the highlight video.

[0538] For example, the illustrator (780) can add camera information that captured an external video to the highlight video.

[0539] As another example, the illustrator (780) can add camera information that captured the internal video to the highlight video.

[0540] As another example, the illustrator (780) can add a map image containing vehicle movement location information to the highlight video.

[0541] As another example, the illustrator (780) can add location information corresponding to an external image to the highlight image.

[0542] As another example, the illustrator (780) can add music or sound corresponding to the emotional data to the highlight video.

[0543] Meanwhile, the multimodal output encoder (781) in the illustrator (780) can encode the received highlight video. Accordingly, the memory (925) can store the encoded highlight video.

[0544] Meanwhile, the analyzer (1152) in the illustrator (780) can generate an analysis report for the highlight video.

[0545] For example, the analyzer (1152) can provide an analysis of the cause of the occurrence of repetitive negative emotions and an emotion timeline.

[0546] Meanwhile, the theme providing unit (1154) within the illustrator (780) can provide a theme corresponding to the highlight video.

[0547] The theme at this time may be music or sound, or may include text or video.

[0548] For example, the theme providing unit (1154) may provide a map image including vehicle movement location information as an example of a theme corresponding to a highlight video.

[0549] As another example, the theme providing unit (1154) can provide location information corresponding to an external image as another example of a theme corresponding to a highlight image.

[0550] As another example, the theme providing unit (1154) can provide music or sound corresponding to an external video as another example of a theme corresponding to a highlight video.

[0551] Meanwhile, the interface (1156) within the illustrator (780) can provide a highlight video to the display (180) based on an event.

[0552] Alternatively, the interface (1156) may provide highlight video to a mobile terminal (600), a server (400), or another vehicle based on an event. Accordingly, emotion-based highlight video can be provided.

[0553] The signal processing device (170) in the system (1100) of Fig. 11 can be controlled to generate and store highlight images that include moments that change according to seasonal changes such as spring, summer, autumn, and winter, or changes in time or weather, even if the path is the same.

[0554] Meanwhile, the signal processing device (170) can generate a highlight video including the journey of the vehicle's driving.

[0555] Meanwhile, the signal processing device (170) can recognize a specific pattern based on psychological events.

[0556] For example, the signal processing device (170) can provide feedback to the occupant or generate analysis data for experts when a repeated stress response situation is captured.

[0557] Meanwhile, the signal processing device (170) can detect the stress repetition interval.

[0558] For example, the signal processing device (170) can detect trauma-like reactions, including cases where repetitive tension reactions and negative negative emotions occur in specific intervals.

[0559] Alternatively, the signal processing device (170) can match an emotional response similar to a previous highlight video.

[0560] Meanwhile, the signal processing device (170) can detect an emotional reversal pattern.

[0561] Meanwhile, the signal processing device (170) can analyze the time from negative emotion to positive reaction and perform a resilience evaluation based on this.

[0562] Meanwhile, the signal processing device (170) can provide a psychological analysis report.

[0563] For example, the signal processing device (170) may provide a daily or weekly summary report of emotions and reactions, or provide reports such as a stress index, an emotional diversity index, and recovery time.

[0564] Meanwhile, the signal processing device (170) can transmit this report to an external server (400) to control it so that it is linked with an expert analysis API or used for remote psychological counseling.

[0565] Meanwhile, the signal processing device (170) can generate a highlight video by collecting only moments of specific emotions (e.g., joy or positivity) after the trip ends.

[0566] Meanwhile, the signal processing device (170) can analyze the cause of the occurrence of repeated negative emotions.

[0567] Meanwhile, the signal processing device (170) can generate and store an entire emotional timeline on a daily or weekly basis.

[0568] Meanwhile, the signal processing device (170) can monitor emotional changes of passengers, such as children, the elderly, and patients.

[0569] Meanwhile, the signal processing device (170) can track the emotional state of the passenger and provide relevant information or highlight video to the mobile terminal (600) or server (400) for self-recognition or use of healing tools.

[0570] Meanwhile, the signal processing device (170) can generate a personalized highlight image or provide a user interface (UI) or user experience (UX) of the vehicle display (180) based on the highlight image.

[0571] Meanwhile, the signal processing device (170) can provide highlighter images for digital healthcare applications, etc.

[0572] Meanwhile, the signal processing device (170) can track the burnout of the passenger or monitor changes in emotion during the commute to and from work.

[0573] Meanwhile, the signal processing device (170) can track the emotions of the passenger, monitor stress such as academics, or monitor changes in emotions during family conversations.

[0574] Meanwhile, the signal processing device (170) can detect depression early or analyze the reduction of facial expression or prolonged blank expression state.

[0575] Meanwhile, the signal processing device (170) can continuously detect similar reactions after a specific event to manage post-traumatic stress disorder (PTSD), etc.

[0576] FIGS. 12a to 26 are drawings referenced in the description of FIG. 10 or FIG. 11.

[0577] Fig. 12a is a drawing referenced in the mode setting.

[0578] Referring to the drawing, the processor (175) can set a mode based on the vehicle's internal sensor data.

[0579] For example, the processor (175) can perform modeling through an embedding model (1211) based on voice, driving path, schedule information, vehicle data information or passenger information, etc., which are examples of vehicle internal sensor data as in (a) of FIG. 12a, and can set a mode through text embeddings (1213) and a vector database (1215).

[0580] Specifically, the processor (175) can control the execution of a destination-based recording mode when, during vehicle driving, it is determined to be a destination-based driving mode such as a trip based on the driver's schedule information, GPS information, navigation information, and conversation information of the driver or passenger.

[0581] As another example, the processor (175) can perform modeling through an embedding model (1211) based on other examples of vehicle interior sensor data, such as voice, emotion, driver gaze information, etc., as in (b) of FIG. 12a, and set a mode through text embeddings (1213) and a vector database (1215).

[0582] Specifically, the processor (175) can control the performance of a psychological analysis mode when, during vehicle operation, the psychological state of the driver or passenger continues to be maintained in a specific psychological state based on the driver's face analysis or voice analysis, or when the change in the psychological state exceeds an allowable limit.

[0583] As another example, the processor (175) can perform modeling through an embedding model (1211) based on another example of vehicle internal sensor data, such as body signals, heart rate variability (HRV) information, health information, health history information, etc., as in (c) of FIG. 12a, and set a mode through text embeddings (1213) and a vector database (1215).

[0584] Specifically, the processor (175) can analyze the health condition based on the body signal data of the driver or passenger while the vehicle is in motion, and control the health mode to be performed based on the analyzed health condition.

[0585] Figure 12b is a drawing referenced for similarity evaluation.

[0586] Referring to the drawing, the processor (175) can receive input data from an input source required for each mode, such as a destination-based recording mode, a psychological analysis mode, or a health mode.

[0587] For example, the processor (175) can receive input data such as vehicle internal sensor data, vehicle external sensor data, and personal storage data as a video signal (1221), a text signal (1223), or a voice signal (1224).

[0588] And, the processor (175) can model various input data through a multimodal embedding model (1211) and set or return similarity by mode through a vector database (1215).

[0589] The multimodal embedding model (1211) at this time may include contrastive language-image pre-training (CLIP), etc.

[0590] Meanwhile, the processor (175) may add information from external data of the user characteristic engine (758) or recognizer (720) based on the passenger's preference or point of interest (POI) or disease information.

[0591] Fig. 12c is a diagram illustrating various scenarios.

[0592] Referring to the drawing, the processor (175) can output a similarity or class based on emotional data corresponding to external scenery, etc. when performing a recording mode as in (a) of FIG. 12c.

[0593] Meanwhile, when performing a recording mode, the processor (175) can recognize vehicle interior situation information based on driver gaze information, facial essential voice, emotion information, heart rate variability (HRV) information, driving path information, vehicle operation data or health information, etc., which are examples of vehicle interior sensing data.

[0594] When performing a recording mode, the processor (175) can recognize external vehicle situation information based on external vehicle camera data, sound, or vehicle event information, which are examples of external vehicle sensing data.

[0595] The processor (175) can recognize passenger personal information based on schedule information, personal characteristic information, or health history information, which are examples of communication data, when performing a recording mode.

[0596] Meanwhile, when performing a recording mode, the processor (175) can output a similarity or class corresponding to the emotion data based on the vehicle interior situation information, vehicle exterior situation information, passenger personal information, and preferred object or landscape information stored in the personal characteristic database (794).

[0597] Meanwhile, the processor (175) can classify appraisal data regarding mountains, seas, rivers, skies, fields, cultural properties, etc., when performing a recording mode, and output a similarity or class (1231) corresponding to the appraisal data.

[0598] Meanwhile, the processor (175) can recognize that a specific psychological state (e.g., anxiety, sadness, anger, surprise, disgust, joy, etc.) is continuously maintained after boarding when performing the psychological analysis mode as in (b) of FIG. 12c.

[0599] And, the processor (175) can output a similarity or class based on changes in emotional data, etc., due to specific external factors during the performance of the psychological analysis mode.

[0600] Meanwhile, the processor (175) can recognize vehicle interior situation information based on driver gaze information, facial essential voice, emotion information, heart rate variability (HRV) information, driving path information, vehicle motion data or health information, etc., which are examples of vehicle interior sensing data when performing a psychological analysis mode.

[0601] When performing a psychological analysis mode, the processor (175) can recognize external vehicle situation information based on external vehicle camera data, sound, or vehicle event information, which are examples of external vehicle sensing data.

[0602] The processor (175) can recognize passenger personal information based on schedule information, personal characteristic information, or health history information, which are examples of communication data, when performing psychological analysis mode.

[0603] Meanwhile, the processor (175) can output a similarity or class corresponding to the emotional data based on the vehicle interior situation information, vehicle exterior situation information, and passenger personal information when performing the psychological analysis mode.

[0604] Meanwhile, the processor (175) can output a similarity or class (1232) corresponding to the psychological state result or the cause of the psychological change when performing the psychological analysis mode.

[0605] The processor (175) can output a class or similarity including abnormal symptom symptoms and the corresponding causes by comparing past data when performing health mode as in (c) of FIG. 12c.

[0606] For example, the processor (175) can output a class or similarity of stress, depression, etc. based on physical signals when performing health mode.

[0607] Meanwhile, the processor (175) can recognize vehicle interior situation information based on driver gaze information, facial essential voice, emotion information, heart rate variability (HRV) information, driving path information, vehicle motion data, or health information, which are examples of vehicle interior sensing data when performing health mode.

[0608] The processor (175) can recognize external vehicle situation information based on external vehicle camera data, sound, or vehicle event information, which are examples of external vehicle sensing data, when performing health mode.

[0609] The processor (175) can recognize passenger personal information based on schedule information, personal characteristic information, or health history information, which are examples of communication data, when performing health mode.

[0610] Meanwhile, the processor (175) can output a similarity or class corresponding to emotional data or symptoms based on vehicle interior situation information, vehicle exterior situation information, passenger personal information and health-related information stored in the personal characteristic database (794) when performing health mode.

[0611] Meanwhile, the processor (175) can output a similarity or class (1233) corresponding to signs of health abnormalities or causes of health changes when performing health mode.

[0612] Fig. 12d is a drawing referenced for the generation of a highlight image.

[0613] Referring to the drawing, the processor (175) may include a data collector (1120) for generating a highlight image.

[0614] Meanwhile, the data collector (1120) may include a data buffer (1121) and a data manager (1129).

[0615] For example, the data collector (1120) collects vehicle exterior camera data or vehicle interior camera data.

[0616] In addition, the data collector (1120) can collect audio data, body signal data, GPS data, artificial intelligence-based response result data or inference result data, etc.

[0617] Next, the data buffer (1121) stores data collected from the data collector (1120).

[0618] For example, the data buffer (1121) can store image data (IMGm) and other data (DTm) collected from the data collector (1120).

[0619] Meanwhile, the data buffer (1121) stores data (IMGm, DTm) in real time in volatile memory (not shown), and when it receives a request to store the current data from the Event Detection Engine, it can store the memory dump in the actual storage space.

[0620] For example, the data buffer (1121) can store camera data, audio data, body signal data, GPS data, artificial intelligence-based response result data or inference result data, etc. when a request to store current data is received.

[0621] The data manager (1129), based on a request to generate a highlight video, stores content clips (CCD1) of a predetermined periodic interval in the data buffer (1121) for a data buffer (1121). <CCD2,CCD3,CCDm) 중 적어도 하나의 컨텐츠 클립(CCDm)을 선택하고, 추출할 수 있다.

[0622] Meanwhile, the data manager (1129) can perform duplicate data management.

[0623] For example, the data manager (1129) can monitor the data storage space in consideration of the limited space constraints of the vehicle storage device and, if necessary, delete duplicate data or data that has been stored for a long time. Accordingly, the data manager (1129) can perform optimization of the storage space.

[0624] Meanwhile, the processor (175) can optimize the index based on the time distribution of class or similarity for each setting mode when setting a data interval of a predetermined periodic interval.

[0625] The figure illustrates an example of an index graph (GRPa) in which levels change over time.

[0626] For example, in recording mode, the processor (175) may exemplify the time-based index distribution of multimodal-based passenger emotion results as shown in the index graph (GRPa) of the drawing.

[0627] Meanwhile, the processor (175) can clip the minimum and maximum values ​​within the index graph (GRPa) by limiting them in recording mode.

[0628] As another example, the processor (175) may exemplify a time-based index distribution of passenger emotion results based on multimodal in a psychological analysis mode.

[0629] Meanwhile, the processor (175) can clip the minimum and maximum values ​​within the index distribution by limiting them in the psychological analysis mode.

[0630] As another example, the processor (175) may exemplify a time-based index distribution of multimodal-based passenger emotion results in health mode.

[0631] Meanwhile, the processor (175) can clip the minimum and maximum values ​​within the index distribution by limiting them in health mode.

[0632] FIG. 12e is a drawing referenced in the description of FIG. 12d. In particular, FIG. 12e illustrates various examples of other data (DTm) stored together with the image data (IMGm) of FIG. 12d.

[0633] Referring to the drawing, the data buffer (1121) can store basic information data (DTm) that can be stored separately for each frame together with image data (IMGm).

[0634] In this case, the basic information data (DTm) may include metadata.

[0635] Meanwhile, buffer data (IMGm or DTm) stored in the data buffer (1121) can be used to return analysis results at specific intervals through the embedding inference model (1137), mode setting unit (1132), or similarity evaluation unit (1134), etc.

[0636] Meanwhile, the event detection engine (1136) can determine that the scene group is meaningful data.

[0637] Meanwhile, the data collector (1120) can store data (IMGm or DTm) related to the corresponding scene group and class and similarity in the data buffer (1121) based on the event detected by the event detection engine (1136).

[0638] In the drawing, as an example of the corresponding scene group, data at a first point in time (1241a), data at a second point in time after the first point in time (1241b), and data at a third point in time after the second point in time (1241c) are respectively exemplified.

[0639] In contrast, other examples of a different scene group may include data at a first time point (1241a), data at a second time point prior to the first time point (1241b), and data at a third time point prior to the second time point (1241c).

[0640] Meanwhile, the processor (175) determines the length of the recording intervals before and after the video clip based on the transmitted data, and can set and store the length of the video clip intervals based on the data storage space.

[0641] For example, the processor (175) can control the generation and storage of a first video clip (1241a) corresponding to the first emotion data.

[0642] As another example, the processor (175) can be controlled to generate and store a second video clip (1241b) corresponding to the second emotion data.

[0643] As another example, the processor (175) can control to generate and store a third video clip (1241c) corresponding to the third emotion data.

[0644] Figure 12f is a drawing referenced in the generation of a highlight image.

[0645] Referring to the drawing, the processor (175) may include a data collector (1120) for collecting data, a content generator (1140) for generating content such as highlight video, and a recording engine (799) for performing recording mode, etc.

[0646] The data collector (1120), content generator (1140), and recording engine (799) can correspond to the data collector (1120), content generator (1140), and recording engine (799) of FIG. 11.

[0647] Meanwhile, the data collector (1120) may include a data buffer (1121) that stores video data, text, or audio data, and a data clip database (1123).

[0648] Meanwhile, the recording engine (799) within the processor (175) may include an embedding inference model (1137) that performs embedding or inference on input data, a retrieval-augmented generation module (RAG) (1139) that performs retrieval-augmented generation, a mode setting evaluation unit (1131) that evaluates mode setting or similarity, and an event detection engine (1136) that detects events.

[0649] Meanwhile, the event detection engine (1136) can send a data collection request to the data collector (1120) when an event is detected.

[0650] Meanwhile, the mode setting evaluation unit (1131) can transmit past modes and past data to the search augmentation generation module (1139).

[0651] Meanwhile, the content generator (1140) generates a highlight video including an external image from camera data outside the vehicle in response to the classified emotion data. Accordingly, it is possible to provide an emotion-based highlight video.

[0652] FIGS. 13a to 13h illustrate various examples of highlight images.

[0653] Figure 13a illustrates an example of a highlight video.

[0654] Referring to the drawings, a signal processing device (170) according to one embodiment of the present disclosure includes a memory (925) and a processor (175) that receives camera data from outside the vehicle and camera data from inside the vehicle.

[0655] Meanwhile, the processor (175) classifies the passenger's emotional data based on camera data inside the vehicle, generates a highlight image including an external image from camera data outside the vehicle corresponding to the classified emotional data, and controls the storage of the generated highlight image in memory (925). Accordingly, it is possible to provide an emotion-based highlight image.

[0656] Meanwhile, the processor (175) can generate a map image (1312) containing vehicle movement location information (1313, 1314, 1316) as shown in the drawing, and a highlight image (1310) containing an external image (1311). Accordingly, it is possible to provide an emotion-based highlight image.

[0657] Meanwhile, the processor (175) can generate a highlight video (1310) that further includes music or sound information (1318) corresponding to the classified emotion data. Accordingly, it is possible to provide an emotion-based highlight video, music, or sound.

[0658] Figure 13b illustrates another example of a highlight image.

[0659] Referring to the drawings, a processor (175) according to one embodiment of the present disclosure can generate a highlight image (1320) including an interior image from camera data inside a vehicle in response to classified emotion data.

[0660] At this time, the processor (175) can control the inclusion of an image containing camera information (1321) of the vehicle interior within the highlight image (1320). Accordingly, the camera position within the highlight image can be intuitively identified.

[0661] Fig. 13c illustrates another example of a highlight video.

[0662] Referring to the drawings, a processor (175) according to one embodiment of the present disclosure can generate a highlight image (1330) including an external image (1332) from camera data outside the vehicle and an internal image (1334) from camera data inside the vehicle, corresponding to classified emotion data.

[0663] At this time, the processor (175) can control the inclusion of an image containing camera information (1331) of the exterior of the vehicle within the highlight image (1330). Accordingly, the camera position within the highlight image can be intuitively identified.

[0664] Fig. 13d illustrates another example of a highlight image.

[0665] Referring to the drawing, the processor (175) can generate a highlight image (1340) including an external image (1432) corresponding to classified emotion data and camera information (1431) that captured the external image. Accordingly, an emotion-based highlight image can be provided.

[0666] Meanwhile, the processor (175) can generate a highlight image (1340) including location information (1343) corresponding to an external image and an external image (1432) corresponding to classified emotion data. Accordingly, an emotion-based highlight image can be provided.

[0667] Fig. 13e illustrates another example of a highlight video.

[0668] Referring to the drawing, the processor (175) can generate a highlight image (1350) including a plurality of external images (1352, 1354, 1356, 1357) corresponding to classified emotion data and camera information (1351) captured by the external images. Accordingly, an emotion-based highlight image can be provided.

[0669] Fig. 13f illustrates another example of a highlight video.

[0670] Referring to the drawings, the processor (175) can generate a highlight image (1360) including a plurality of external images (1352, 1364, 1366, 1367) corresponding to classified emotion data, an internal image (1368), and camera information (1361) captured by the external images. Accordingly, an emotion-based highlight image can be provided.

[0671] Fig. 13g illustrates another example of a highlight video.

[0672] Referring to the drawing, the processor (175) can generate a highlight image (1370) including an external image (1372) corresponding to classified emotion data, camera information (1371) that captured the external image, and location information (1373) corresponding to the external image. Accordingly, an emotion-based highlight image can be provided.

[0673] Fig. 13h illustrates another example of a highlight video.

[0674] Referring to the drawings, the processor (175) can generate a map image (1384) containing vehicle movement location information (1381, 1382), an external image (1382) corresponding to classified emotion data, and a highlight image (1380) containing camera information (1381) of the external image. Accordingly, an emotion-based highlight image can be provided.

[0675] Combining FIGS. 13a to 13h, the processor (175) can extract a first external image at a time corresponding to the first emotional data from camera data outside the vehicle when the passenger's emotional data is the first emotional data, and generate a first highlight image including the first external image. Accordingly, an emotion-based highlight image can be provided.

[0676] Meanwhile, the processor (175), when the passenger's emotional data is the first emotional data and the second emotional data, can extract a first external image at a time corresponding to the first emotional data from the camera data outside the vehicle and generate a first highlight image including the first external image, and extract a second external image at a time corresponding to the second emotional data from the camera data outside the vehicle and generate a second highlight image including the second external image. Accordingly, it is possible to provide an emotion-based highlight image.

[0677] Meanwhile, the processor (175) can generate a separate highlight video for each of the multiple emotional data when the passenger's emotional data consists of multiple emotional data. Accordingly, it is possible to provide an emotion-based highlight video.

[0678] Meanwhile, the processor (175) can generate a single highlight video including an external video corresponding to the multiple emotional data when the passenger's emotional data is multiple emotional data. Accordingly, an emotion-based highlight video can be provided.

[0679] Meanwhile, the processor (175) can receive more voice signals of the passenger, extract facial expression information of the passenger from camera data inside the vehicle, and classify the passenger's emotion data based on the facial expression information and the passenger's voice signals.

[0680] Meanwhile, the processor (175) can extract a first external image at a time corresponding to the first emotional data during driving on a first driving path at a first time point, and generate a first highlight image including the first external image.

[0681] Meanwhile, the processor (175) can extract a second external image at a time corresponding to the second emotion data and generate a second highlight image including the second external image when driving on a first driving path at a second time point different from the first time point. Accordingly, different emotion-based highlight images can be provided even on the same driving path.

[0682] Meanwhile, the processor (175) can play a highlight video based on an event or transmit it to a mobile terminal (600) or another vehicle.

[0683] For example, the processor (175) can generate a highlight video while the vehicle is in motion and play the highlight video based on an event where the vehicle is in motion.

[0684] As another example, the processor (175) can generate a highlight video while driving a vehicle along a first path, and when an event occurs in which the same first path driving information is received from a mobile terminal (600) or another vehicle, the generated highlight video can be transmitted to the mobile terminal (600) or another vehicle. Accordingly, the emotion-based highlight video can be shared with other mobile terminals, etc.

[0685] As another example, the processor (175) can play a highlight video when an event occurs where the same path is traveled at a different time. Accordingly, it is possible to provide an emotion-based highlight video.

[0686] Meanwhile, the processor (175) can detect the position of the occupant's gaze based on camera data inside the vehicle.

[0687] And, the processor (175) can extract an external image containing a point of interest corresponding to the passenger's gaze position and generate a highlight image containing the extracted external image. Accordingly, a highlight image containing a point of interest can be provided based on emotion.

[0688] Meanwhile, the processor (175) detects the gaze position of the occupant based on camera data inside the vehicle, performs voice recognition based on the occupant's voice signal, and if the performed voice recognition content includes a confirmation request such as "What is that?", extracts an external image including a point of interest corresponding to the gaze position of the occupant at the time of the confirmation request utterance, and generates a highlight image including the extracted external image. Accordingly, it is possible to provide a highlight image including a point of interest based on emotion.

[0689] Meanwhile, a signal processing device (170) according to one embodiment of the present disclosure may further include a neural processor (179).

[0690] Meanwhile, the processor (175) can generate multimodal context data based on camera data outside the vehicle, camera data inside the vehicle, and voice signals of the occupant, and generate a prompt for classifying the occupant's emotion data based on the multimodal context data.

[0691] Meanwhile, the processor (175) receives an inference result or response result based on a prompt generated from the neural processor (179), and can extract the passenger's emotion data based on the inference result or response result based on the generated prompt. Accordingly, it is possible to provide an emotion-based highlight video.

[0692] Meanwhile, the processor (175) can run a hypervisor (505) as in FIG. 5 and run a display virtualization machine (530) for a display on the hypervisor (505).

[0693] Meanwhile, the display virtualization machine (530) can extract emotion data of the occupant based on the inference result or response result based on the generated prompt, and generate a highlight video including an external image from camera data outside the vehicle based on the extracted emotion data. Accordingly, it is possible to provide an emotion-based highlight video.

[0694] Meanwhile, the processor (175) may further run a gateway virtualization machine (510) for gateway operation on the hypervisor (505) as shown in FIG. 5.

[0695] Meanwhile, the gateway virtualization machine (510) receives camera data from outside the vehicle, camera data from inside the vehicle, and a voice signal of the passenger, and can transmit the camera data from outside the vehicle, camera data from inside the vehicle, and the voice signal of the passenger to the display virtualization machine (530) through the shared memory (508) within the hypervisor (505).

[0696] Meanwhile, the display virtualization machine (530) can receive camera data outside the vehicle, camera data inside the vehicle, and a voice signal of a passenger through a shared memory (508) within the hypervisor (505), and perform signal processing of the camera data outside the vehicle, camera data inside the vehicle, or the voice signal of a passenger.

[0697] For example, the IVI virtualization machine (530b) within the display virtualization machine (530) can perform signal processing of the passenger's voice signal.

[0698] Specifically, the IVI virtual machine (530b) within the display virtual machine (530) can perform voice recognition based on the passenger's voice signal.

[0699] A signal processing device (170) according to another embodiment of the present disclosure includes a memory (925), camera data from outside the vehicle, camera data from inside the vehicle, and a processor (175) for receiving a voice signal of a passenger.

[0700] Meanwhile, the processor (175) classifies the passenger's emotional data based on the passenger's voice signal and camera data inside the vehicle, generates a highlight image including an external image from camera data outside the vehicle corresponding to the classified emotional data, and controls the storage of the generated highlight image in memory (925). Accordingly, an emotion-based highlight image can be provided.

[0701] FIG. 14a illustrates setting the search range based on the occupant's line of sight.

[0702] Referring to the drawing, the processor (175) can identify an object placed within a predetermined distance or predetermined range based on the passenger's line of sight when a confirmation request such as "What's that?" is received, and provide an Application Programming Interface (API) for executing an AI model (767 or 769) based on the identified object.

[0703] The drawing illustrates that the angle between the vehicle driving direction (DRm) and the first direction reference is θα, and the driver's (1020) line of sight angle is θβ based on the in-vehicle image from the in-vehicle camera (195i) inside the vehicle (200).

[0704] The processor (175) can calculate a gaze vector (DRe) based on the angle (θα) between the vehicle driving direction (DRm) and the first direction reference and the gaze angle (θβ) of the driver (1020).

[0705] And, the processor (175) can set a search range (FOV) corresponding to a predetermined angle (θm) based on the view vector (DRe).

[0706] For example, the processor (175) can control the field of view (FOV) to become smaller as the vehicle's speed increases. Accordingly, the field of view can be adaptedly adjusted according to the vehicle's speed.

[0707] In the drawing, the field of view (FOV) is illustrated as being cone-shaped relative to the driver (1020).

[0708] Meanwhile, the processor (175) can extract multiple objects (OBJma, OBJmb) within a predetermined search distance (RAS).

[0709] Meanwhile, the processor (175) can control the search distance (RAS) to increase as the vehicle speed increases. Accordingly, the search distance (RAS) can be adaptively adjusted according to the vehicle speed.

[0710] Meanwhile, the processor (175) can extract multiple map tiles within the search range (FOV).

[0711] Meanwhile, the processor (175) can calculate the distance between the location of a driver (1020), which is an example of a passenger, and a plurality of objects (OBJma, OBJmb).

[0712] Meanwhile, the processor (175) can set at least one of the plurality of objects (OBJma, OBJmb) as a point of interest and collect name information, address information, distance information, etc. of the point of interest.

[0713] At this time, the processor (175) can collect information through image signal processing for multiple objects (OBJma, OBJmb) or collect information through an external server (400).

[0714] FIG. 14b is a diagram illustrating the search distance (RAS) within the map data (1717).

[0715] Referring to the drawing, meanwhile, the processor (175) can detect an object within a predetermined search distance (RAS) of the map data (1717).

[0716] Meanwhile, the processor (175) controls the search distance (RAS) to increase as the vehicle speed increases and to decrease as the vehicle speed decreases, thereby enabling the search distance (RAS) to be adaptively adjusted according to the vehicle speed.

[0717] FIG. 14c is a diagram illustrating multiple map tiles within a field of view (FOV).

[0718] Referring to the drawing, the processor (175) can set a plurality of map tiles (720) within the search range (FOV).

[0719] In the drawing, the shape of each map tile is square, but unlike this, various shapes such as circles, hexagons, and pentagons are possible.

[0720] And, the processor (175) can set a search range (FOV) corresponding to a predetermined angle (θm) based on the view vector (VTm).

[0721] That is, the processor (175) can set the field of view (FOV) between the first line (LNa) and the second line (Lnb) based on the line of view vector (VTm). Meanwhile, the angle between the first line (LNa) and the second line (Lnb) may be θm.

[0722] Meanwhile, the processor (175) can detect an object or point of interest within the search range (FOV) and within the search distance (RAS).

[0723] In the drawing, there are six objects or points of interest within the field of view (FOV) and within the search distance (RAS).

[0724] Meanwhile, the processor (175) identifies multiple map tiles that overlap with the area when the gaze search area of ​​the driver (1020), which is an example of a passenger, is specified.

[0725] For example, if the map tile at the default zoom level of the procedural modeling of the 3D map contains data attributes of a point of interest, the processor (175) can use that zoom level as is and reuse the map tile data loaded into memory (140) for modeling. Accordingly, performance can be optimized.

[0726] Meanwhile, the processor (175) can perform additional separate exploration based on the changed zoom level if the 3D landmark zoom level is different from the exploration zoom level of the point of interest.

[0727] Meanwhile, the processor (175) optimizes performance by reusing 3D landmarks when they are already being used for modeling the 3D map or are already loaded into memory (140).

[0728] Meanwhile, the processor (175) can select at least one of the objects as a point of interest based on the occupant's characteristic data when multiple objects are located in the occupant's line of sight. This is described with reference to FIG. 14d.

[0729] Fig. 14d illustrates selecting an object hidden from the line of sight as a point of interest.

[0730] Referring to the drawing, the processor (175) can detect the line of sight of a driver (1020), which is an example of a passenger, and detect a plurality of objects (BDma, BDmb, BDmc) placed on the line of sight.

[0731] The processor (175) can detect a first object (BDma) based on front image data from the front camera (195).

[0732] Meanwhile, the processor (175) can detect a second object (BDmb) and a third object (BDmc) that are hidden and positioned behind the first object (BDma) based on map data stored in memory (140) and location information.

[0733] And, the processor (175) detects the line of sight of a driver (1020), which is an example of a passenger, and among a plurality of objects (BDma, BDmb, BDmc) placed on the line of sight, it can select at least one of the plurality of objects as a point of interest based on characteristic data of the driver (1020), which is an example of a passenger.

[0734] In the drawing, the processor (175) is shown selecting a second object (BDmb) among a plurality of objects (BDma, BDmb, BDmc) as the first point of interest (PTma) and selecting a third object (BDmc) as the second point of interest (PTmb). Accordingly, the point of interest can be accurately selected.

[0735] In particular, when a confirmation request such as "What's that?" is received, the processor (175) can select a second object (BDmb) among a plurality of objects (BDma, BDmb, BDmc) as a first point of interest (PTma) and select a third object (BDmc) as a second point of interest (PTmb) based on the line of sight of a driver (1020), which is an example of a passenger. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.

[0736] Meanwhile, the processor (175) can perform a test to determine whether multiple objects (BDma, BDmb, BDmc) are obscured.

[0737] For example, the processor (175) can determine whether each object is obscured by the height difference of each object within the line of sight of a driver (1020), which is an example of a passenger, based on the height information of each object (BDma, BDmb, BDmc).

[0738] The processor (175) can determine that the second object (BDmb) is obscured when the height of the first object (BDma) that is closer, as shown in the drawing, is greater than the height of the second object (BDmb) that is further away.

[0739] Meanwhile, the processor (175) determines whether the first object (BDma) is matched based on the occupant's characteristic data, and if it is not matched, determines whether the second object (BDmb) is matched, and if it is matched, can select it as the first point of interest (PTma).

[0740] Similarly, the processor (175) can determine whether the third object (BDmc) is a match based on the occupant's characteristic data, and if it is a match, select it as the second point of interest (PTmb).

[0741] Meanwhile, the processor (175) can select a fixed object placed behind the moving object as a point of interest based on map data when the moving object is located in the line of sight of the occupant. This is described with reference to FIG. 14e.

[0742] FIG. 14e illustrates a moving object positioned in the occupant's line of sight.

[0743] Referring to the drawing, the processor (175) can perform object detection based on the front camera image (1730) and detect the occupant's line of sight (VTm) based on the interior camera (195i).

[0744] In the drawing, the angle between the occupant's line of sight (VTm) and the vehicle's direction of travel (Drm) is θβ.

[0745] Meanwhile, the processor (175) can detect a moving object (TRa) based on the front camera image (1730).

[0746] Meanwhile, the processor (175) can select a fixed object placed behind the moving object as a point of interest based on map data when the moving object (TRa) is located in the line of sight of the occupant.

[0747] FIG. 14f illustrates selecting a fixed object behind the moving object (TRa) of FIG. 14e as the point of interest.

[0748] Referring to the drawing, the processor (175) can obtain an inference result or a response result based on multimodal context data including the moving object when the moving object is located in the occupant's line of sight (VTm), and can select a point of interest based on the obtained inference result or response result.

[0749] In particular, when a confirmation request such as "What's that?" is received, the processor (175) detects an object based on the line of sight (VTm) of a driver (1020), which is an example of a passenger, and when a moving object (TRa) is detected, it can select a fixed object placed behind the moving object (TRa) as a point of interest based on map data in memory (140).

[0750] In the drawing, an example is provided of selecting the building (BDmb) behind the moving object (TRa) as the first point of interest (PTma), and then selecting the next building (BDmc) as the second point of interest (PTmb).

[0751] Meanwhile, the processor (175) may ultimately select a second point of interest (PTmb) rather than a first point of interest (PTma) based on the passenger's characteristic data.

[0752] Meanwhile, the processor (175) can collect information about a selected point of interest from an external server (400) and output the collected information.

[0753] Meanwhile, the processor (175) can select the fixed object (BDmc) as a point of interest (PTmb) when the fixed object (BDmc) is located in the passenger's line of sight (VTm), collect information about the point of interest (PTmb) based on map data stored in memory (140), and output the collected information. Accordingly, the point of interest can be accurately selected in response to a confirmation request.

[0754] Meanwhile, the processor (175) can select the fixed object (BDmc) as the point of interest (PTmb) when the fixed object (BDmc) is located in the passenger's line of sight (VTm), and if there is no information about the point of interest (PTmb) in the map data stored in memory (140), collect information about the selected point of interest (PTmb) from an external server (400) and output the collected information. Accordingly, the point of interest can be accurately selected in response to a confirmation request.

[0755] Fig. 15a illustrates an example of a vehicle front image.

[0756] Referring to the drawing, the processor (175) can detect an object in the vehicle front image (1735) based on the occupant's line of sight (VTm).

[0757] For example, the processor (175) can detect a first object (1738) based on a passenger's line of sight (VTm) that forms a predetermined angle (θβ) with the vehicle's driving direction (Drm).

[0758] Meanwhile, the processor (175) may also detect a second object (1736) corresponding to the vehicle driving direction (Drm).

[0759] Meanwhile, the processor (175) can perform segmentation for object detection.

[0760] FIG. 15b illustrates a segmented image (1740) in which segmentation has been performed on the vehicle front image (1735) of FIG. 15a.

[0761] Referring to the drawing, the processor (175) can perform segmentation on the vehicle front image (1735) to obtain a segmentation image (1740).

[0762] Meanwhile, the segmentation image (1740) may include a first segmentation object (1748) corresponding to the first object (1738) and a second segmentation object (1746) corresponding to the second object (1736).

[0763] FIG. 15c is a diagram illustrating separate signal processing for moving objects or stationary objects.

[0764] Referring to the drawing, the processor (175) can determine whether the object the occupant is looking at is the first object (1738) or the building behind the first object (1738) based on the first segmentation object (1748) of FIG. 15b.

[0765] For example, the processor (175) can output a question message to determine whether the object the occupant is looking at is the first object (1738) or the building behind the first object (1738) based on the first segmentation object (1748) of FIG. 15b, and can determine whether the object the occupant is looking at is the first object (1738) or the building behind the first object (1738) based on the response message of the occupant (1020) corresponding to the question message.

[0766] Meanwhile, the processor (175) can control the execution of a multimodal LLM (765) for the analysis of a response message when the object being watched by the occupant is a moving object such as a vehicle, a first object (1738), or a person.

[0767] That is, the processor (175) can control the execution of an AI model set (765) for the analysis of a response message when the object being watched by the occupant is a moving object such as a vehicle, a first object (1738), or a person.

[0768] And, the processor (175) can select the first object (1738) as the object that the passenger is watching, based on the inference result or response result from the AI ​​model set (765). Accordingly, the point of interest can be accurately selected.

[0769] As another example, the processor (175) can check whether a point of interest registration (1745) is set in the map data to check whether the object being watched by the occupant is a building behind the first object (1738).

[0770] For example, if the processor (175) has a point of interest (1745) registered in the map data for a building behind the first object (1738), the processor (175) can select the building behind the first object (1738), rather than the first object (1738), as the object that the occupant is watching. Accordingly, the point of interest can be accurately selected.

[0771] FIG. 16a illustrates the location of multiple points of interest within a building.

[0772] Referring to the drawing, when multiple points of interest (PTna, PTnb, PTnc, PTnd) are located corresponding to the gaze position of the occupant (1020), the processor (175) can select at least one point of interest among the multiple points of interest (PTna, PTnb, PTnc, PTnd) based on the occupant's characteristic data, collect information about the selected point of interest, and output the collected information. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.

[0773] For example, the processor (175) can select at least one point of interest (PTna, PTnb, PTnc, PTnd) among the points of interest (PTna, PTnb, PTnc, PTnd) based on the occupant's characteristic data, when the building object (BDna) is located on the occupant's (1020) line of sight (VTn) and multiple points of interest (PTna, PTnb, PTnc, PTnd) are located within the building object (BDna) in the map data.

[0774] For example, if the passenger characteristic data is focused on a cafe, the processor (175) can select a cafe (PTnd) within a building object (BDna) as a point of interest and collect and provide information about the selected cafe (PTnd). Accordingly, it is possible to select a point of interest that matches the passenger characteristics and provide information.

[0775] As another example, the processor (175) can select a hospital (PTnb) within a building object (BDna) as a point of interest when the passenger characteristic data is focused on a hospital, and collect and provide information about the selected hospital (PTnb). Accordingly, it is possible to select a point of interest that matches the passenger characteristics and provide information.

[0776] Meanwhile, the processor (175) can adjust the prompt to inquire which of the multiple points of interest to provide information through the AI ​​model set (765) when multiple points of interest are identified based on the passenger's characteristic data or passenger's preference data. Accordingly, it is possible to select a point of interest that matches the passenger's characteristics and provide information.

[0777] FIG. 16b illustrates the identification of one of the multiple points of interest within the building of FIG. 16a.

[0778] Referring to the drawing, the processor (175) can detect multiple objects (1752, 1753, 1754) among the vehicle front image (1750).

[0779] Meanwhile, the processor (175) can select a building object (1754) based on a passenger's gaze vector (VTn) that forms a predetermined angle (θβ) with the vehicle's direction of travel (DRn).

[0780] And, the processor (175) can crop a portion of the image (1755) of the building object (1754) corresponding to the line of sight vector (VTn) to obtain a cropped image (1760).

[0781] And, the processor (175) can recognize text or brand logos within an image (1760) through a multimodal AO agent (1777) or an AI model set (765), and perform secondary filtering (1774) based on the recognized text or brand logos (1772).

[0782] Accordingly, the processor (175) is able to distinguish multiple points of interest (PTna, PTnb, PTnc, PTnd) within the building object (1754).

[0783] FIG. 17a illustrates selecting a moving object as a point of interest.

[0784] Referring to the drawing, the processor (175) can obtain a vehicle front image (1820) from a vehicle camera (195mA, 195mb, 195mc).

[0785] Meanwhile, the processor (175) can detect an object placed within a predetermined distance or predetermined range based on the passenger's line of sight when a confirmation request (1810) such as "What's that?" is received.

[0786] For example, the processor (175) can detect that the passenger's gaze is located on the vehicle ahead (1825) within the vehicle front image (1820).

[0787] Figure 17b is a drawing referenced in the description of Figure 17a.

[0788] Referring to the drawing, the processor (175) receives a voice signal corresponding to the passenger's utterance (1842), such as "What's that?"

[0789] Next, the processor (175) can extract the occupant's line of sight (1845) from the internal camera (195i).

[0790] And, the processor (175) can extract the passenger's gaze information (1843) based on the passenger's gaze angle (1845).

[0791] Next, the processor (175) can generate a gaze image or gaze information (1852) based on the vehicle front image (1850) and gaze information (1843) or gaze angle (1845).

[0792] The processor (175) can crop or extract the vehicle front image (1850) into a specific range based on the gaze image or gaze information (1852) to generate the gaze image or gaze information (1854) of the occupant.

[0793] Next, the processor (175) can control the execution of an object classification / recognition AI model (1856) based on the passenger's gaze image or gaze information (1854).

[0794] The processor (175) determines (1864) whether the identified object is a dynamic point of interest or a static point of interest based on the result of the object identification AI model (1856), and if it is a dynamic point of interest, controls the execution of the AI ​​model set (765).

[0795] At this time, the AI ​​model set (765) can be executed based on the gaze object identification information generated from the object identification AI model (1856).

[0796] And, the processor (175) can generate speech-based descriptive text data (1869) for dynamic points of interest based on the inference results or response results of the AI ​​model set (765).

[0797] Meanwhile, the processor (175) can, based on the result of the object identification AI model (1856), search for point of interest information (1867) within the map data if the identified object is a static point of interest, and control the execution of the AI ​​model set (765) based on the searched point of interest information (1867).

[0798] And, the processor (175) can generate colloquial-based descriptive text data (1869) for static points of interest based on the inference results or response results of the AI ​​model set (765).

[0799] FIG. 17c illustrates selecting a fixed object as a point of interest.

[0800] Referring to the drawing, the processor (175) can obtain an image of the vehicle exterior from the vehicle camera (195mA, 195mb, 195mc).

[0801] For example, the processor (175) can acquire an image of the vehicle's exterior based on a occupant's gaze vector (VTna) that forms a predetermined angle (θn) with the vehicle's direction of travel (DRn).

[0802] Meanwhile, the processor (175) can select at least one of the plurality of vehicle cameras (195ma, 195mb, 195mc) based on the angle of view information (190) of the plurality of vehicle cameras (195ma, 195mb, 195mc) and the occupant's gaze vector (VTna), and acquire an image (1912) from the selected camera.

[0803] Meanwhile, the processor (175) can calculate a gaze point based on the horizontal angle of the passenger's gaze and the vertical angle of the passenger's gaze.

[0804] Meanwhile, the processor (175) can perform semantic segmentation based on a specific size window or AI model based on the gaze point.

[0805] Meanwhile, the processor (175) can calculate the coordinates required for Image Cropping (1914) based on the coordinate values ​​of the result of a specific size or semantic segmentation, and perform Image Cropping (1914) based on the calculated coordinates.

[0806] That is, the processor (175) can generate a cropped image (1916) based on the calculated coordinates and calculate the average depth (1920) of the image.

[0807] FIG. 17d is a drawing referenced in the description of FIG. 17c.

[0808] Referring to the drawing, the occupant's line of sight (VTnc) may be lower than the horizon (LVN).

[0809] That is, the occupant's line of sight (VTnc) can be lower than the horizon (LVN) by θn.

[0810] The processor (175) can select a vehicle object (1945) within a camera image (1940) as a point of interest (PTnc) in response to the passenger's gaze.

[0811] FIG. 18a illustrates an example of a method for identifying objects at a dynamic point of interest.

[0812] Referring to the drawing, the processor (175) can control the execution of an object identification artificial intelligence model (765) based on a cropped image (1940) corresponding to the passenger's line of sight among the vehicle exterior images.

[0813] The cropped image (1940) at this time may be an image including a vehicle.

[0814] And, the processor (175) can obtain inference result data (1942) including type, category, etc. as inference result data of the object identification artificial intelligence model (765).

[0815] FIG. 18b illustrates an example of a method for identifying objects at a static point of interest.

[0816] Referring to the drawing, the processor (175) can control the execution of an object identification artificial intelligence model (765) based on a cropped image (1950) corresponding to the passenger's line of sight among the vehicle exterior images.

[0817] The cropped image (1950) at this time may be an image of a sign inside a building.

[0818] And, the processor (175) can obtain inference result data (1952) including type, category, etc. as inference result data of the object identification artificial intelligence model (765).

[0819] FIG. 19 is a drawing referenced in the explanation of guide point information for vehicle movement.

[0820] Referring to the drawing, the processor (175) in the signal processing unit (170) can recognize the road being driven (DRm) or the intersection (CDm) connected to the road being driven (DRm) based on location data and camera data among the sensing data outside the vehicle.

[0821] Meanwhile, the processor (175) in the signal processing device (170) can control the execution of a navigation service based on the destination data of the vehicle among the sensing data inside the vehicle.

[0822] Meanwhile, the processor (175) in the signal processing device (170) can receive turn-by-turn (TBT) information for each section of the road (DRm) or intersection (CDm) being driven on, based on map data, during the execution of the navigation service.

[0823] For example, the turn-by-turn information for each section may include angle information based on a first direction. The first direction may correspond to the true north direction or the vehicle driving direction.

[0824] Meanwhile, the processor (175) in the signal processing device (170) can calculate the first exit angle (θa) from the road (DRm) to the first branch road (BRa) based on map data, location data, and camera data during the execution of the navigation service.

[0825] At this time, the first exit angle (θa) may be the angle between the road (DRm) currently in motion and the first branch road (BRa).

[0826] Meanwhile, the angle between the second branch road (BRb) and the first branch road (BRa) of the intersection (CDm) may be θb as the second exit angle.

[0827] Meanwhile, the processor (175) within the signal processing device (170) can extract information on multiple points of interest (POI) or buildings or landmarks existing within the search distance (RAS) from map tiles within the map data.

[0828] Meanwhile, the processor (175) in the signal processing device (170) can generate multimodal context data including map data, sensing data inside the vehicle, and sensing data outside the vehicle, and generate a prompt based on the multimodal context data.

[0829] And, the processor (175) in the signal processing unit (170) outputs guide point information for vehicle progress based on the inference result or response result obtained based on the prompt.

[0830] At this time, meanwhile, the processor (175) in the signal processing device (170) can generate a plurality of guide points, select at least one guide point among the plurality of guide points, and output guide information corresponding to the selected guide point.

[0831] For example, a processor (175) within a signal processing device (170) may select information of a point of interest or a building or landmark closest to the first branch road (BRa) or the first exit angle (θa) of the first branch road (BRa) as guide information.

[0832] Specifically, the processor (175) in the signal processing device (170) can select the first branch road (BRa) or the first building (BDma) located within the first exit angle (θa) of the first branch road (BRa) as guide information.

[0833] And, the processor (175) in the signal processing unit (170) can be controlled to output guide information such as "turn right around the first building (BDma)."

[0834] In this way, instead of providing guidance information that is not intuitive, such as the conventional "turn right 200m ahead," guidance information such as "turn right around Building 1 (BDma)" is output, thereby enabling the accurate selection of a point of interest in response to a confirmation request. Furthermore, in response to a confirmation request, a point of interest that matches the characteristics of the occupant can be selected and information provided.

[0835] As another example, the processor (175) in the signal processing unit (170) can select the first branch road (BRa) or the second building (BDmb) closest to the first branch road (BRa) as guide information.

[0836] And, the processor (175) within the signal processing unit (170) can be controlled to output guide information such as "turn right when the second building (BDmb) is visible." Accordingly, the point of interest can be accurately selected in response to a confirmation request.

[0837] Figure 20a illustrates an example of map data.

[0838] Referring to the drawing, the processor (175) in the signal processing unit (170) can receive map data (1010) while the vehicle is in motion.

[0839] In particular, the processor (175) within the signal processing device (170) can receive map data (1010) stored in memory (140) based on location data, which is an example of sensing data outside the vehicle, while the vehicle is in motion.

[0840] FIG. 20b illustrates an example of context data based on the map data of FIG. 20a.

[0841] Referring to the drawing, the processor (175) in the signal processing device (170) can generate navigation context data (1020) for an AI model based on the map data (1010) of FIG. 20a.

[0842] Context data (1020) may include, as shown in the drawing, a driving road, a building, a building location (1012), a building direction (1014), an original message based on map data, etc.

[0843] FIG. 20c exemplifies a natural language phrase based on the context data of FIG. 20b.

[0844] Referring to the drawing, the processor (175) in the signal processing unit (170) can generate a natural language phrase (1032 or 1035 or 1037) for an AI model based on the context data (1020) of FIG. 20b.

[0845] A processor (175) within a signal processing unit (170) can execute an AI model (767) based on natural language phrases (1032, 1035, 1037) and generate guide point information (1033 or 1036) for vehicle progress based on the execution result of the AI ​​model (767). Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.

[0846] Figure 21 illustrates the generation of voice-based guide information based on camera data.

[0847] Referring to the drawing, the processor (175) in the signal processing device (170) can receive sensing data inside the vehicle and sensing data outside the vehicle.

[0848] In particular, the processor (175) within the signal processing device (170) can receive camera image data (1105), which is an example of sensing data outside the vehicle, from a plurality of cameras (106mA, 195mb, 195mc).

[0849] Meanwhile, the processor (175) can receive turn-by-turn (TBT) information (1109) for each section of the road (DRm) or intersection (CDm) being driven on, based on location data (e.g., GPS data), which is an example of sensing data from outside the vehicle, and map data.

[0850] Meanwhile, the processor (175) in the signal processing device (170) can generate context data (1110) for an AI model based on segment-by-segment turn-by-turn information (1109) and camera image data (1105).

[0851] And, the processor (175) can generate dynamic points of interest (1112) based on context data (1110).

[0852] That is, the processor (175) can generate dynamic points of interest (1112) based on camera image data (1105), location data, and map data.

[0853] Meanwhile, the processor (175) in the signal processing device (170) can perform a point of interest search (1114) based on the turn-by-turn information (1109) for each section.

[0854] And, the processor (175) can generate a static point of interest (1115) based on the point of interest search (1114).

[0855] That is, the processor (175) can generate a static point of interest (1115) based on location data and map data.

[0856] Meanwhile, the dynamic point of interest (1112) may include vehicles, motorcycles, people, animals, robots, etc. That is, the dynamic point of interest (1112) may correspond to the moving objects described above.

[0857] Meanwhile, the static point of interest (1115) may include buildings, signboards, roads, etc. That is, the static point of interest (1115) may correspond to the fixed object described above.

[0858] Next, the processor (175) can generate a natural language-based phrase or prompt based on the dynamic point of interest (1112) and the static point of interest (1115), and control the execution of the AI ​​model (1122) based on the natural language-based phrase or prompt.

[0859] The AI ​​model (1122) at this time may be the above-described on-device AI model (767) or the AI ​​model (769) within the server (400).

[0860] Meanwhile, the processor (175) can generate a 3D navigation graphic effect (1130) based on a dynamic point of interest (1112) and a static point of interest (1115).

[0861] Next, the processor (175) can output text-based guide point information (1124) based on the inference result or response result according to the execution of the AI ​​model (1122).

[0862] And, the processor (175) can convert text-based guide point information (1124) into voice-based guide point information (1126) and output it for a driver who is driving.

[0863] Accordingly, sound corresponding to voice-based guide point information (1126) is output through the audio output unit (185) in the vehicle (200). Accordingly, it is possible to accurately select a point of interest in response to a confirmation request. Furthermore, in response to a confirmation request, it is possible to select a point of interest that matches the characteristics of the occupant and provide information.

[0864] FIG. 22 illustrates an example of voice-based guide point information of FIG. 21.

[0865] Referring to the drawing, the processor (175) can control the navigation screen (1210) to be displayed on the display when the navigation service is executed.

[0866] And, the processor (175) can select a gas station (1216) within the navigation screen (1210) as a guide point (1217) and output voice-based audio guide information (1225) corresponding to the selected guide point (1217).

[0867] That is, as shown in the drawing, natural language-based audio guide information (1225) such as "Go straight, and when you see the sparkling gas station over there, turn right and take Yangcheon-ro" can be output.

[0868] Accordingly, it becomes possible to accurately select points of interest in response to confirmation requests. Furthermore, in response to confirmation requests, it becomes possible to select points of interest that match passenger characteristics and provide information.

[0869] FIG. 23 illustrates the generation of dynamic object-based guide information based on camera data.

[0870] Referring to the drawing, the processor (175) in the signal processing device (170) can receive sensing data inside the vehicle and sensing data outside the vehicle.

[0871] In particular, the processor (175) within the signal processing device (170) can receive camera image data (IMGa), which is an example of sensing data outside the vehicle, from a plurality of cameras (106mA, 195mb, 195mc).

[0872] Meanwhile, the processor (175) can receive turn-by-turn (TBT) information (1107) for each section of the road being driven on, based on location data (e.g., GPS data), which is an example of sensing data from outside the vehicle, and map data.

[0873] Meanwhile, the processor (175) in the signal processing device (170) can generate context data (1110) for an AI model based on segment-by-segment turn-by-turn information (1107) and camera image data (IMGa).

[0874] And, the processor (175) can extract dynamic points of interest or moving objects based on context data (1110).

[0875] For example, dynamic points of interest or moving objects may include people, vehicles ahead, etc., within the camera image (IMGa).

[0876] The processor (175) can select a guide point based on a dynamic point of interest or a moving object. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.

[0877] For example, the processor (175) can generate context data (1310) based on a dynamic point of interest or a moving object, generate a prompt based on the context data, and, when executing a navigation service, select a guide point based on the inference result obtained based on the prompt. Accordingly, the point of interest can be accurately selected in response to a confirmation request.

[0878] FIGS. 24a to 24d illustrate the selection of guide points based on multiple conditions.

[0879] FIG. 24a illustrates a road (DRm) in motion or an intersection (CDm) connected to a road (DRm) in motion, as in FIG. 19.

[0880] Referring to the drawing, the processor (175) in the signal processing unit (170) can recognize the road being driven (DRm) or the intersection (CDm) connected to the road being driven (DRm) based on location data and camera data among the sensing data outside the vehicle.

[0881] Meanwhile, the processor (175) in the signal processing device (170) can receive turn-by-turn (TBT) information for each section of the road (DRm) or intersection (CDm) being driven on, based on map data, during the execution of the navigation service.

[0882] For example, the turn-by-turn information for each section may include angle information based on a first direction. The first direction may correspond to the true north direction or the vehicle driving direction.

[0883] Meanwhile, the processor (175) in the signal processing device (170) can calculate the first exit angle (θa) from the road (DRm) to the first branch road (BRa) based on map data, location data, and camera data during the execution of the navigation service.

[0884] At this time, the first exit angle (θa) may be the angle between the road (DRm) currently in motion and the first branch road (BRa).

[0885] Meanwhile, the angle between the second branch road (BRb) and the first branch road (BRa) of the intersection (CDm) may be θb as the second exit angle.

[0886] Meanwhile, the processor (175) in the signal processing device (170) can set an arc-shaped search distance (RAS) based on the first exit angle, the second exit angle, the latitude and longitude of the turn-by-turn point, and the search distance.

[0887] Meanwhile, the processor (175) in the signal processing device (170) can search for multiple points of interest, buildings, or landmarks after setting the search distance (RAS).

[0888] For example, the processor (175) in the signal processing device (170) can generate at least one building or landmark that is a plurality of points of interest searched based on the first exit angle as a first condition, as shown in FIG. 24b, as a guide point.

[0889] Next, the processor (175) in the signal processing device (170) can generate at least one building or landmark that is a plurality of points of interest searched based on the second exit angle as a second condition, as shown in FIG. 24c, when the first condition is not met, as a guide point.

[0890] Next, the processor (175) in the signal processing device (170) can create a guide point as a blank space as a third condition, as shown in FIG. 24d, when the first and second conditions are not met.

[0891] FIG. 25 is an example of the creation of guide points in a rotation section.

[0892] Referring to the drawing, the guide information generation unit (1120) in the processor (175) in the signal processing unit (170) can receive a plurality of context data (1512, 1514) based on camera data.

[0893] Meanwhile, the processor (175) can generate a static point of interest (1516) based on a search for a point of interest based on camera data.

[0894] That is, the processor (175) can generate a static point of interest (1516) based on location data and map data.

[0895] Meanwhile, the static point of interest (1516) may include buildings, signboards, roads, etc. That is, the static point of interest (1516) may correspond to the fixed object described above.

[0896] Next, the processor (175) can generate a natural language-based phrase or prompt based on a static point of interest (1516) and control the execution of an AI model (1122) based on the natural language-based phrase or prompt.

[0897] The AI ​​model (1122) at this time may be the above-described on-device AI model (767) or the AI ​​model (769) within the server (400).

[0898] Meanwhile, the processor (175) can generate 3D navigation graphic effects (1130) based on static points of interest (1516).

[0899] Next, the processor (175) can output text-based guide point information (1124) based on the inference result or response result (1518) resulting from the execution of the AI ​​model (1122). Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request. Furthermore, it becomes possible to select a point of interest that matches the passenger characteristics and provide information in response to a confirmation request.

[0900] Figure 26 is another example of the creation of guide points in a rotation section.

[0901] Referring to the drawing, the guide information generation unit (1120) in the processor (175) in the signal processing unit (170) can receive a plurality of context data (1612, 1614) based on camera data.

[0902] Meanwhile, the processor (175) can generate dynamic points of interest (1616) based on a search for points of interest based on camera data.

[0903] That is, the processor (175) can generate dynamic points of interest (1616) based on camera data.

[0904] Meanwhile, the dynamic point of interest (1616) may include vehicles, motorcycles, people, animals, robots, etc. That is, the dynamic point of interest (1616) may correspond to the moving objects described above.

[0905] Next, the processor (175) can generate a natural language-based phrase or prompt based on a dynamic point of interest (1616) and control the execution of an AI model (1122) based on the natural language-based phrase or prompt.

[0906] The AI ​​model (1122) at this time may be the above-described on-device AI model (767) or the AI ​​model (769) within the server (400).

[0907] Meanwhile, the processor (175) can generate 3D navigation graphic effects (1130) based on dynamic points of interest (1616).

[0908] At this time, the processor (175) can generate text-based 3D navigation graphic effects (1130) based on a 3D model (1132) and a 3D generation AI model (1134).

[0909] Next, the processor (175) can output text-based guide point information (1124) based on the inference result or response result (1518) resulting from the execution of the AI ​​model (1122). Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request. Furthermore, it becomes possible to select a point of interest that matches the passenger characteristics and provide information in response to a confirmation request.

[0910] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. Various modifications are possible by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.

Claims

1. Memory; A processor that receives camera data from outside the vehicle and camera data from inside the vehicle; The above processor is, Classifying passenger emotion data based on camera data inside the vehicle, and Corresponding to the above-described emotion data, a highlight image including an external image from camera data outside the vehicle is generated, and A signal processing device that controls the storage of the generated highlight image in the memory.

2. In Paragraph 1, The above processor is, If the emotional data of the above-mentioned passenger is the first emotional data, A signal processing device that extracts a first external image at a time corresponding to the first emotional data from camera data outside the vehicle and generates a first highlight image including the first external image.

3. In Paragraph 1, The above processor is, If the emotional data of the above-mentioned passenger is the first emotional data and the second emotional data, A first external image at a point in time corresponding to the first emotional data is extracted from the camera data outside the vehicle, and a first highlight image including the first external image is generated. A signal processing device that extracts a second external image at a time corresponding to the second emotional data from camera data outside the vehicle and generates a second highlight image including the second external image.

4. In Paragraph 1, The above processor is, A signal processing device that generates a highlight image for each of the multiple emotional data when the emotional data of the passenger is multiple emotional data.

5. In Paragraph 1, The above processor is, A signal processing device that generates a single highlight image including an external image corresponding to the plurality of emotional data when the emotional data of the above-mentioned passenger is a plurality of emotional data.

6. In Paragraph 1, The above processor is, A signal processing device that generates the highlight image including the external image and camera information captured by the external image.

7. In Paragraph 1, The above processor is, A signal processing device that generates a map image including vehicle movement location information and a highlight image including the external image.

8. In Paragraph 1, The above processor is, A signal processing device that generates a highlight image including position information corresponding to the external image and the external image.

9. In Paragraph 1, The above processor is, A signal processing device that generates a highlight image including an external image from camera data outside the vehicle and an internal image from camera data inside the vehicle, corresponding to the above-described classified emotion data.

10. In Paragraph 1, The above processor is, Receive more of the passenger's voice signal, A signal processing device that extracts facial expression information of a passenger from camera data inside the vehicle and classifies the passenger's emotion data based on the facial expression information and the passenger's voice signal.

11. In Paragraph 1, The above processor is, When driving on a first driving path at a first time point, a first external image at a time point corresponding to first emotion data is extracted, and a first highlight image including the first external image is generated. A signal processing device that, when driving on the first driving path at a second time point different from the first time point, extracts a second external image at a time point corresponding to the second emotional data and generates a second highlight image including the second external image.

12. In Paragraph 1, The above processor is, A signal processing device that plays the highlight video or transmits it to a mobile terminal or another vehicle based on an event.

13. In Paragraph 1, The above processor is, Based on camera data from outside the vehicle, camera data from inside the vehicle, and the passenger's voice signal, multimodal context data is generated, and Based on the above multimodal context data, a prompt for classifying the passenger's emotion data is generated, and A signal processing device that extracts emotion data of the occupant based on an inference result or response result based on the generated prompt.

14. In Paragraph 1, The above processor is, Based on camera data inside the vehicle, the position of the occupant's gaze is detected, and Extract the external image including a point of interest corresponding to the gaze position of the occupant, and A signal processing device that generates the highlight image including the extracted external image.

15. In Paragraph 1, The above processor is, Based on camera data inside the vehicle, the position of the occupant's gaze is detected, and Voice recognition is performed based on the voice signal of the passenger, and if a confirmation request is included in the content of the voice recognition performed, the external image including a point of interest corresponding to the passenger's gaze position at the time of the utterance of the confirmation request is extracted. A signal processing device that generates the highlight image including the extracted external image.

16. In Paragraph 12, It further includes a neural processor; and The above processor is, Based on camera data from outside the vehicle, camera data from inside the vehicle, and the passenger's voice signal, multimodal context data is generated, and Based on the above multimodal context data, a prompt for classifying the passenger's emotion data is generated, and Receive the inference result or response result based on the generated prompt from the above neural processor, and A signal processing device that extracts emotion data of the occupant based on an inference result or response result based on the generated prompt.

17. In Paragraph 1, The above processor is, Run a hypervisor, and on the hypervisor, run a display virtualization machine for a display, and The above display virtualization machine is, A signal processing device that extracts emotion data of the occupant based on an inference result or response result based on the generated prompt, and generates a highlight image including an external image from camera data outside the vehicle based on the extracted emotion data.

18. In Paragraph 17, The above processor is, On the above hypervisor, a gateway virtualization machine for gateway operation is further run, and The above gateway virtual machine is, Receiving camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of the occupant, A signal processing device that transmits camera data outside the vehicle, camera data inside the vehicle, and a passenger's voice signal to the display virtualization machine through shared memory within the hypervisor.

19. Memory; A processor that receives camera data from outside the vehicle, camera data from inside the vehicle, and a voice signal from the occupant; The above processor is, Classifying the passenger's emotion data based on the passenger's voice signal and the camera data inside the vehicle, and Corresponding to the above-described emotion data, a highlight image including an external image from camera data outside the vehicle is generated, and A signal processing device that controls the storage of the generated highlight image in the memory.

20. A vehicle control device comprising a signal processing device according to any one of claims 1 to 19.