Signal processing device and vehicle control device including same
The signal processing device uses camera and voice recognition to select and provide passenger-relevant points of interest based on gaze detection and multimodal context, addressing the limitations of existing systems by ensuring information aligns with passenger interests.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2025-11-05
- Publication Date
- 2026-05-15
AI Technical Summary
Existing vehicle control systems struggle to accurately select and provide information about points of interest relevant to the passenger during driving, as they typically only display pre-registered information and fail to account for the passenger's actual interests or characteristics.
A signal processing device that utilizes camera data, voice recognition, and multimodal context data to detect the passenger's gaze position and select points of interest based on inference results, considering passenger characteristics and line of sight, with the ability to collect and output relevant information from internal and external sources.
Enables accurate selection and provision of information about passenger-relevant points of interest in response to confirmation requests, enhancing the passenger experience by aligning displayed content with their actual interests and needs.
Smart Images

Figure KR2025017988_15052026_PF_FP_ABST
Abstract
Description
Signal processing device, and vehicle control device having the same
[0001] The present disclosure relates to a signal processing device and a vehicle control device equipped with the same, and more specifically, to a signal processing device capable of accurately selecting a point of interest in response to a confirmation request and a vehicle control device equipped with the same.
[0002] A vehicle is a device that moves the user in the desired direction. A typical example is an automobile.
[0003] Meanwhile, for the convenience of users, a vehicle control device is installed inside the vehicle.
[0004] The vehicle control unit includes a signal processing unit and can perform signal processing based on sensor data from various internal sensor devices.
[0005] Meanwhile, the signal processing device can output information related to the passenger's point of interest during vehicle driving based on signal processing.
[0006] On the other hand, when outputting information related to points of interest, only information regarding pre-registered or set points of interest is displayed, so there is a disadvantage in that it is difficult to provide information about points of interest that the passenger is interested in during actual driving.
[0007] The problem that the present disclosure aims to solve is to provide a signal processing device capable of accurately selecting a point of interest in response to a confirmation request, and a vehicle control device equipped with the same.
[0008] Another problem that the present disclosure aims to solve is to provide a signal processing device capable of selecting a point of interest corresponding to occupant characteristics and providing information in response to a confirmation request, and a vehicle control device equipped with the same.
[0009] To solve the above technical problem, a signal processing device according to one embodiment of the present disclosure and a vehicle control device equipped therewith include a processor that receives camera data outside the vehicle, camera data inside the vehicle, and a voice signal of a passenger. The processor performs object detection based on camera data outside the vehicle, detects the passenger's gaze position based on camera data inside the vehicle, performs voice recognition based on the passenger's voice signal, and if a confirmation request is included in the performed voice recognition content, extracts multimodal context data corresponding to the confirmation request, and selects a point of interest corresponding to the passenger's gaze position based on an inference result or a response result obtained based on the extracted multimodal context data.
[0010] Meanwhile, the processor can collect information about the selected point of interest and output the collected information.
[0011] Meanwhile, the processor can select a point of interest corresponding to the passenger's gaze position at the time when the voice related to the confirmation request is spoken, based on an inference result or response result obtained based on the extracted multimodal context data, and output information about the selected point of interest.
[0012] Meanwhile, the processor can recognize the occupant's gaze vector information based on camera data inside the vehicle.
[0013] Meanwhile, the processor detects multiple objects based on camera data outside the vehicle, recognizes gaze vector information of the occupant based on camera data inside the vehicle, and can select an object among the multiple objects that corresponds to the gaze vector information as a point of interest.
[0014] Meanwhile, when multiple objects are located in the occupant's line of sight, the processor can select at least one of the multiple objects as a point of interest based on the occupant's characteristic data.
[0015] Meanwhile, the processor can select a fixed object placed behind the moving object as a point of interest based on map data when a moving object is located in the occupant's line of sight.
[0016] Meanwhile, when a moving object is located in the occupant's line of sight, the processor can obtain an inference result or a response result based on multimodal context data including the moving object, and select a point of interest based on the obtained inference result or response result.
[0017] Meanwhile, the processor can collect information about a selected point of interest from an external server and output the collected information.
[0018] Meanwhile, if a fixed object is located in the occupant's line of sight, the processor selects the fixed object as a point of interest, collects information about the point of interest based on map data stored in memory, and outputs the collected information.
[0019] Meanwhile, the processor selects a fixed object as a point of interest when a fixed object is located in the occupant's line of sight, and if there is no information about the point of interest in the map data stored in memory, it collects information about the selected point of interest from an external server and outputs the collected information.
[0020] Meanwhile, when multiple points of interest are located corresponding to the passenger's gaze position, the processor can select at least one point of interest among the multiple points of interest based on the passenger's characteristic data, collect information about the selected point of interest, and output the collected information.
[0021] Meanwhile, the processor performs voice recognition based on the passenger's voice signal, extracts context information corresponding to a confirmation request from a database among the performed voice recognition content, and can extract context data corresponding to a confirmation request based on the context information.
[0022] Meanwhile, the processor can generate a prompt corresponding to a passenger characteristic based on context data including context information, a portion of image data, and map data, select a point of interest corresponding to the passenger's gaze position based on an inference result or response result obtained based on the prompt, and output information about the selected point of interest.
[0023] Meanwhile, the processor can generate a prompt corresponding to passenger characteristics based on context data including context information, a portion of image data, and map data, control the transmission of the generated prompt to an external server, receive an inference result or a response result from the external server, select a point of interest corresponding to the passenger's gaze position based on the received inference result or response result, and output information about the selected point of interest.
[0024] A signal processing device according to one embodiment of the present disclosure and a vehicle control device equipped therewith further include a neural processor, and the processor generates a prompt corresponding to a passenger characteristic based on context data including context information, a portion of image data, and map data, receives an inference result or response result based on the prompt generated from the neural processor, selects a point of interest corresponding to the passenger's gaze position based on the received inference result or response result, and outputs information about the selected point of interest.
[0025] Meanwhile, the processor executes a hypervisor and, on the hypervisor, executes a display virtualization machine for a display, and the display virtualization machine selects a point of interest corresponding to the occupant's gaze position based on an inference result or response result obtained based on extracted context data, and can output information about the selected point of interest.
[0026] Meanwhile, the processor further runs a gateway virtualization machine for gateway operation on the hypervisor, and the gateway virtualization machine receives camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of the occupant, and can transmit the camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of the occupant to a display virtualization machine through shared memory within the hypervisor.
[0027] Meanwhile, the processor further executes a driving control virtualization machine on the hypervisor, and the driving control virtualization machine receives camera data outside the vehicle, camera data inside the vehicle, and voice signals of the occupant through shared memory within the hypervisor, and can execute a driving control service or a driving control application based on the camera data outside the vehicle, camera data inside the vehicle, and voice signals of the occupant.
[0028] A signal processing device according to one embodiment of the present disclosure and a vehicle control device equipped therewith include a processor that receives camera data from outside the vehicle, camera data from inside the vehicle, and a voice signal of a passenger. The processor performs object detection based on camera data from outside the vehicle, detects the passenger's gaze position based on camera data from inside the vehicle, performs voice recognition based on the passenger's voice signal, and if a confirmation request is included in the performed voice recognition content, extracts multimodal context data corresponding to the confirmation request, and selects a point of interest corresponding to the passenger's gaze position based on an inference result or a response result obtained based on the extracted multimodal context data. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request. Furthermore, it is possible to select a point of interest that matches the passenger characteristics in response to a confirmation request and provide information.
[0029] Meanwhile, the processor can collect information about the selected point of interest and output the collected information. Accordingly, it becomes possible to accurately select the point of interest in response to a confirmation request.
[0030] Meanwhile, the processor can select a point of interest corresponding to the passenger's gaze position at the time the voice related to the confirmation request is uttered, based on an inference result or response result obtained based on the extracted multimodal context data, and output information regarding the selected point of interest. Accordingly, it becomes possible to accurately select a point of interest in response to the confirmation request.
[0031] Meanwhile, the processor can recognize the occupant's gaze vector information based on camera data inside the vehicle. Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request.
[0032] Meanwhile, the processor detects multiple objects based on camera data outside the vehicle, recognizes the gaze vector information of the occupant based on camera data inside the vehicle, and can select the object among the multiple objects that corresponds to the gaze vector information as the point of interest. Accordingly, it is possible to accurately select the point of interest in response to a verification request.
[0033] Meanwhile, when multiple objects are located in the occupant's line of sight, the processor can select at least one of the multiple objects as a point of interest based on the occupant's characteristic data. Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request.
[0034] Meanwhile, when a moving object is located in the occupant's line of sight, the processor can select a fixed object positioned behind the moving object as a point of interest based on map data. Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request.
[0035] Meanwhile, when a moving object is located in the occupant's line of sight, the processor obtains an inference result or a response result based on multimodal context data including the moving object, and can select a point of interest based on the obtained inference result or response result. Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request.
[0036] Meanwhile, the processor can collect information about a selected point of interest from an external server and output the collected information. Accordingly, it becomes possible to accurately select a point of interest in response to a verification request.
[0037] Meanwhile, if a fixed object is located in the occupant's line of sight, the processor selects the fixed object as a point of interest, collects information about the point of interest based on map data stored in memory, and outputs the collected information. Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request.
[0038] Meanwhile, the processor selects a fixed object as a point of interest when the fixed object is located in the occupant's line of sight, and if information regarding the point of interest does not exist in the map data stored in memory, it collects information regarding the selected point of interest from an external server and outputs the collected information. Accordingly, it becomes possible to accurately select a point of interest in response to a verification request.
[0039] Meanwhile, when multiple points of interest are located corresponding to the passenger's gaze position, the processor can select at least one point of interest among the multiple points of interest based on the passenger's characteristic data, collect information regarding the selected point of interest, and output the collected information. Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request.
[0040] Meanwhile, the processor performs voice recognition based on the passenger's voice signal, extracts context information corresponding to a confirmation request from a database among the performed voice recognition content, and can extract context data corresponding to the confirmation request based on the context information. Accordingly, it becomes possible to accurately select a point of interest in response to the confirmation request.
[0041] Meanwhile, the processor can generate a prompt corresponding to passenger characteristics based on context data including context information, a portion of image data, and map data, select a point of interest corresponding to the passenger's gaze position based on an inference result or response result obtained based on the prompt, and output information regarding the selected point of interest. Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request.
[0042] Meanwhile, the processor can generate a prompt corresponding to passenger characteristics based on context data including context information, a portion of image data, and map data, control the transmission of the generated prompt to an external server, receive an inference result or a response result from the external server, select a point of interest corresponding to the passenger's gaze position based on the received inference result or response result, and output information regarding the selected point of interest. Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request.
[0043] A signal processing device according to one embodiment of the present disclosure and a vehicle control device equipped therewith further include a neural processor. The processor generates a prompt corresponding to occupant characteristics based on context data including context information, a portion of image data, and map data, receives an inference result or response result based on the prompt generated by the neural processor, selects a point of interest corresponding to the occupant's gaze position based on the received inference result or response result, and outputs information regarding the selected point of interest. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0044] Meanwhile, the processor executes a hypervisor and, on the hypervisor, executes a display virtualization machine for a display, and the display virtualization machine selects a point of interest corresponding to the occupant's gaze position based on an inference result or response result obtained based on extracted context data, and can output information about the selected point of interest. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0045] Meanwhile, the processor further runs a gateway virtualization machine for gateway operation on the hypervisor, and the gateway virtualization machine receives camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of the occupants, and transmits the camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of the occupants to a display virtualization machine through shared memory within the hypervisor. Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request.
[0046] Meanwhile, the processor further executes a driving control virtualization machine on the hypervisor, and the driving control virtualization machine receives camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of passengers through shared memory within the hypervisor, and can execute a driving control service or a driving control application based on the camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of passengers. Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request.
[0047] Figure 1 is a drawing illustrating an example of the exterior and interior of a vehicle.
[0048] Figure 2 is a diagram illustrating an example of the architecture of a vehicle control device.
[0049] FIG. 3a is a drawing illustrating an example of the arrangement of displays inside a vehicle.
[0050] Figure 3b is a drawing illustrating another example of the arrangement of displays inside a vehicle.
[0051] FIG. 4 is an example of an internal block diagram of a vehicle control device according to an embodiment of the present disclosure.
[0052] FIG. 5 is an example of a configuration diagram of a signal processing device according to an embodiment of the present disclosure.
[0053] FIG. 6 is an example of a block diagram of a vehicle control device according to an embodiment of the present disclosure.
[0054] FIGS. 8a to 8c are drawings referenced in the description of FIG. 7.
[0055] FIG. 9 is an example of a block diagram of a signal processing system according to an embodiment of the present disclosure.
[0056] FIGS. 10a to 22 are drawings referenced in the description of FIG. 9.
[0057] The present disclosure will be described in more detail below with reference to the drawings.
[0058] The suffixes "module" and "part" for components used in the following description are assigned solely for the ease of drafting this specification and do not inherently confer any particularly significant meaning or role. Accordingly, the terms "module" and "part" may be used interchangeably.
[0059] Figure 1 is a drawing illustrating an example of the exterior and interior of a vehicle.
[0060] Referring to the drawing, the vehicle (200) is operated by a plurality of wheels (103FR, 103FL, 103RL,...) that rotate by a power source, and a steering wheel (150) for controlling the direction of travel of the vehicle (200).
[0061] Meanwhile, the vehicle (200) may further be equipped with a camera (195), etc., for acquiring an image of the front of the vehicle.
[0062] Meanwhile, the vehicle (200) may be equipped with a plurality of displays (180a, 180b) for displaying images, information, etc. inside.
[0063] In FIG. 1, a cluster display (180a) and an IVI (In-Vehicle Infotainment) display (180b) are exemplified as multiple displays (180a, 180b). In addition, a HUD (Head Up Display) and the like are also possible.
[0064] Meanwhile, the IVI (In-Vehicle Infotainment) display (180b) may also be named a Center Information Display or an AVN (Audio Video Navigation) display.
[0065] Meanwhile, the vehicle (200) described in this specification may be a concept that includes all of the following: a vehicle equipped with an engine as a power source, a hybrid vehicle equipped with an engine and an electric motor as a power source, an electric vehicle equipped with an electric motor as a power source, etc.
[0066] Figure 2 is a diagram illustrating an example of the architecture of a vehicle control device.
[0067] Referring to the drawing, the architecture (300a) of the vehicle control device can correspond to a zone-based architecture.
[0068] Accordingly, sensor devices and processors inside the vehicle may be placed in each of the multiple zones (Z1 to Z4), and a signal processing device (170a) including a gateway (GWDa) may be placed in the central area of the multiple zones (Z1 to Z4).
[0069] Meanwhile, the signal processing device (170a) may additionally include an autonomous driving control module (ACC), a cockpit control module (CPG), etc., in addition to the gateway (GWDa).
[0070] The gateway (GWDa) within the signal processing device (170a) may be a High Performance Computing (HPC) gateway.
[0071] That is, the signal processing device (170a) of FIG. 2 is an integrated HPC and can exchange data with an external communication module (not shown) or a processor (not shown) in a plurality of zones (Z1 to Z4).
[0072] FIG. 3a is a drawing illustrating an example of the arrangement of displays inside a vehicle.
[0073] Referring to the drawing, the vehicle interior may be equipped with a cluster display (180a), an IVI (In-Vehicle Infotainment) display (180b), a rear seat entertainment display (180c, 180d), a rearview mirror display (not shown), etc.
[0074] Meanwhile, in addition to the display, an interior camera (195i) may be installed inside the vehicle.
[0075] Figure 3b is a drawing illustrating another example of the arrangement of displays inside a vehicle.
[0076] A vehicle control device (100) according to an embodiment of the present disclosure may include a plurality of displays (180a to 180b) and a signal processing device (170) that performs signal processing for displaying images, information, etc. on the plurality of displays (180a to 180b) and outputs an image signal to at least one display (180a to 180b).
[0077] Among the plurality of displays (180a to 180b), the first display (180a) is a cluster display (180a) for displaying driving status, operation information, etc., and the second display (180b) may be an IVI (In-Vehicle Infotainment) display (180b) for displaying vehicle operation information, navigation map, various entertainment information or video.
[0078] The signal processing device (170) has a processor (175) inside and can execute a first virtualization machine to a third virtualization machine (not shown) on a hypervisor (not shown) within the processor (175).
[0079] A second virtualization machine (not shown) operates for the first display (180a), and a third virtualization machine (not shown) can operate for the second display (180b).
[0080] Meanwhile, the first virtualization machine (not shown) within the processor (175) can be controlled to set up a shared memory (508) based on a hypervisor (505) for the same data transmission to the second virtualization machine (not shown) and the third virtualization machine (not shown). Accordingly, the same information or the same image can be synchronized and displayed on the first display (180a) and the second display (180b) within the vehicle.
[0081] Meanwhile, the first virtualization machine (not shown) within the processor (175) shares at least a portion of the data with the second virtualization machine (not shown) and the third virtualization machine (not shown) for data sharing processing. Accordingly, data can be shared and processed by multiple virtualization machines for multiple displays within the vehicle.
[0082] Meanwhile, the first virtualization machine (not shown) within the processor (175) can receive and process wheel speed sensor data of the vehicle and transmit the processed wheel speed sensor data to at least one of the second virtualization machine (not shown) or the third virtualization machine (not shown). Accordingly, the wheel speed sensor data of the vehicle can be shared with at least one virtualization machine, etc.
[0083] Meanwhile, the vehicle control device (100) according to the embodiment of the present disclosure may further include a rear seat entertainment display (180c) for displaying driving status information, simple navigation information, various entertainment information or images.
[0084] The signal processing device (170) can control the RSE display (180c) by running a fourth virtualization machine (not shown) in addition to the first to third virtualization machines (not shown) on a hypervisor (not shown) within the processor (175).
[0085] Accordingly, various displays (180a to 180c) can be controlled using a single signal processing device (170).
[0086] Meanwhile, some of the multiple displays (180a to 180c) operate under a Linux OS, and others can operate under a Web OS.
[0087] A signal processing device (170) according to an embodiment of the present disclosure can control displays (180a to 180c) operating under various operating systems (OS) to synchronize and display the same information or the same image.
[0088] Meanwhile, FIG. 3b illustrates that a vehicle speed indicator (212a) and a vehicle interior temperature indicator (213a) are displayed on a first display (180a), a home screen (222) including a plurality of applications, a vehicle speed indicator (212b), and a vehicle interior temperature indicator (213b) is displayed on a second display (180b), and a second home screen (222b) including a plurality of applications and a vehicle interior temperature indicator (213c) is displayed on a third display (180c).
[0089] FIG. 4 is an example of an internal block diagram of a vehicle control device according to an embodiment of the present disclosure.
[0090] Referring to the drawings, a vehicle control device (100) according to an embodiment of the present disclosure may include an input unit (110), a communication unit (120) for communication with an external device, a plurality of communication modules (EMa~EMd) for internal communication, a memory (140), a signal processing unit (170), a plurality of displays (180a~180c), an audio output unit (185), and a power supply unit (190).
[0091] Multiple communication modules (EMa~EMd) can be placed in each of the multiple zones (Z1~Z4) of FIG. 2, for example.
[0092] Meanwhile, the signal processing device (170) may have a communication switch (736b) inside for data communication with each communication module (EM1~EM4).
[0093] Each communication module (EM1~EM4) can perform data communication with a plurality of sensor devices (SN), ECU (770), or area signal processing device (170Z).
[0094] Meanwhile, a plurality of sensor devices (SN) may include a camera (195), lidar (196), radar (197), or position sensor (198).
[0095] The input unit (110) may be equipped with physical buttons, pads, etc. for button input, touch input, etc.
[0096] Meanwhile, the input unit (110) may be equipped with a microphone (not shown) for user voice input.
[0097] The communication unit (120) can exchange data wirelessly with a mobile terminal (600) or a server (400).
[0098] In particular, the communication unit (120) can wirelessly exchange data with the vehicle driver's mobile terminal. Various data communication methods are possible as wireless data communication methods, such as Bluetooth, WiFi, WiFi Direct, and APiX.
[0099] The communication unit (120) can receive weather information, road traffic condition information, for example, TPEG (Transport Protocol Expert Group) information from a mobile terminal (600) or a server (400). To this end, the communication unit (120) may be equipped with a mobile communication module (not shown).
[0100] A plurality of communication modules (EM1~EM4) can receive sensor data, etc. from an ECU (770), a sensor device (SN), or a region signal processing device (170Z), and transmit the received sensor data to the signal processing device (170).
[0101] Here, the sensor data may include at least one of vehicle direction data, vehicle location data (GPS data), vehicle angle data, vehicle speed data, vehicle acceleration data, vehicle tilt data, vehicle forward / reverse data, battery data, fuel data, tire data, vehicle lamp data, vehicle interior temperature data, and vehicle interior humidity data.
[0102] Such sensor data can be obtained from a heading sensor, a yaw sensor, a gyro sensor, a position module, a vehicle forward / reverse sensor, a wheel sensor, a vehicle speed sensor, a vehicle body inclination sensor, a battery sensor, a fuel sensor, a tire sensor, a steering sensor based on steering wheel rotation, a vehicle interior temperature sensor, a vehicle interior humidity sensor, etc.
[0103] Meanwhile, the position module may include a GPS module or a position sensor (198) for receiving GPS information.
[0104] Meanwhile, at least one of the multiple communication modules (EM1 to EM4) can transmit location information data sensed from a GPS module or a location sensor (198) to a signal processing device (170).
[0105] Meanwhile, at least one of the plurality of communication modules (EM1 to EM4) can receive vehicle front image data, vehicle side image data, vehicle rear image data, and obstacle distance information around the vehicle from a camera (195), lidar (196), radar (197), etc., and transmit the received information to a signal processing device (170).
[0106] The memory (140) can store various data for the overall operation of the vehicle control device (100), such as a program for processing or controlling the signal processing device (170).
[0107] For example, memory (140) can store data regarding a hypervisor, a first virtualization machine to a third virtualization machine, for execution within a processor (175).
[0108] The audio output unit (185) converts an electrical signal from the signal processing device (170) into an audio signal and outputs it. To do this, a speaker or the like may be provided.
[0109] The power supply unit (190) can supply power necessary for the operation of each component under the control of the signal processing unit (170). In particular, the power supply unit (190) can receive power from a battery inside the vehicle, etc.
[0110] The signal processing unit (170) controls the overall operation of each unit within the vehicle control unit (100).
[0111] For example, the signal processing device (170) may include a processor (175) that performs signal processing for a vehicle display (180a, 180b).
[0112] The processor (175) can run a first virtualization machine to a third virtualization machine (not shown) on a hypervisor (505 in FIG. 5) within the processor (175).
[0113] Among the first to third virtual machines (not shown), the first virtual machine (not shown) may be named a Server Virtual Machine, and the second to third virtual machines (not shown) may be named a Guest Virtual Machine.
[0114] For example, a first virtualization machine (not shown) within a processor (175) can receive sensor data from a plurality of sensor devices, such as vehicle sensor data, location information data, camera image data, audio data, or touch input data, and process or modify it to output it.
[0115] In this way, by performing most of the data processing in the first virtualization machine (not shown), 1:N data sharing becomes possible.
[0116] As another example, the first virtualization machine (not shown) can directly receive and process CAN data, Ethernet data, audio data, radio data, USB data, and wireless communication data for the second virtualization machine to the third virtualization machine (not shown).
[0117] And, the first virtualization machine (not shown) can transmit the processed data to the second virtualization machine to the third virtualization machine (not shown).
[0118] Accordingly, among the first to third virtualization machines (not shown), only the first virtualization machine (not shown) receives sensor data, communication data, or external input data from a plurality of sensor devices and performs signal processing, thereby reducing the signal processing burden on other virtualization machines and enabling 1:N data communication, which enables synchronization when sharing data.
[0119] Meanwhile, the first virtualization machine (not shown) can control the sharing of the same data with the second virtualization machine (not shown) and the third virtualization machine (not shown) by writing data to the shared memory (508 in FIG. 5).
[0120] For example, the first virtualization machine (not shown) can record vehicle sensor data, the location information data, the camera image data, or the touch input data in a shared memory (508) and control the sharing of the same data with the second virtualization machine (not shown) and the third virtualization machine (not shown). Accordingly, data sharing in a 1:N manner becomes possible.
[0121] Ultimately, by performing most of the data processing on the first virtualization machine (not shown), 1:N data sharing becomes possible.
[0122] Meanwhile, the first virtualization machine (not shown) within the processor (175) can control the second virtualization machine (not shown) and the third virtualization machine (not shown) to set up a shared memory (508) based on the hypervisor (505) for the same data transmission.
[0123] Meanwhile, the signal processing device (170) can process various signals such as audio signals, video signals, and data signals. To this end, the signal processing device (170) can be implemented in the form of a System On Chip (SOC).
[0124] Meanwhile, the signal processing device (170) in the display device (100) of FIG. 4 may be the same as the signal processing device (170) of the vehicle control device of FIG. 5 and below.
[0125] FIG. 5 is an example of a configuration diagram of a signal processing device according to an embodiment of the present disclosure.
[0126] Referring to the drawings, a signal processing device (170) according to an embodiment of the present disclosure includes a processor (175).
[0127] Meanwhile, the signal processing device (170) may be named as an HPC (High Performance Computing) signal processing device.
[0128] Meanwhile, the signal processing device (170) according to an embodiment of the present disclosure may further include a neural processor (179).
[0129] Meanwhile, the neural processor (179) may also be referred to as an on-device-based learning processor.
[0130] Meanwhile, the processor (175) in the signal processing device (170) can execute the hypervisor (505) and execute the first to third virtualization machines (510 to 530) on the hypervisor (505).
[0131] The first virtualization machine (510) may be a gateway virtualization machine corresponding to the gateway (GWDa) of FIG. 2.
[0132] The second virtualization machine (520) may be a driving control virtualization machine corresponding to the autonomous driving control module (ACC) of FIG. 2.
[0133] The driving control virtualization machine (520) at this time can control the vehicle driving assistance (ADAS) or the autonomous driving (AD).
[0134] The third virtualization machine (530) may be a display virtualization machine corresponding to the cockpit control module (CPG) of FIG. 2 or the display (180a, 180b, 180c) of FIG. 3.
[0135] For example, the third virtualization machine (530) may include a cluster virtualization machine (530b) for a cluster display (180a) and an IVI virtualization machine (530b) for an IVI display (180b).
[0136] Meanwhile, the third virtualization machine (530) may further include a HUD virtualization machine (not shown) for a HUD display (180c).
[0137] Meanwhile, the processor (175) can share data between each virtualization machine (510~530) through shared memory (505) within the hypervisor (505).
[0138] Meanwhile, the processor (175) can share data with a plurality of sensor devices (SN) or ECUs (770) or area signal processing devices (170Z) through shared memory (505) within the hypervisor (505).
[0139] Meanwhile, the processor (175) can exchange data with an external server (400) or mobile terminal (600) through the communication device (120) of FIG. 4.
[0140] Meanwhile, the data received by the processor (175) within the signal processing device (170) may include camera data or sensor data.
[0141] For example, sensor data within the vehicle may include at least one of vehicle wheel speed data, vehicle direction data, vehicle location data (GPS data), vehicle angle data, vehicle speed data, vehicle acceleration data, vehicle tilt data, vehicle forward / reverse data, battery data, fuel data, tire data, vehicle lamp data, vehicle interior temperature data, vehicle interior humidity data, vehicle external radar data, and vehicle external lidar data.
[0142] Meanwhile, camera data may include external vehicle camera data and internal vehicle camera data.
[0143] Meanwhile, the processor (175) within the signal processing device (170) can execute multiple virtualization machines (510 to 530) based on safety standards.
[0144] Meanwhile, the processor (175) in the signal processing unit (170a) can execute the hypervisor (505) and, on the hypervisor (505), execute the first to third virtualization machines (510 to 530) according to the automotive safety integrity level (Automotive SIL and ASIL).
[0145] For example, the first virtualization machine (510) may be a virtualization machine corresponding to ASIL C or ASIL D, in which the sum of Severity, Exposure, and Controllability in the Automotive Safety Integrity Level (ASIL) is 9 or 10.
[0146] Meanwhile, ASIL D can correspond to the grade requiring the highest safety level.
[0147] The first virtualization machine (510) can run a safety operating system (not shown) and an application (not shown) on the safety operating system.
[0148] Meanwhile, the first virtualization machine (510) may run a container runtime (not shown) and a container runtime (not shown) on a safety operating system.
[0149] Meanwhile, unlike the drawing, the first virtualization machine (510) may also be executed through a separate processor core instead of the processor (175).
[0150] Meanwhile, the second virtualization machine (520) may be a virtualization machine corresponding to ASIL A or ASIL B, in which the sum of Severity, Exposure, and Controllability in the Automotive Safety Integrity Level (ASIL) is 7 or 8.
[0151] Meanwhile, the second virtualization machine (520) can run an operating system (not shown), a container runtime (not shown) on the operating system (not shown), and a container (not shown) on the container runtime.
[0152] Alternatively, the second virtualization machine (520) can run an operating system (not shown) and an application on the operating system (not shown).
[0153] Meanwhile, the third virtualization machine (530) may be a virtualization machine corresponding to Quality Management (QM), which is the lowest safety level and non-mandatory grade in the Automotive Safety Integrity Level (ASIL).
[0154] Meanwhile, the third virtualization machine (530) can run an operating system (not shown), a container runtime (not shown) on the operating system (not shown), and a container (not shown) on the container runtime.
[0155] Alternatively, the third virtualization machine (530) can run an operating system (not shown) and an application on the operating system (not shown).
[0156] FIG. 6 is an example of a block diagram of a vehicle control device according to an embodiment of the present disclosure.
[0157] Referring to the drawings, a vehicle control device (900) according to an embodiment of the present disclosure comprises a signal processing device (170).
[0158] A vehicle control device (900) according to an embodiment of the present disclosure may further include at least one display.
[0159] In the drawing, at least one display is exemplified as a cluster display (180a) and an IVI display (180b).
[0160] Meanwhile, the vehicle control device (900) may further include a plurality of area signal processing devices (170Z1 to 170Z4).
[0161] The signal processing device (170) at this time is a high-performance centralized signal processing and control device having a plurality of CPUs (175), GPUs (178), NPUs (179), etc., and can be named as a High Performance Computing (HPC) signal processing device or a central signal processing device.
[0162] Multiple area signal processing devices (170Z1~170Z4) and signal processing device (170) are connected by wired cables (CB1~CB4).
[0163] Meanwhile, multiple area signal processing devices (170Z1~170Z4) can be connected to each other by wired cables (CBa~CBd).
[0164] The wired cable (CBa~CBd) at this time may include a CAN communication cable, an Ethernet communication cable, or a PCI Express cable.
[0165] Meanwhile, the signal processing device (170) according to the embodiment of the present disclosure may have at least one processor (175, 178, 177) and a large-capacity storage device (925).
[0166] For example, a signal processing device (170) according to an embodiment of the present disclosure may include a central processor (175, 177), a graphics processor (178), and a neural processor (179).
[0167] Meanwhile, sensor data can be transmitted from at least one of the multiple area signal processing devices (170Z1 to 170Z4) to the signal processing device (170). In particular, the sensor data can be stored in a storage device (925) within the signal processing device (170).
[0168] The sensor data at this time may include at least one of camera data, lidar data, radar data, vehicle direction data, vehicle position data (GPS data), vehicle angle data, vehicle speed data, vehicle acceleration data, vehicle tilt data, vehicle forward / reverse data, battery data, fuel data, tire data, vehicle lamp data, vehicle interior temperature data, and vehicle interior humidity data.
[0169] In the drawing, camera data from a camera (195a) and lidar data from a lidar sensor (196) are input to a first area signal processing device (170Z1), and the camera data and lidar data are transmitted to a signal processing device (170) via a second area signal processing device (170Z2) and a third area signal processing device (170Z3), etc.
[0170] Meanwhile, since the data reading or writing speed to the storage device (925) is faster than the network speed when sensor data is transmitted from at least one of the multiple area signal processing devices (170Z1~170Z4) to the signal processing device (170), it is desirable to perform multipath routing so that network bottlenecks do not occur.
[0171] To this end, the signal processing device (170) according to an embodiment of the present disclosure can perform multipath routing based on a Software Defined Network (SDN). Accordingly, a stable network environment can be secured when reading or writing data of the storage device (925). Furthermore, since data can be transmitted to the storage device (925) using multiple paths, data can be transmitted by dynamically changing the network configuration.
[0172] Data communication between a plurality of area signal processing devices (170Z1~170Z4) and a signal processing device (170) within a vehicle control device (900) according to an embodiment of the present disclosure is preferably Peripheral Component Interconnect Express communication or Ethernet communication for high-bandwidth, low-latency communication.
[0173] FIG. 7 is an example of a block diagram of a signal processing device according to an embodiment of the present disclosure.
[0174] Referring to the drawings, the signal processing system (700) according to an embodiment of the present disclosure may be referred to as an IVEX (In Vehicle Experience) system.
[0175] Meanwhile, the signal processing system (700) according to an embodiment of the present disclosure may include an edge sensor group (SN), a signal processing device (170) in a vehicle, and a server (400).
[0176] Meanwhile, the signal processing device (170) according to an embodiment of the present disclosure includes a recognizer (720), an insightor (750), and an illustrator (780).
[0177] Meanwhile, the edge sensor group (SN) can correspond to a plurality of sensor devices (SN) of FIG. 2.
[0178] Meanwhile, the recognizer (720), insightor (750), and illustrator (780) may be included in the signal processing device (170) of the vehicle display device (100) of FIG. 4.
[0179] In particular, the recognizer (720), the insightor (750), and the illustrator (780) may be included in the processor (175) of the signal processing device (170).
[0180] Meanwhile, the edge sensor group (SN) may include vehicle interior sensors (710) and vehicle exterior sensors (715).
[0181] Meanwhile, the edge sensor group (SN) may further include a data receiving unit (718) for receiving vehicle internal data or vehicle external data. The data receiving unit (718) may correspond to the communication device (120) of FIG. 2.
[0182] Meanwhile, the vehicle interior sensors (710) are sensors placed inside the vehicle (200) and may include a front camera (195), lidar (196), radar (197), interior camera (195i), vehicle interior temperature sensor, or vehicle interior humidity sensor.
[0183] The vehicle external sensors (715) are sensors positioned on the exterior of the vehicle (200) and may include an external camera, lidar (196), radar (197), heading sensor, yaw sensor, gyro sensor, position module, vehicle forward / reverse sensor, wheel sensor, vehicle speed sensor, vehicle body inclination sensor, battery sensor, fuel sensor, tire sensor, or steering sensor based on steering wheel rotation. The position module may include a GPS module or a position sensor (198) for receiving GPS information.
[0184] Sensing data may include vehicle interior sensing data and vehicle exterior sensing data.
[0185] The vehicle interior sensing data may be data sensed by the vehicle interior sensors (710).
[0186] Vehicle interior sensing data may include at least one of vehicle interior temperature data, vehicle interior humidity data, battery data, fuel data, vehicle lamp data, tire data, or vehicle interior camera data, or audio data received through a microphone.
[0187] The vehicle external sensing data may be data sensed by the vehicle external sensors (715).
[0188] Vehicle external sensing data may include at least one of vehicle location data (GPS), vehicle direction data, vehicle angle data, vehicle speed data, vehicle acceleration data, vehicle tilt data, whether the vehicle is moving forward or backward, front camera data, or rear camera data.
[0189] Meanwhile, the recognizer (720) can recognize the vehicle situation based on sensing data received from the edge sensor group (SN). In this regard, the recognizer (720) may be referred to as a situation recognition unit.
[0190] The vehicle situation at this time may include external vehicle conditions and internal vehicle conditions.
[0191] Meanwhile, the recognizer (720) can recognize the vehicle situation based on the vehicle internal sensing data or the vehicle internal sensing data, and can generate vehicle situation information regarding the recognized vehicle situation.
[0192] For example, the vehicle situation information (740) may include at least one of the external situation information of the vehicle and the situation information of the vehicle's occupants.
[0193] As another example, the vehicle situation information (740) may include at least one of the vehicle's external situation information, the vehicle's internal situation information, and the vehicle's occupant situation information.
[0194] Meanwhile, the recognizer (720) can recognize external situation information of the vehicle based on external sensing data from the edge sensor group (SN), recognize internal situation information of the vehicle based on internal sensing data of the vehicle, or recognize situation information of the vehicle's occupant based on internal sensing data of the vehicle.
[0195] Meanwhile, the recognizer (720) can output external situation information of the vehicle based on internal vehicle sensing data or external vehicle sensing data, output internal situation information of the vehicle based on internal vehicle sensing data or external vehicle sensing data, or output situation information of the vehicle's occupant based on internal vehicle sensing data or external vehicle sensing data.
[0196] Meanwhile, the recognizer (720) can perform preprocessing and calibration of the vehicle interior sensing data or the vehicle interior sensing data, and can obtain vehicle situation information based on the preprocessed and calibrated sensing data.
[0197] For example, the internal recognition unit (725) within the recognition unit (720) can perform preprocessing and calibration (721) on the vehicle internal sensing data, and can obtain internal situation information of the vehicle based on the preprocessed and calibrated vehicle internal sensing data.
[0198] Specifically, the internal recognizer (725) can perform preprocessing and calibration (721) on the vehicle interior sensing data, perform primitive detection (722) or sensor fusion (723), and based on this, obtain information about the vehicle's interior situation.
[0199] For example, an external recognizer (730) within the recognizer (720) can perform preprocessing and calibration (721) on the external vehicle sensing data, and can obtain external situation information of the vehicle based on the preprocessed and calibrated external vehicle sensing data.
[0200] Specifically, the external recognizer (730) can perform preprocessing and calibration (731) on external vehicle sensing data, execute a perception engine (732), perform sensor fusion (733), or perform segmentation (734), and based on this, obtain external situation information of the vehicle.
[0201] Meanwhile, the passenger recognition device (730) within the recognition device (720) can perform facial recognition (741) of the passenger based on the vehicle interior sensing data, recognize distraction (742), recognize gaze, position, gesture (743), recognize emotion (744), recognize whether the passenger is drinking or taking medication (746), recognize drowsiness (747), or recognize information (748) regarding other actions, and obtain situational information of the vehicle's passenger based thereon.
[0202] Meanwhile, the situational information of the occupant may include at least one of the following: a face identifier (Face ID) of the occupant inside the vehicle (200), distraction information, gaze information, position information, gesture information, emotion information, information on whether alcohol or drugs have been taken, drowsiness information, or other information regarding actions.
[0203] Meanwhile, the recognizer (720) can transmit the generated vehicle situation information to the insightor (750).
[0204] Meanwhile, the insightor (750) can generate an inference result or a response result based on the situation information of the vehicle received from the recognizer (720).
[0205] Meanwhile, the insightor (750) can generate an inference result or a response result based on the vehicle situation information received from the recognizer (720), and can generate or control a service to be executed based on the inference result or the response result.
[0206] Accordingly, the insightor (750) may be named a response result generation unit or a service generation unit.
[0207] Meanwhile, the insightor (750) may include a multimodal context engine (755), a context controller (754), a safety assistance engine (757), a user characteristic engine (758), an AI orchestrator (760), and an AI model set (765).
[0208] Meanwhile, the multimodal context engine (755) can generate multimodal context data based on the vehicle situation information received from the recognizer (720).
[0209] For example, multimodal context data may include at least one of text data or image data describing a vehicle situation generated based on vehicle situation information.
[0210] Meanwhile, the context controller (754) may be a component included in the multimodal context engine (755) or provided separately from the multimodal context engine (755).
[0211] Meanwhile, the context controller (754) can execute the process of generating multimodal context data when it receives a passenger query from the AI orchestrator (760).
[0212] For example, the passenger curie may be a voice recognition result corresponding to a voice command spoken by the passenger.
[0213] Meanwhile, the safety assist engine (757) can determine whether the current situation is a safe situation or a dangerous situation based on the vehicle situation information received from the recognizer (720).
[0214] Meanwhile, the safety assist engine (757) can transmit a driver assistance control command or a warning notification output command to the safety application (790) of the illustrator (780) if the current situation is determined to be a dangerous situation.
[0215] Meanwhile, the user characteristic engine (758) can generate user context data or passenger context data based on user information or passenger information.
[0216] Meanwhile, user information or passenger information may be referred to as user persona or passenger persona.
[0217] Meanwhile, user information or passenger information may include at least one of the user or passenger's nationality, age, gender, occupation, personality, or psychological type (Myers-Briggs Type Indicator, MBTI).
[0218] Meanwhile, the AI orchestrator (760) can generate a prompt based on at least one of multimodal context data or passenger context data.
[0219] Meanwhile, the AI orchestrator (760) can send the generated prompt to the AI model set (765).
[0220] Meanwhile, the AI model set (765) may include at least one AI model.
[0221] For example, the AI model set (765) may include an on-device AI model (767).
[0222] Meanwhile, the on-device AI model (767) may include at least one AI model. For example, the on-device AI model (767) may include a Small Language Model (LLM).
[0223] Meanwhile, the AI model set (765) may include an interface (766) for data exchange with the AI model (769) in the server (400).
[0224] Meanwhile, the AI model (769) within the server (400) may include at least one AI model. For example, the AI model (769) within the server (400) may include a Large Language Model (LLM).
[0225] That is, the on-device AI model (767) may have a smaller capacity or size than the AI model (769) in the server (400).
[0226] Meanwhile, the AI model set (765) can output an inference result in response to a prompt received from the AI orchestrator (760) and can transmit the inference result to the AI orchestrator (760).
[0227] Meanwhile, the AI orchestrator (760) can generate additional prompts based on the inference results and send the additional prompts to the AI model set (765).
[0228] Meanwhile, the AI model set (765) that receives the additional prompt can output an additional inference result in response to the additional prompt and can transmit the additional inference result to the AI orchestrator (760).
[0229] The AI orchestrator (760) can obtain an inference result or additional inference result received from the AI model set (765) as a response result, and can output the obtained response result to the illustrator (780).
[0230] Meanwhile, the illustrator (780) can execute a service, output service information, execute an application, or output application information based on the response result output from the insightor (750).
[0231] For example, the illustrator (780) can execute a navigation service, execute a vehicle driving assistance control service, execute an autonomous driving service, or execute a display-related service based on the response result output from the insightor (750).
[0232] As another example, the illustrator (780) can run a navigation application, run a vehicle driving assistance control application, run an autonomous driving application, or run a display-related application based on the response result output from the insightor (750).
[0233] Meanwhile, the illustrator (780) may include a multimodal output encoder (781), a visual interface (782), an audio interface (785), and a safety application (790).
[0234] Meanwhile, the multimodal output encoder (781) can encode the response result output from the insightor (750) and output the encoded response result data to the visual interface (782) or audio interface (785).
[0235] Meanwhile, the visual interface (782) or audio interface (785) may be referred to as an output interface.
[0236] Meanwhile, the visual interface (782) can output response result data output from the multimodal output encoder (781).
[0237] Accordingly, at least one of the plurality of displays (180a to 180c) of FIG. 4 can display an image based on response result data from the visual interface (782).
[0238] Meanwhile, the visual interface (782) can output augmented reality (AR) video based on response result data or mixed reality (MR) video data based on response result data.
[0239] Meanwhile, the audio interface (785) can output response result data in the form of audio.
[0240] Accordingly, the audio output unit (185) of FIG. 4 can output a sound corresponding to the response result data from the audio interface (785).
[0241] Meanwhile, the safety application (790) can perform Advanced Driver Assistance System (ADAS) control based on the driver assistance control command or the response result received from the insightor (750).
[0242] For example, the safety application (790) can output warning notification data according to a warning notification output command.
[0243] In response to this, at least one of the electronic control unit (770) of FIG. 4 or a plurality of displays (180a to 180c) can output warning notification data.
[0244] FIGS. 10a to 22 are drawings referenced in the description of FIG. 9.
[0245] First, FIG. 8a is a drawing referenced in the description of the insight of FIG. 7.
[0246] Referring to the drawing, the insightor (750) may include a multimodal context engine (755), an AI orchestrator (760), and a multimodal LLM (767).
[0247] Meanwhile, the multimodal context engine (755) may include a multimodal signal adapter (751), a multimodal indexer (812), a multimodal context buffer (752), a multimodal context retriever (814), a multimodal context descriptor (815), a multimodal event monitor (811), and a context controller (754).
[0248] Unlike Fig. 7, the multimodal context engine (755) may include a context controller (754).
[0249] Meanwhile, the multimodal signal adapter (751) can generate pre-processed multimodal data by filtering, cleaning, synchronizing, and reformulating data received from various sensors or vehicle situation information received from the recognizer (720).
[0250] Meanwhile, the multimodal signal adapter (751) can receive at least one of audio data received through a microphone, front image data captured through a front camera, ADAS information obtained from the front image data or vehicle sensor, location data, IVI (In-Vehicle Infotainment) system data, IVI display information, DMS (Driver Monitoring System) information or IMS (Interior Monitoring System) information based on image data captured through an interior camera (195i), and biometric data obtained from a biometric sensor.
[0251] Meanwhile, the multimodal indexer (812) can generate multimodal processing data by dividing the preprocessed multimodal data into chunks and can index the multimodal processing data.
[0252] Meanwhile, the multimodal indexer (812) can obtain an encoding vector or keyword representing the attribute (or meaning) of the multimodal processed data divided into chunks as an index.
[0253] Meanwhile, the multimodal context buffer (752) can store multimodal processing data and an index corresponding to the multimodal processing data.
[0254] Meanwhile, the multimodal context retriever (814) can search for multimodal processing data most related to the passenger query through an index from the multimodal context buffer (752), and can select the searched multimodal processing data as a context candidate for the creation of multimodal context data.
[0255] Meanwhile, the multimodal context descriptor (815) can reconfigure multimodal processing data selected as a context candidate into multimodal context data having a prompt form that the multimodal LLM (767) can interpret.
[0256] Meanwhile, the multimodal event monitor (811) can monitor whether an index matching the trigger condition is entered when a trigger condition for multimodal data is registered. If an index matching the trigger condition is entered, the multimodal event monitor (811) can generate a trigger event to generate a proactive service query.
[0257] Meanwhile, the context controller (754) can control the overall operation of the multimodal context engine (755).
[0258] Meanwhile, the context controller (754) can execute the process of generating multimodal context data when it receives a passenger query from the AI orchestrator (760).
[0259] Meanwhile, the context controller (754) can generate a preemptive service query when it receives a trigger event from the multimodal event monitor (811).
[0260] Meanwhile, the AI orchestrator (760) can generate a prompt by combining the passenger query and the multimodal context, and can send the generated prompt to the multimodal LLM (767).
[0261] Meanwhile, the passenger curry may be a query in the form of recognized text based on a voice command spoken by the passenger. The voice command may be received through a microphone, and the voice command may be converted into text through an Automatic Speech Recognition (ASR) process.
[0262] Meanwhile, passenger curie may be text converted through the Automatic Speech Recognition (ASR) process.
[0263] Meanwhile, multimodal LLM (767) may be an example of an AI model included in the set of AI models (765) of FIG. 6a.
[0264] Meanwhile, the multimodal LLM (767) can output an inference result from a prompt received from the AI orchestrator (760) and can transmit the inference result to the AI orchestrator (760).
[0265] Meanwhile, the inference result may include a result representing a response to the prompt.
[0266] For example, the inference result may include an API call result regarding whether an API call corresponding to a driving assistance function was successfully performed, and a feedback generation result regarding whether feedback corresponding to the API call was successfully generated.
[0267] Next, Fig. 8b is a drawing referenced in the description of the AI orchestrator of Fig. 7.
[0268] Referring to the drawing, the AI orchestrator (760) may include a task arbitrator (771), a prompt manager (763), a sub-agent set (761), a knowledge database (762), a workflow controller (772), and a tool set / adapter set (764).
[0269] Meanwhile, the task arbitrator (771) can determine one of the multiple sub-agents based on the voice recognition result and the multimodal context data output from the multimodal context engine (755).
[0270] Meanwhile, the prompt manager (763) can generate a prompt for the operation of the determined sub-agent. The prompt manager (763) can generate a prompt based on multimodal context data and passenger queries.
[0271] Meanwhile, the prompt manager (763) can generate a prompt based on information about the functions that the determined sub-agent can perform, multimodal context data, the results of previously performed functions, conversation history and passenger query.
[0272] Meanwhile, the sub-agent set (761) may include multiple sub-agents.
[0273] For example, the sub-agent set (761) may include a navigation agent for navigation services, a vehicle function agent for providing vehicle functions, a telephony agent for automated telephone answering services, and a Q&A agent for providing response services to queries.
[0274] Meanwhile, the sub-agent determined by the task arbitrator (771) can call the cloud AI model (769) or on-device AI model (767) assigned to it.
[0275] Meanwhile, the cloud AI model (769) or on-device AI model (767) may be a Large Language Model (LLM).
[0276] Meanwhile, a sub-agent within a sub-agent set (761) can call a cloud AI model (769) or an on-device AI model (767) assigned to it to obtain an inference result corresponding to a prompt from the model.
[0277] Meanwhile, a sub-agent in the sub-agent set (761) can call the knowledge database (762) to provide additional information based on the acquired inference result, obtain additional information from the knowledge database (762), specify the name of the function to be executed and the parameter value of the function, and determine the feedback phrase to be provided to the user.
[0278] Meanwhile, a sub-agent in the sub-agent set (761) can transmit to the workflow controller (772) a parsing result including additional information called from the knowledge database (762), the name of the function to be executed, the parameter value of the function, and a feedback phrase, which is generated by parsing the inference result received from the model.
[0279] Meanwhile, the knowledge database (762) can store additional information and information about functions. The information about functions may include the name of the function and the parameter values of the function.
[0280] Meanwhile, the workflow controller (772) can generate control commands to perform the corresponding function based on the parsing results and can store the results of the conversation and function performed by the AI model and sub-agent.
[0281] Meanwhile, the tool set / adapter set (764) can call an API corresponding to a control command generated by the workflow controller (772). The tool set / adapter set (764) can convert the control command into an execution command of an actual function or an API call command that the IVI system can understand and execute, and can execute the converted command.
[0282] Fig. 8c is an example of an internal block diagram of the server of Fig. 7.
[0283] Referring to the drawing, the server (400) may represent a device that trains an artificial neural network using a machine learning algorithm or uses a trained artificial neural network.
[0284] Here, the server (400) may be composed of multiple servers to perform distributed processing. Alternatively, the server (400) may be defined as a 5G network.
[0285] The server (400) may be included as part of the vehicle (200) and may perform at least some of the AI processing together.
[0286] The server (400) may include a communication interface (410), memory (430), a learning processor (440), and a processor (470).
[0287] The communication interface (410) can transmit and receive data with the vehicle (200) or an external device.
[0288] The memory (430) may include a model storage unit (431). The model storage unit (431) may store a model (or artificial neural network, 531a) that is being learned or has been learned through the learning processor (440).
[0289] The learning processor (440) can train the artificial neural network (431a) using the training data.
[0290] For example, the learning processor (440) can execute a learning model. The learning model may include a cloud AI model (769) such as FIG. 7.
[0291] The learning model may be used while mounted on the server (400) of the artificial neural network, or may be used while mounted on an external device such as a vehicle (200).
[0292] The learning model may be implemented in hardware, software, or a combination of hardware and software. If part or all of the learning model is implemented in software, one or more instructions constituting the learning model may be stored in memory (430).
[0293] The processor (470) can infer a result value for new input data using a learning model and generate a response or control command based on the inferred result value.
[0294] Meanwhile, the processor (470) can control the learning processor (440) to execute the cloud AI model (769) based on requests from the signal processing device (170) in the vehicle (200), etc.
[0295] FIG. 9 is an example of a block diagram of a signal processing system according to an embodiment of the present disclosure.
[0296] Referring to the drawings, the signal processing system (900) according to an embodiment of the present disclosure may be referred to as an IVEX (In Vehicle Experience) system or a vehicle control device.
[0297] Meanwhile, the signal processing system (900) according to an embodiment of the present disclosure includes an edge sensor group (SN) and a signal processing device (170) in a vehicle.
[0298] Meanwhile, the signal processing system (900) according to an embodiment of the present disclosure may further include a server (400).
[0299] Meanwhile, the signal processing device (170) according to an embodiment of the present disclosure includes a recognizer (720), an insightor (750), and an illustrator (780).
[0300] Meanwhile, the edge sensor group (SN) can correspond to a plurality of sensor devices (SN) of FIG. 2.
[0301] Meanwhile, the edge sensor group (SN) may include a microphone (112) which is an example of an input device (110), a vehicle interior camera (195i), a vehicle front camera (195), and a position sensor (198) that receives GPS data, etc.
[0302] Among these, the microphone (112) and the vehicle interior camera (195i) may be included in the vehicle interior sensors (710) of FIG. 7.
[0303] Meanwhile, the vehicle front camera (195) or position sensor (198) may be included in the vehicle external sensors (715).
[0304] Meanwhile, the recognizer (720) can recognize the vehicle situation based on sensing data received from the edge sensor group (SN). In this regard, the recognizer (720) may be referred to as a situation recognition unit.
[0305] The vehicle situation at this time may include external vehicle conditions and internal vehicle conditions.
[0306] Meanwhile, the recognizer (720) can recognize the vehicle situation based on the vehicle internal sensing data or the vehicle internal sensing data, and can generate vehicle situation information regarding the recognized vehicle situation.
[0307] For example, the vehicle situation information (740) may include at least one of the external situation information of the vehicle and the situation information of the vehicle's occupants.
[0308] As another example, the vehicle situation information (740) may include at least one of the vehicle's external situation information, the vehicle's internal situation information, and the vehicle's occupant situation information.
[0309] Meanwhile, the recognizer (720) can recognize external situation information of the vehicle based on external sensing data from the edge sensor group (SN), recognize internal situation information of the vehicle based on internal sensing data of the vehicle, or recognize situation information of the vehicle's occupant based on internal sensing data of the vehicle.
[0310] Meanwhile, the recognizer (720) can output external situation information of the vehicle based on internal vehicle sensing data or external vehicle sensing data, output internal situation information of the vehicle based on internal vehicle sensing data or external vehicle sensing data, or output situation information of the vehicle's occupant based on internal vehicle sensing data or external vehicle sensing data.
[0311] Meanwhile, the recognizer (720) can perform preprocessing and calibration of the vehicle interior sensing data or the vehicle interior sensing data, and can obtain vehicle situation information based on the preprocessed and calibrated sensing data.
[0312] For example, the recognizer (720) can perform preprocessing and calibration on the vehicle interior sensing data, and can obtain information on the vehicle's interior situation based on the preprocessed and calibrated vehicle interior sensing data.
[0313] For example, the recognizer (720) can perform preprocessing and calibration on the vehicle external sensing data, and can obtain external situation information of the vehicle based on the preprocessed and calibrated vehicle external sensing data.
[0314] Meanwhile, the recognition device (720) can perform facial identification of the occupant based on the vehicle interior sensing data, recognize distraction, recognize gaze, position, gesture, recognize emotion, recognize whether alcohol or drugs have been taken, recognize drowsiness, or recognize information (748) about other actions, and obtain situational information of the vehicle occupant based thereon.
[0315] Meanwhile, the situational information of the occupant may include at least one of the following: a face identifier (Face ID) of the occupant inside the vehicle (200), distraction information, gaze information, position information, gesture information, emotion information, information on whether alcohol or drugs have been taken, drowsiness information, or other information regarding actions.
[0316] Meanwhile, the recognizer (720) can transmit the generated vehicle situation information to the insightor (750).
[0317] Meanwhile, the recognizer (720) performs voice recognition based on the passenger's voice signal from the microphone (112), and if a confirmation request is included in the performed voice recognition content, it extracts the confirmation request (724).
[0318] For example, the confirmation request (724) may include phrases such as "What's that?"
[0319] And, the recognizer (720) can send a confirmation request to the AI orchestrator (760).
[0320] Meanwhile, the recognizer (720) receives camera data from outside the vehicle from the vehicle's external camera (195) and performs object detection based on the camera data from outside the vehicle.
[0321] For example, the recognizer (720) can receive vehicle front image data from the vehicle's external camera (195), perform segmentation on the vehicle front image data, and detect multiple objects (735).
[0322] Meanwhile, the recognizer (720) receives location information (737) from the location sensor (178) and can transmit the location information (737) to the insightor (750).
[0323] Meanwhile, the recognizer (720) receives camera data inside the vehicle from the vehicle's internal camera (195i) and detects the occupant's gaze position (743) based on the camera data inside the vehicle.
[0324] For example, the recognizer (720) can receive image data of the interior of the vehicle from the vehicle's interior camera (195i) and detect the gaze position (743) of the occupant, particularly the driver, among the image data of the interior of the vehicle.
[0325] And, the recognizer (720) can select an object (743b) corresponding to the occupant's gaze position based on the extracted gaze position (743) and a plurality of objects (735).
[0326] Meanwhile, the recognizer (720) can crop or extract an area corresponding to the passenger's gaze position (743) from the vehicle's front image data from the vehicle's external camera (195) based on the passenger's gaze position (743).
[0327] For example, the recognizer (720) can crop or extract an area containing an object corresponding to the passenger's gaze position (743) among a plurality of objects (735).
[0328] Meanwhile, the recognizer (720) can transmit an object (743b) corresponding to the passenger's gaze position to the insightor (750).
[0329] Additionally, the recognizer (720) can further transmit multiple objects (735), a gaze position (743), and location information (737) to the insightor (750), respectively.
[0330] That is, the recognizer (720) can transmit an object (743b) corresponding to the passenger's gaze position, a plurality of objects (735), a gaze position (743), and location information (737) to the insightor (750).
[0331] Meanwhile, the insightor (750) can generate an inference result or a response result based on the situation information of the vehicle received from the recognizer (720).
[0332] Meanwhile, the insightor (750) can generate an inference result or a response result based on the vehicle situation information received from the recognizer (720), and can generate or control a service to be executed based on the inference result or the response result.
[0333] Accordingly, the insightor (750) may be named a response result generation unit or a service generation unit.
[0334] Meanwhile, the insightor (750) may include a multimodal context engine (755), an AI orchestrator (760), and an AI model set (765).
[0335] Meanwhile, the multimodal context engine (755) can generate multimodal context data based on the vehicle situation information received from the recognizer (720).
[0336] For example, multimodal context data may include at least one of text data, image data, or audio data describing a vehicle situation generated based on vehicle situation information.
[0337] Meanwhile, the multimodal context engine (755) may include a multimodal signal adapter (751), a multimodal context buffer (752), a multimodal context retriever (814), etc.
[0338] The multimodal signal adapter (751) can receive an object (743b) corresponding to the passenger's gaze position, a plurality of objects (735), a gaze position (743), and location information (737).
[0339] Additionally, the multimodal signal adapter (751) can control the synchronization of received data, such as an object (743b) corresponding to the passenger's gaze position, a plurality of objects (735), a gaze position (743), and location information (737), and convert it into synchronized vector data to be stored in a database within the multimodal context buffer (752).
[0340] In particular, the multimodal signal adapter (751) can be controlled to store synchronized vector data for a predetermined period (e.g., 10 seconds) in a database within the multimodal context buffer (752).
[0341] Accordingly, context analysis before and after the receipt of a voice recognition-based confirmation request becomes possible.
[0342] The multimodal context buffer (752) can store synchronized vector data.
[0343] The multimodal context retriever (814) can receive a voice recognition-based confirmation request from the AI orchestrator (760).
[0344] When a voice recognition-based confirmation request is received, the multimodal context retriever (814) can search for multimodal processing data related to the voice recognition-based confirmation request through the index of the multimodal context buffer (752) based on the confirmation request, and can select the searched multimodal processing data as a context candidate for the creation of multimodal context data.
[0345] Meanwhile, the multimodal context retriever (814) can receive voice recognition-based confirmation requests and map data from the AI orchestrator (760).
[0346] Meanwhile, when a voice recognition-based confirmation request and map data are received, the multimodal context retriever (814) can search for multimodal processing data related to the voice recognition-based confirmation request through the index of the multimodal context buffer (752) based on the confirmation request or map data, and can select the searched multimodal processing data as a context candidate for the creation of multimodal context data.
[0347] Meanwhile, the multimodal context retriever (814) can reconfigure multimodal processing data selected as a context candidate into multimodal context data having a prompt form that the multimodal LLM (767) can interpret.
[0348] For example, the multimodal context retriever (814) can extract context information corresponding to a voice recognition-based confirmation request from a vector database within the multimodal context buffer (752), and extract attribute information for multiple extracted objects and the context corresponding to the confirmation request.
[0349] And, the multimodal context retriever (814) can transmit, based on the extracted context, an object (743b) corresponding to the passenger's gaze position, multiple objects (735), gaze position (743), location information (737), and map data to the prompt manager (763) in the AI orchestrator (760).
[0350] Meanwhile, the AI orchestrator (760) can receive a voice recognition-based confirmation request from the recognizer (720) and control the generation of multimodal context data based on the confirmation request.
[0351] To this end, the AI orchestrator (760) may include a confirmation request agent (761w), a prompt manager (763), a user attribute database (794), etc.
[0352] The AI orchestrator (760) can transmit the voice recognition-based confirmation request to the multimodal context retriever (814) when the voice recognition-based confirmation request is received from the recognizer (720).
[0353] In particular, when the AI orchestrator (760) receives a voice recognition-based confirmation request from the recognizer (720), it can transmit the voice recognition-based confirmation request and map data to the multimodal context retriever (814).
[0354] Meanwhile, the prompt manager (763) within the AI orchestrator (760) can receive from the multimodal context retriever (814) an object (743b) corresponding to the rider's gaze position, multiple objects (735), a gaze position (743), location information (737), and map data to the prompt manager (763) within the AI orchestrator (760).
[0355] Meanwhile, the prompt manager (763) within the AI orchestrator (760) can generate a confirmation request and a prompt corresponding to the user characteristics by referring to the vector database within the user characteristics database (794), and transmit the generated prompt to the confirmation request agent (761w).
[0356] The confirmation request agent (761w) can control the operation of an on-device AI model (767) in the AI model set (765) or an AI model (769) in the server (400) based on the received prompt.
[0357] At this time, the AI model (769) in the server (400) can share data with the AI model set (765) through the interface (766) in the AI model set (765).
[0358] For example, the confirmation request agent (761w) can distribute the load of the on-device AI model (767) in the AI model set (765) or the AI model (769) in the server (400) based on the received prompt, and control the on-device AI model (767) or the AI model (769) in the server (400) to operate respectively based on the load distribution.
[0359] Meanwhile, the confirmation request agent (761w) can receive an inference result or response result corresponding to the prompt from an on-device AI model (767) or an AI model (769) within the server (400).
[0360] And, the confirmation request agent (761w) can transmit the inference result or response result to the illustrator (780).
[0361] Meanwhile, the illustrator (780) can execute a service, output service information, execute an application, or output application information based on the inference result or response result output from the insightor (750).
[0362] For example, the illustrator (780) can execute a navigation service or a display-related service based on the response result output from the insightor (750).
[0363] Meanwhile, the illustrator (780) may include a multimodal output encoder (781) including a modality manager (795), a visual interface (782), an audio interface (785), etc.
[0364] Meanwhile, the multimodal output encoder (781) can encode the inference result or response result output from the insightor (750) and output the encoded inference result data or response result data to the visual interface (782) or audio interface (785).
[0365] Meanwhile, the visual interface (782) or audio interface (785) may be referred to as an output interface.
[0366] Meanwhile, the visual interface (782) can output response result data output from the multimodal output encoder (781).
[0367] Accordingly, at least one of the plurality of displays (180a to 180c) of FIG. 4 can display an image based on response result data from the visual interface (782).
[0368] Meanwhile, the visual interface (782) can output augmented reality (AR) video data (783) based on response result data or mixed reality (MR) video data (784) based on response result data.
[0369] Meanwhile, the audio interface (785) can output response result data in the form of audio.
[0370] For example, the audio interface (785) can convert text-based response result data into audio data and output the converted audio data (798).
[0371] Accordingly, the audio output unit (185) of FIG. 4 can output a sound corresponding to the response result data from the audio interface (785).
[0372] In summary, FIG. 9, the processor (175) in the signal processing device (170) receives camera data from outside the vehicle, camera data from inside the vehicle, and a voice signal of the passenger, performs object detection based on the camera data from outside the vehicle, detects the passenger's gaze position based on the camera data inside the vehicle, and performs voice recognition based on the passenger's voice signal.
[0373] Meanwhile, the processor (175), when a confirmation request is included in the voice recognition content performed, extracts multimodal context data corresponding to the confirmation request, and selects a point of interest corresponding to the passenger's gaze position based on an inference result or response result obtained based on the extracted multimodal context data.
[0374] Accordingly, it becomes possible to accurately select points of interest in response to confirmation requests. Furthermore, in response to confirmation requests, it becomes possible to select points of interest that match passenger characteristics and provide information.
[0375] Meanwhile, the processor (175) can collect information about the selected point of interest and output the collected information. Accordingly, the point of interest can be accurately selected in response to a confirmation request.
[0376] Meanwhile, the processor (175) within the signal processing device (170) can select a fixed object-based point of interest or a moving object-based point of interest when selecting at least one point of interest. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0377] Meanwhile, fixed objects may include buildings, signboards, roads, etc.
[0378] Meanwhile, moving objects may include vehicles, motorcycles, people, animals, robots, etc.
[0379] Meanwhile, the processor (175) in the signal processing device (170) can vary the point of interest based on the vehicle's driving path or the vehicle's situation information.
[0380] For example, the processor (175) within the signal processing device (170) can control the vehicle's driving path based on road traffic during driving and output point of interest information in response to the driving path of the vehicle being changed.
[0381] As another example, the processor (175) within the signal processing unit (170) can vary the driving path of the vehicle or vary the selected point of interest information based on situational information of the occupant, such as urgent business of the occupant inside the vehicle during driving. Accordingly, the point of interest can be accurately selected based on the situational information of the occupant.
[0382] Meanwhile, the processor (175) in the signal processing device (170) can select a group of points of interest including a plurality of points of interest for vehicle progress.
[0383] Meanwhile, the interest point group may include fixed object-based interest points and moving object-based interest points.
[0384] Specifically, the processor (175) in the signal processing device (170) can combine a first point of interest corresponding to a fixed object and a second point of interest corresponding to a moving object to set a first point of interest group.
[0385] For example, the processor (175) within the signal processing unit (170) can be controlled to output composite point of interest information, such as "There is a cafe in the yellow building where the woman walking with a black umbrella is located," based on first point of interest information corresponding to a yellow building in front of the vehicle and second point of interest information corresponding to a woman walking with a black umbrella on the right side in front of the vehicle. Accordingly, information regarding the points of interest can be accurately output.
[0386] Meanwhile, the interest point group may include road-based interest points and object-based interest points.
[0387] Alternatively, the interest point group may include map data-based interest points and object-based interest points.
[0388] Alternatively, the interest point group may include road-based interest points and object-based interest points.
[0389] Meanwhile, the processor (175) can select a point of interest corresponding to the position of the passenger's gaze at the time when the voice related to the confirmation request is spoken, based on the inference result or response result obtained based on the extracted multimodal context data, and output information about the selected point of interest. Accordingly, it is possible to accurately select a point of interest in response to the confirmation request.
[0390] Specifically, if the confirmation request includes a phrase such as "What's that?", the processor (175) can select an external point of interest corresponding to the passenger's gaze position, collect information about the selected point of interest, and control the output of the collected information.
[0391] For example, based on an external object corresponding to a confirmation request and the passenger's gaze position, a first building on the front right of the vehicle is selected, and if the passenger's characteristic data is focused on restaurant data, a first restaurant within the first building can be selected as a point of interest, and information about the selected first restaurant can be collected and provided.
[0392] As another example, based on an external object corresponding to a confirmation request and the passenger's gaze position, a first building on the front right of the vehicle is selected, and if the passenger's characteristic data is focused on hospital data, a first hospital within the first building can be selected as a point of interest, and information regarding the selected first hospital can be collected and provided.
[0393] As another example, based on an external object corresponding to a confirmation request and the passenger's gaze position, a first vehicle on the front left of the vehicle is selected, and the processor (175) selects the first vehicle as a point of interest when the passenger's characteristic data is focused on cost data, collects purchase information regarding price, displacement, etc. related to cost data among the information about the first vehicle, and can provide the collected purchase information.
[0394] That is, the processor (175) can select a moving point of interest, such as the first vehicle, based on a confirmation request and the passenger's gaze position, and can control to output information about the selected moving point of interest.
[0395] Meanwhile, when a moving object is selected as a point of interest, the processor (175) may collect information about the moving object through additional analysis or detailed analysis within the acquired image data in addition to collecting information from an external server (400).
[0396] Meanwhile, the processor (175) can select a fixed object as a point of interest, such as the first building, based on a confirmation request and the occupant's line of sight position, and can control to output information about the selected fixed object.
[0397] Meanwhile, the processor (175) can collect information about a selected point of interest through an external server (400) and output the collected information about the selected point of interest as image data or audio data. Accordingly, it is possible to output information with improved accuracy in response to a verification request.
[0398] For example, the processor (175) can control the information about the collected points of interest to be displayed on the display (180).
[0399] As another example, the processor (175) can control the information about the collected point of interest to be output through the audio output unit (185).
[0400] Meanwhile, the processor (175) can detect multiple objects based on camera data from outside the vehicle and recognize the class and attributes of each object. Accordingly, it is possible to output information with improved accuracy in response to a verification request.
[0401] Meanwhile, the processor (175) can recognize the gaze vector information of the occupant based on camera data inside the vehicle. Accordingly, it is possible to output information with improved accuracy in response to a verification request.
[0402] Meanwhile, the processor (175) detects multiple objects based on camera data outside the vehicle, recognizes the gaze vector information of the occupant based on camera data inside the vehicle, and can select an object among the multiple objects that corresponds to the gaze vector information as a point of interest. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0403] Meanwhile, the processor (175) receives location information from the location sensor (198) and can control the synchronization of the location information with the object detected based on camera data outside the vehicle and the occupant's gaze vector information recognized based on camera data inside the vehicle, so that the location information is stored in a database within the multimodal context buffer (752). Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0404] Meanwhile, the processor (175) performs voice recognition based on the passenger's voice signal, extracts context information corresponding to a confirmation request from the database within the multimodal context buffer (752) among the performed voice recognition contents, and can extract context data corresponding to a confirmation request based on the context information. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0405] Meanwhile, the processor (175) can generate a prompt corresponding to the occupant characteristics based on context data including context information, a portion of image data, and map data.
[0406] And, the processor (175) can select an object corresponding to the occupant's gaze position based on the inference result or response result obtained based on the prompt, and output information about the selected object. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0407] Meanwhile, the processor (175) can generate a prompt corresponding to the occupant characteristics based on context data including context information, a portion of image data, and map data.
[0408] And, the processor (175) controls the transmission of the generated prompt to an external server (400), receives an inference result or a response result from the external server (400), selects a point of interest corresponding to the passenger's gaze position based on the received inference result or response result, and outputs information about the selected point of interest. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0409] Meanwhile, the signal processing device (170) according to an embodiment of the present disclosure may further include a neural processor (179).
[0410] Meanwhile, the processor (175) can generate a prompt corresponding to the occupant characteristics based on context data including context information, part of image data, and map data, and transmit the prompt to the neural processor (179).
[0411] The neural processor (179) can output an inference result or a response result based on a prompt from the processor (175).
[0412] Meanwhile, the processor (175) receives an inference result or response result based on a prompt generated by the neural processor (179), selects a point of interest corresponding to the passenger's gaze position based on the received inference result or response result, and outputs information about the selected point of interest. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0413] Meanwhile, the processor (175) can run a hypervisor (505) as in FIG. 5 and run a display virtualization machine (530) for a display on the hypervisor (505).
[0414] Meanwhile, the display virtualization machine (530) can select a point of interest corresponding to the passenger's gaze position based on an inference result or response result obtained based on the extracted context data, and output information about the selected point of interest. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0415] Meanwhile, the processor (175) may further run a gateway virtualization machine (510) for gateway operation on the hypervisor (505) as shown in FIG. 5.
[0416] Meanwhile, the gateway virtualization machine (510) receives camera data from outside the vehicle, camera data from inside the vehicle, and a voice signal of the passenger, and can transmit the camera data from outside the vehicle, camera data from inside the vehicle, and the voice signal of the passenger to the display virtualization machine (530) through the shared memory (508) within the hypervisor (505). Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request.
[0417] Meanwhile, the processor (175) can further run a driving control virtualization machine (520) on the hypervisor (505) as shown in FIG. 5.
[0418] Meanwhile, the driving control virtualization machine (520) receives camera data outside the vehicle, camera data inside the vehicle, and a passenger's voice signal through a shared memory (508) within the hypervisor (505), and can execute a driving control service or a driving control application based on the camera data outside the vehicle, camera data inside the vehicle, and the passenger's voice signal. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0419] FIG. 10a illustrates setting the search range based on the occupant's line of sight.
[0420] Referring to the drawing, the processor (175) can identify an object placed within a predetermined distance or predetermined range based on the passenger's line of sight when a confirmation request such as "What's that?" is received, and provide an Application Programming Interface (API) for executing an AI model (767 or 769) based on the identified object.
[0421] The drawing illustrates that the angle between the vehicle driving direction (DRm) and the first direction reference is θα, and the driver's (1020) line of sight angle is θβ based on the in-vehicle image from the in-vehicle camera (195i) inside the vehicle (200).
[0422] The processor (175) can calculate a gaze vector (DRe) based on the angle (θα) between the vehicle driving direction (DRm) and the first direction reference and the gaze angle (θβ) of the driver (1020).
[0423] And, the processor (175) can set a search range (FOV) corresponding to a predetermined angle (θm) based on the view vector (DRe).
[0424] For example, the processor (175) can control the field of view (FOV) to become smaller as the vehicle's speed increases. Accordingly, the field of view can be adaptedly adjusted according to the vehicle's speed.
[0425] In the drawing, the field of view (FOV) is illustrated as being cone-shaped relative to the driver (1020).
[0426] Meanwhile, the processor (175) can extract multiple objects (OBJma, OBJmb) within a predetermined search distance (RAS).
[0427] Meanwhile, the processor (175) can control the search distance (RAS) to increase as the vehicle speed increases. Accordingly, the search distance (RAS) can be adaptively adjusted according to the vehicle speed.
[0428] Meanwhile, the processor (175) can extract multiple map tiles within the search range (FOV).
[0429] Meanwhile, the processor (175) can calculate the distance between the location of a driver (1020), which is an example of a passenger, and a plurality of objects (OBJma, OBJmb).
[0430] Meanwhile, the processor (175) can set at least one of the plurality of objects (OBJma, OBJmb) as a point of interest and collect name information, address information, distance information, etc. of the point of interest.
[0431] At this time, the processor (175) can collect information through image signal processing for multiple objects (OBJma, OBJmb) or collect information through an external server (400).
[0432] FIG. 10b is a diagram illustrating the search distance (RAS) within the map data (1717).
[0433] Referring to the drawing, meanwhile, the processor (175) can detect an object within a predetermined search distance (RAS) of the map data (1717).
[0434] Meanwhile, the processor (175) controls the search distance (RAS) to increase as the vehicle speed increases and to decrease as the vehicle speed decreases, thereby enabling the search distance (RAS) to be adaptively adjusted according to the vehicle speed.
[0435] FIG. 10c is a diagram illustrating multiple map tiles within a field of view (FOV).
[0436] Referring to the drawing, the processor (175) can set a plurality of map tiles (720) within the search range (FOV).
[0437] In the drawing, the shape of each map tile is square, but unlike this, various shapes such as circles, hexagons, and pentagons are possible.
[0438] And, the processor (175) can set a search range (FOV) corresponding to a predetermined angle (θm) based on the view vector (VTm).
[0439] That is, the processor (175) can set the field of view (FOV) between the first line (LNa) and the second line (Lnb) based on the line of view vector (VTm). Meanwhile, the angle between the first line (LNa) and the second line (Lnb) may be θm.
[0440] Meanwhile, the processor (175) can detect an object or point of interest within the search range (FOV) and within the search distance (RAS).
[0441] In the drawing, there are six objects or points of interest within the field of view (FOV) and within the search distance (RAS).
[0442] Meanwhile, the processor (175) identifies multiple map tiles that overlap with the area when the gaze search area of the driver (1020), which is an example of a passenger, is specified.
[0443] For example, if the map tile at the default zoom level of the procedural modeling of the 3D map contains data attributes of a point of interest, the processor (175) can use that zoom level as is and reuse the map tile data loaded into memory (140) for modeling. Accordingly, performance can be optimized.
[0444] Meanwhile, the processor (175) can perform additional separate exploration based on the changed zoom level if the 3D landmark zoom level is different from the exploration zoom level of the point of interest.
[0445] Meanwhile, the processor (175) optimizes performance by reusing 3D landmarks when they are already being used for modeling the 3D map or are already loaded into memory (140).
[0446] Meanwhile, the processor (175) can select at least one of the objects as a point of interest based on the occupant's characteristic data when multiple objects are located in the occupant's line of sight. This is described with reference to FIG. 10d.
[0447] Fig. 10d illustrates selecting an object hidden from the line of sight as a point of interest.
[0448] Referring to the drawing, the processor (175) can detect the line of sight of a driver (1020), which is an example of a passenger, and detect a plurality of objects (BDma, BDmb, BDmc) placed on the line of sight.
[0449] The processor (175) can detect a first object (BDma) based on front image data from the front camera (195).
[0450] Meanwhile, the processor (175) can detect a second object (BDmb) and a third object (BDmc) that are hidden and positioned behind the first object (BDma) based on map data stored in memory (140) and location information.
[0451] And, the processor (175) detects the line of sight of a driver (1020), which is an example of a passenger, and among a plurality of objects (BDma, BDmb, BDmc) placed on the line of sight, it can select at least one of the plurality of objects as a point of interest based on characteristic data of the driver (1020), which is an example of a passenger.
[0452] In the drawing, the processor (175) is shown selecting a second object (BDmb) among a plurality of objects (BDma, BDmb, BDmc) as the first point of interest (PTma) and selecting a third object (BDmc) as the second point of interest (PTmb). Accordingly, the point of interest can be accurately selected.
[0453] In particular, when a confirmation request such as "What's that?" is received, the processor (175) can select a second object (BDmb) among a plurality of objects (BDma, BDmb, BDmc) as a first point of interest (PTma) and select a third object (BDmc) as a second point of interest (PTmb) based on the line of sight of a driver (1020), which is an example of a passenger. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0454] Meanwhile, the processor (175) can perform a test to determine whether multiple objects (BDma, BDmb, BDmc) are obscured.
[0455] For example, the processor (175) can determine whether each object is obscured by the height difference of each object within the line of sight of a driver (1020), which is an example of a passenger, based on the height information of each object (BDma, BDmb, BDmc).
[0456] The processor (175) can determine that the second object (BDmb) is obscured when the height of the first object (BDma) that is closer, as shown in the drawing, is greater than the height of the second object (BDmb) that is further away.
[0457] Meanwhile, the processor (175) determines whether the first object (BDma) is matched based on the occupant's characteristic data, and if it is not matched, determines whether the second object (BDmb) is matched, and if it is matched, can select it as the first point of interest (PTma).
[0458] Similarly, the processor (175) can determine whether the third object (BDmc) is a match based on the occupant's characteristic data, and if it is a match, select it as the second point of interest (PTmb).
[0459] Meanwhile, the processor (175) can select a fixed object placed behind the moving object as a point of interest based on map data when the moving object is located in the line of sight of the occupant. This is described with reference to FIG. 10e.
[0460] FIG. 10e illustrates a moving object positioned in the occupant's line of sight.
[0461] Referring to the drawing, the processor (175) can perform object detection based on the front camera image (1730) and detect the occupant's line of sight (VTm) based on the interior camera (195i).
[0462] In the drawing, the angle between the occupant's line of sight (VTm) and the vehicle's direction of travel (Drm) is θβ.
[0463] Meanwhile, the processor (175) can detect a moving object (TRa) based on the front camera image (1730).
[0464] Meanwhile, the processor (175) can select a fixed object placed behind the moving object as a point of interest based on map data when the moving object (TRa) is located in the line of sight of the occupant.
[0465] FIG. 10f illustrates selecting a fixed object behind the moving object (TRa) of FIG. 10e as the point of interest.
[0466] Referring to the drawing, the processor (175) can obtain an inference result or a response result based on multimodal context data including the moving object when the moving object is located in the occupant's line of sight (VTm), and can select a point of interest based on the obtained inference result or response result.
[0467] In particular, when a confirmation request such as "What's that?" is received, the processor (175) detects an object based on the line of sight (VTm) of a driver (1020), which is an example of a passenger, and when a moving object (TRa) is detected, it can select a fixed object placed behind the moving object (TRa) as a point of interest based on map data in memory (140).
[0468] In the drawing, an example is provided of selecting the building (BDmb) behind the moving object (TRa) as the first point of interest (PTma), and then selecting the next building (BDmc) as the second point of interest (PTmb).
[0469] Meanwhile, the processor (175) may ultimately select a second point of interest (PTmb) rather than a first point of interest (PTma) based on the passenger's characteristic data.
[0470] Meanwhile, the processor (175) can collect information about a selected point of interest from an external server (400) and output the collected information.
[0471] Meanwhile, the processor (175) can select the fixed object (BDmc) as a point of interest (PTmb) when the fixed object (BDmc) is located in the passenger's line of sight (VTm), collect information about the point of interest (PTmb) based on map data stored in memory (140), and output the collected information. Accordingly, the point of interest can be accurately selected in response to a confirmation request.
[0472] Meanwhile, the processor (175) can select the fixed object (BDmc) as the point of interest (PTmb) when the fixed object (BDmc) is located in the passenger's line of sight (VTm), and if there is no information about the point of interest (PTmb) in the map data stored in memory (140), collect information about the selected point of interest (PTmb) from an external server (400) and output the collected information. Accordingly, the point of interest can be accurately selected in response to a confirmation request.
[0473] Figure 11a illustrates an example of a vehicle front image.
[0474] Referring to the drawing, the processor (175) can detect an object in the vehicle front image (1735) based on the occupant's line of sight (VTm).
[0475] For example, the processor (175) can detect a first object (1738) based on a passenger's line of sight (VTm) that forms a predetermined angle (θβ) with the vehicle's driving direction (Drm).
[0476] Meanwhile, the processor (175) may also detect a second object (1736) corresponding to the vehicle driving direction (Drm).
[0477] Meanwhile, the processor (175) can perform segmentation for object detection.
[0478] FIG. 11b illustrates a segmented image (1740) in which segmentation has been performed on the vehicle front image (1735) of FIG. 11a.
[0479] Referring to the drawing, the processor (175) can perform segmentation on the vehicle front image (1735) to obtain a segmentation image (1740).
[0480] Meanwhile, the segmentation image (1740) may include a first segmentation object (1748) corresponding to the first object (1738) and a second segmentation object (1746) corresponding to the second object (1736).
[0481] FIG. 11c is a diagram illustrating separate signal processing for moving objects or stationary objects.
[0482] Referring to the drawing, the processor (175) can determine whether the object the occupant is looking at is the first object (1738) or the building behind the first object (1738) based on the first segmentation object (1748) of FIG. 11b.
[0483] For example, the processor (175) can output a question message to determine whether the object the occupant is looking at is the first object (1738) or the building behind the first object (1738) based on the first segmentation object (1748) of FIG. 11b, and can determine whether the object the occupant is looking at is the first object (1738) or the building behind the first object (1738) based on the response message of the occupant (1020) corresponding to the question message.
[0484] Meanwhile, the processor (175) can control the execution of a multimodal LLM (765) for the analysis of a response message when the object being watched by the occupant is a moving object such as a vehicle, a first object (1738), or a person.
[0485] That is, the processor (175) can control the execution of an AI model set (765) for the analysis of a response message when the object being watched by the occupant is a moving object such as a vehicle, a first object (1738), or a person.
[0486] And, the processor (175) can select the first object (1738) as the object that the passenger is watching, based on the inference result or response result from the AI model set (765). Accordingly, the point of interest can be accurately selected.
[0487] As another example, the processor (175) can check whether a point of interest registration (1745) is set in the map data to check whether the object being watched by the occupant is a building behind the first object (1738).
[0488] For example, if the processor (175) has a point of interest (1745) registered in the map data for a building behind the first object (1738), the processor (175) can select the building behind the first object (1738), rather than the first object (1738), as the object that the occupant is watching. Accordingly, the point of interest can be accurately selected.
[0489] FIG. 12a illustrates the location of multiple points of interest within a building.
[0490] Referring to the drawing, when multiple points of interest (PTna, PTnb, PTnc, PTnd) are located corresponding to the gaze position of the occupant (1020), the processor (175) can select at least one point of interest among the multiple points of interest (PTna, PTnb, PTnc, PTnd) based on the occupant's characteristic data, collect information about the selected point of interest, and output the collected information. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0491] For example, the processor (175) can select at least one point of interest (PTna, PTnb, PTnc, PTnd) among the points of interest (PTna, PTnb, PTnc, PTnd) based on the occupant's characteristic data, when the building object (BDna) is located on the occupant's (1020) line of sight (VTn) and multiple points of interest (PTna, PTnb, PTnc, PTnd) are located within the building object (BDna) in the map data.
[0492] For example, if the passenger characteristic data is focused on a cafe, the processor (175) can select a cafe (PTnd) within a building object (BDna) as a point of interest and collect and provide information about the selected cafe (PTnd). Accordingly, it is possible to select a point of interest that matches the passenger characteristics and provide information.
[0493] As another example, the processor (175) can select a hospital (PTnb) within a building object (BDna) as a point of interest when the passenger characteristic data is focused on a hospital, and collect and provide information about the selected hospital (PTnb). Accordingly, it is possible to select a point of interest that matches the passenger characteristics and provide information.
[0494] Meanwhile, the processor (175) can adjust the prompt to inquire which of the multiple points of interest to provide information through the AI model set (765) when multiple points of interest are identified based on the passenger's characteristic data or the passenger's preference data. Accordingly, it is possible to select a point of interest that matches the passenger's characteristics and provide information.
[0495] FIG. 12b illustrates the identification of one of the multiple points of interest within the building of FIG. 12a.
[0496] Referring to the drawing, the processor (175) can detect multiple objects (1752, 1753, 1754) among the vehicle front image (1750).
[0497] Meanwhile, the processor (175) can select a building object (1754) based on a passenger's gaze vector (VTn) that forms a predetermined angle (θβ) with the vehicle's direction of travel (DRn).
[0498] And, the processor (175) can crop a portion of the image (1755) of the building object (1754) corresponding to the line of sight vector (VTn) to obtain a cropped image (1760).
[0499] And, the processor (175) can recognize text or brand logos within an image (1760) through a multimodal AO agent (1777) or an AI model set (765), and perform secondary filtering (1774) based on the recognized text or brand logos (1772).
[0500] Accordingly, the processor (175) is able to distinguish multiple points of interest (PTna, PTnb, PTnc, PTnd) within the building object (1754).
[0501] Figure 13a illustrates selecting a moving object as a point of interest.
[0502] Referring to the drawing, the processor (175) can obtain a vehicle front image (1820) from a vehicle camera (195mA, 195mb, 195mc).
[0503] Meanwhile, the processor (175) can detect an object placed within a predetermined distance or predetermined range based on the passenger's line of sight when a confirmation request (1810) such as "What's that?" is received.
[0504] For example, the processor (175) can detect that the passenger's gaze is located on the vehicle ahead (1825) within the vehicle front image (1820).
[0505] Figure 13b is a drawing referenced in the description of Figure 13a.
[0506] Referring to the drawing, the processor (175) receives a voice signal corresponding to the passenger's utterance (1842), such as "What's that?"
[0507] Next, the processor (175) can extract the occupant's line of sight (1845) from the internal camera (195i).
[0508] And, the processor (175) can extract the passenger's gaze information (1843) based on the passenger's gaze angle (1845).
[0509] Next, the processor (175) can generate a gaze image or gaze information (1852) based on the vehicle front image (1850) and gaze information (1843) or gaze angle (1845).
[0510] The processor (175) can crop or extract the vehicle front image (1850) into a specific range based on the gaze image or gaze information (1852) to generate the gaze image or gaze information (1854) of the occupant.
[0511] Next, the processor (175) can control the execution of an object classification / recognition AI model (1856) based on the passenger's gaze image or gaze information (1854).
[0512] The processor (175) determines (1864) whether the identified object is a dynamic point of interest or a static point of interest based on the result of the object identification AI model (1856), and if it is a dynamic point of interest, controls the execution of the AI model set (765).
[0513] At this time, the AI model set (765) can be executed based on the gaze object identification information generated from the object identification AI model (1856).
[0514] And, the processor (175) can generate speech-based descriptive text data (1869) for dynamic points of interest based on the inference results or response results of the AI model set (765).
[0515] Meanwhile, the processor (175) can, based on the result of the object identification AI model (1856), search for point of interest information (1867) within the map data if the identified object is a static point of interest, and control the execution of the AI model set (765) based on the searched point of interest information (1867).
[0516] And, the processor (175) can generate colloquial-based descriptive text data (1869) for static points of interest based on the inference results or response results of the AI model set (765).
[0517] Fig. 13c illustrates selecting a fixed object as a point of interest.
[0518] Referring to the drawing, the processor (175) can obtain an image of the vehicle exterior from the vehicle camera (195mA, 195mb, 195mc).
[0519] For example, the processor (175) can acquire an image of the vehicle's exterior based on a occupant's gaze vector (VTna) that forms a predetermined angle (θn) with the vehicle's direction of travel (DRn).
[0520] Meanwhile, the processor (175) can select at least one of the plurality of vehicle cameras (195ma, 195mb, 195mc) based on the angle of view information (190) of the plurality of vehicle cameras (195ma, 195mb, 195mc) and the occupant's gaze vector (VTna), and acquire an image (1912) from the selected camera.
[0521] Meanwhile, the processor (175) can calculate a gaze point based on the horizontal angle of the passenger's gaze and the vertical angle of the passenger's gaze.
[0522] Meanwhile, the processor (175) can perform semantic segmentation based on a specific size window or AI model based on the gaze point.
[0523] Meanwhile, the processor (175) can calculate the coordinates required for Image Cropping (1914) based on the coordinate values of the result of a specific size or semantic segmentation, and perform Image Cropping (1914) based on the calculated coordinates.
[0524] That is, the processor (175) can generate a cropped image (1916) based on the calculated coordinates and calculate the average depth (1920) of the image.
[0525] FIG. 13d is a drawing referenced in the description of FIG. 13c.
[0526] Referring to the drawing, the occupant's line of sight (VTnc) may be lower than the horizon (LVN).
[0527] That is, the occupant's line of sight (VTnc) can be lower than the horizon (LVN) by θn.
[0528] The processor (175) can select a vehicle object (1945) within a camera image (1940) as a point of interest (PTnc) in response to the passenger's gaze.
[0529] FIG. 14a illustrates an example of a method for identifying objects at a dynamic point of interest.
[0530] Referring to the drawing, the processor (175) can control the execution of an object identification artificial intelligence model (765) based on a cropped image (1940) corresponding to the passenger's line of sight among the vehicle exterior images.
[0531] The cropped image (1940) at this time may be an image including a vehicle.
[0532] And, the processor (175) can obtain inference result data (1942) including type, category, etc. as inference result data of the object identification artificial intelligence model (765).
[0533] FIG. 14b illustrates an example of a method for identifying objects at a static point of interest.
[0534] Referring to the drawing, the processor (175) can control the execution of an object identification artificial intelligence model (765) based on a cropped image (1950) corresponding to the passenger's line of sight among the vehicle exterior images.
[0535] The cropped image (1950) at this time may be an image of a sign inside a building.
[0536] And, the processor (175) can obtain inference result data (1952) including type, category, etc. as inference result data of the object identification artificial intelligence model (765).
[0537] Figure 15 is a drawing referenced in the explanation of guide point information for vehicle movement.
[0538] Referring to the drawing, the processor (175) in the signal processing unit (170) can recognize the road being driven (DRm) or the intersection (CDm) connected to the road being driven (DRm) based on location data and camera data among the sensing data outside the vehicle.
[0539] Meanwhile, the processor (175) in the signal processing device (170) can control the execution of a navigation service based on the destination data of the vehicle among the sensing data inside the vehicle.
[0540] Meanwhile, the processor (175) in the signal processing device (170) can receive turn-by-turn (TBT) information for each section of the road (DRm) or intersection (CDm) being driven on, based on map data, during the execution of the navigation service.
[0541] For example, the turn-by-turn information for each section may include angle information based on a first direction. The first direction may correspond to the true north direction or the vehicle driving direction.
[0542] Meanwhile, the processor (175) in the signal processing device (170) can calculate the first exit angle (θa) from the road (DRm) to the first branch road (BRa) based on map data, location data, and camera data during the execution of the navigation service.
[0543] At this time, the first exit angle (θa) may be the angle between the road (DRm) currently in motion and the first branch road (BRa).
[0544] Meanwhile, the angle between the second branch road (BRb) and the first branch road (BRa) of the intersection (CDm) may be θb as the second exit angle.
[0545] Meanwhile, the processor (175) within the signal processing device (170) can extract information on multiple points of interest (POI) or buildings or landmarks existing within the search distance (RAS) from map tiles within the map data.
[0546] Meanwhile, the processor (175) in the signal processing device (170) can generate multimodal context data including map data, sensing data inside the vehicle, and sensing data outside the vehicle, and generate a prompt based on the multimodal context data.
[0547] And, the processor (175) in the signal processing unit (170) outputs guide point information for vehicle progress based on the inference result or response result obtained based on the prompt.
[0548] At this time, meanwhile, the processor (175) in the signal processing device (170) can generate a plurality of guide points, select at least one guide point among the plurality of guide points, and output guide information corresponding to the selected guide point.
[0549] For example, a processor (175) within a signal processing device (170) may select the point of interest or building or landmark information closest to the first branch road (BRa) or the first exit angle (θa) of the first branch road (BRa) as guide information.
[0550] Specifically, the processor (175) in the signal processing device (170) can select the first branch road (BRa) or the first building (BDma) located within the first exit angle (θa) of the first branch road (BRa) as guide information.
[0551] And, the processor (175) in the signal processing unit (170) can be controlled to output guide information such as "turn right around the first building (BDma)."
[0552] In this way, instead of providing guidance information that is not intuitive, such as the conventional "turn right 200m ahead," guidance information such as "turn right around Building 1 (BDma)" is output, thereby enabling the accurate selection of a point of interest in response to a confirmation request. Furthermore, in response to a confirmation request, a point of interest that matches the characteristics of the occupant can be selected and information provided.
[0553] As another example, the processor (175) in the signal processing unit (170) can select the first branch road (BRa) or the second building (BDmb) closest to the first branch road (BRa) as guide information.
[0554] And, the processor (175) within the signal processing unit (170) can be controlled to output guide information such as "turn right when the second building (BDmb) is visible." Accordingly, the point of interest can be accurately selected in response to a confirmation request.
[0555] Figure 16a illustrates an example of map data.
[0556] Referring to the drawing, the processor (175) in the signal processing unit (170) can receive map data (1010) while the vehicle is in motion.
[0557] In particular, the processor (175) within the signal processing device (170) can receive map data (1010) stored in memory (140) based on location data, which is an example of sensing data outside the vehicle, while the vehicle is in motion.
[0558] FIG. 16b illustrates an example of context data based on the map data of FIG. 16a.
[0559] Referring to the drawing, the processor (175) in the signal processing device (170) can generate navigation context data (1020) for an AI model based on the map data (1010) of FIG. 16a.
[0560] Context data (1020) may include, as shown in the drawing, a driving road, a building, a building location (1012), a building direction (1014), an original message based on map data, etc.
[0561] FIG. 16c exemplifies a natural language phrase based on the context data of FIG. 16b.
[0562] Referring to the drawing, the processor (175) in the signal processing unit (170) can generate a natural language phrase (1032 or 1035 or 1037) for an AI model based on the context data (1020) of FIG. 16b.
[0563] A processor (175) within a signal processing unit (170) can execute an AI model (767) based on natural language phrases (1032, 1035, 1037) and generate guide point information (1033 or 1036) for vehicle progress based on the execution result of the AI model (767). Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0564] Figure 17 illustrates the generation of voice-based guide information based on camera data.
[0565] Referring to the drawing, the processor (175) in the signal processing device (170) can receive sensing data inside the vehicle and sensing data outside the vehicle.
[0566] In particular, the processor (175) within the signal processing device (170) can receive camera image data (1105), which is an example of sensing data outside the vehicle, from a plurality of cameras (106mA, 195mb, 195mc).
[0567] Meanwhile, the processor (175) can receive turn-by-turn (TBT) information (1109) for each section of the road (DRm) or intersection (CDm) being driven on, based on location data (e.g., GPS data), which is an example of sensing data from outside the vehicle, and map data.
[0568] Meanwhile, the processor (175) in the signal processing device (170) can generate context data (1110) for an AI model based on segment-by-segment turn-by-turn information (1109) and camera image data (1105).
[0569] And, the processor (175) can generate dynamic points of interest (1112) based on context data (1110).
[0570] That is, the processor (175) can generate dynamic points of interest (1112) based on camera image data (1105), location data, and map data.
[0571] Meanwhile, the processor (175) in the signal processing device (170) can perform a point of interest search (1114) based on the turn-by-turn information (1109) for each section.
[0572] And, the processor (175) can generate a static point of interest (1115) based on the point of interest search (1114).
[0573] That is, the processor (175) can generate a static point of interest (1115) based on location data and map data.
[0574] Meanwhile, the dynamic point of interest (1112) may include vehicles, motorcycles, people, animals, robots, etc. That is, the dynamic point of interest (1112) may correspond to the moving objects described above.
[0575] Meanwhile, the static point of interest (1115) may include buildings, signboards, roads, etc. That is, the static point of interest (1115) may correspond to the fixed object described above.
[0576] Next, the processor (175) can generate a natural language-based phrase or prompt based on the dynamic point of interest (1112) and the static point of interest (1115), and control the execution of the AI model (1122) based on the natural language-based phrase or prompt.
[0577] The AI model (1122) at this time may be the above-described on-device AI model (767) or the AI model (769) within the server (400).
[0578] Meanwhile, the processor (175) can generate a 3D navigation graphic effect (1130) based on a dynamic point of interest (1112) and a static point of interest (1115).
[0579] Next, the processor (175) can output text-based guide point information (1124) based on the inference result or response result according to the execution of the AI model (1122).
[0580] And, the processor (175) can convert text-based guide point information (1124) into voice-based guide point information (1126) and output it for a driver who is driving.
[0581] Accordingly, sound corresponding to voice-based guide point information (1126) is output through the audio output unit (185) in the vehicle (200). Accordingly, it is possible to accurately select a point of interest in response to a confirmation request. Furthermore, in response to a confirmation request, it is possible to select a point of interest that matches the characteristics of the occupant and provide information.
[0582] FIG. 18 illustrates an example of voice-based guide point information of FIG. 17.
[0583] Referring to the drawing, the processor (175) can control the navigation screen (1210) to be displayed on the display when the navigation service is executed.
[0584] And, the processor (175) can select a gas station (1216) within the navigation screen (1210) as a guide point (1217) and output voice-based audio guide information (1225) corresponding to the selected guide point (1217).
[0585] That is, as shown in the drawing, natural language-based audio guide information (1225) such as "Go straight, and when you see the sparkling gas station over there, turn right and take Yangcheon-ro" can be output.
[0586] Accordingly, it becomes possible to accurately select points of interest in response to confirmation requests. Furthermore, in response to confirmation requests, it becomes possible to select points of interest that match passenger characteristics and provide information.
[0587] FIG. 19 illustrates the generation of dynamic object-based guide information based on camera data.
[0588] Referring to the drawing, the processor (175) in the signal processing device (170) can receive sensing data inside the vehicle and sensing data outside the vehicle.
[0589] In particular, the processor (175) within the signal processing device (170) can receive camera image data (IMGa), which is an example of sensing data outside the vehicle, from a plurality of cameras (106mA, 195mb, 195mc).
[0590] Meanwhile, the processor (175) can receive turn-by-turn (TBT) information (1107) for each section of the road being driven on, based on location data (e.g., GPS data), which is an example of sensing data from outside the vehicle, and map data.
[0591] Meanwhile, the processor (175) in the signal processing device (170) can generate context data (1110) for an AI model based on segment-by-segment turn-by-turn information (1107) and camera image data (IMGa).
[0592] And, the processor (175) can extract dynamic points of interest or moving objects based on context data (1110).
[0593] For example, dynamic points of interest or moving objects may include people, vehicles ahead, etc., within the camera image (IMGa).
[0594] The processor (175) can select a guide point based on a dynamic point of interest or a moving object. Accordingly, it is possible to accurately select a point of interest in response to a confirmation request.
[0595] For example, the processor (175) can generate context data (1310) based on a dynamic point of interest or a moving object, generate a prompt based on the context data, and, when executing a navigation service, select a guide point based on the inference result obtained based on the prompt. Accordingly, the point of interest can be accurately selected in response to a confirmation request.
[0596] FIGS. 20a to 20d illustrate the selection of guide points based on multiple conditions.
[0597] FIG. 20a illustrates a road (DRm) in motion or an intersection (CDm) connected to the road (DRm) in motion, as in FIG. 15.
[0598] Referring to the drawing, the processor (175) in the signal processing unit (170) can recognize the road being driven (DRm) or the intersection (CDm) connected to the road being driven (DRm) based on location data and camera data among the sensing data outside the vehicle.
[0599] Meanwhile, the processor (175) in the signal processing device (170) can receive turn-by-turn (TBT) information for each section of the road (DRm) or intersection (CDm) being driven on, based on map data, during the execution of the navigation service.
[0600] For example, the turn-by-turn information for each section may include angle information based on a first direction. The first direction may correspond to the true north direction or the vehicle driving direction.
[0601] Meanwhile, the processor (175) in the signal processing device (170) can calculate the first exit angle (θa) from the road (DRm) to the first branch road (BRa) based on map data, location data, and camera data during the execution of the navigation service.
[0602] At this time, the first exit angle (θa) may be the angle between the road (DRm) currently in motion and the first branch road (BRa).
[0603] Meanwhile, the angle between the second branch road (BRb) and the first branch road (BRa) of the intersection (CDm) may be θb as the second exit angle.
[0604] Meanwhile, the processor (175) in the signal processing device (170) can set an arc-shaped search distance (RAS) based on the first exit angle, the second exit angle, the latitude and longitude of the turn-by-turn point, and the search distance.
[0605] Meanwhile, the processor (175) in the signal processing device (170) can search for multiple points of interest, buildings, or landmarks after setting the search distance (RAS).
[0606] For example, the processor (175) in the signal processing device (170) can generate at least one building or landmark that is a plurality of points of interest searched based on the first exit angle as a first condition, as shown in FIG. 20b, as a guide point.
[0607] Next, the processor (175) in the signal processing device (170) can generate at least one building or landmark that is a plurality of points of interest searched based on the second exit angle as a second condition, as shown in FIG. 20c, when the first condition is not met, as a guide point.
[0608] Next, the processor (175) in the signal processing device (170) can create a guide point as a blank space as a third condition, as shown in FIG. 20d, when the first and second conditions are not met.
[0609] Figure 21 is an example of the creation of guide points in a rotation section.
[0610] Referring to the drawing, the guide information generation unit (1120) in the processor (175) in the signal processing unit (170) can receive a plurality of context data (1512, 1514) based on camera data.
[0611] Meanwhile, the processor (175) can generate a static point of interest (1516) based on a search for a point of interest based on camera data.
[0612] That is, the processor (175) can generate a static point of interest (1516) based on location data and map data.
[0613] Meanwhile, the static point of interest (1516) may include buildings, signboards, roads, etc. That is, the static point of interest (1516) may correspond to the fixed object described above.
[0614] Next, the processor (175) can generate a natural language-based phrase or prompt based on a static point of interest (1516) and control the execution of an AI model (1122) based on the natural language-based phrase or prompt.
[0615] The AI model (1122) at this time may be the above-described on-device AI model (767) or the AI model (769) within the server (400).
[0616] Meanwhile, the processor (175) can generate 3D navigation graphic effects (1130) based on static points of interest (1516).
[0617] Next, the processor (175) can output text-based guide point information (1124) based on the inference result or response result (1518) resulting from the execution of the AI model (1122). Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request. Furthermore, it becomes possible to select a point of interest that matches the passenger characteristics and provide information in response to a confirmation request.
[0618] Figure 22 is another example of the generation of guide points in a rotation section.
[0619] Referring to the drawing, the guide information generation unit (1120) in the processor (175) in the signal processing unit (170) can receive a plurality of context data (1612, 1614) based on camera data.
[0620] Meanwhile, the processor (175) can generate dynamic points of interest (1616) based on a search for points of interest based on camera data.
[0621] That is, the processor (175) can generate dynamic points of interest (1616) based on camera data.
[0622] Meanwhile, the dynamic point of interest (1616) may include vehicles, motorcycles, people, animals, robots, etc. That is, the dynamic point of interest (1616) may correspond to the moving objects described above.
[0623] Next, the processor (175) can generate a natural language-based phrase or prompt based on a dynamic point of interest (1616) and control the execution of an AI model (1122) based on the natural language-based phrase or prompt.
[0624] The AI model (1122) at this time may be the above-described on-device AI model (767) or the AI model (769) within the server (400).
[0625] Meanwhile, the processor (175) can generate 3D navigation graphic effects (1130) based on dynamic points of interest (1616).
[0626] At this time, the processor (175) can generate text-based 3D navigation graphic effects (1130) based on a 3D model (1132) and a 3D generation AI model (1134).
[0627] Next, the processor (175) can output text-based guide point information (1124) based on the inference result or response result (1518) resulting from the execution of the AI model (1122). Accordingly, it becomes possible to accurately select a point of interest in response to a confirmation request. Furthermore, it becomes possible to select a point of interest that matches the passenger characteristics and provide information in response to a confirmation request.
[0628] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. Various modifications are possible by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.
Claims
1. A processor that receives camera data from outside the vehicle, camera data from inside the vehicle, and a voice signal from a passenger; and The above processor is, Object detection is performed based on camera data from outside the vehicle, and Based on camera data inside the vehicle, the position of the occupant's gaze is detected, and Voice recognition is performed based on the voice signal of the above-mentioned passenger, and if a confirmation request is included in the content of the voice recognition performed, multimodal context data corresponding to the confirmation request is extracted, and A signal processing device that selects a point of interest corresponding to the gaze position of the occupant based on an inference result or response result obtained based on the extracted multimodal context data.
2. In Paragraph 1, The above processor is, A signal processing device that collects information about the selected point of interest and outputs the collected information.
3. In Paragraph 1, The above processor is, A signal processing device that, based on the inference result or response result obtained based on the extracted multimodal context data, selects a point of interest corresponding to the gaze position of the occupant at the time when the voice related to the confirmation request is spoken, and outputs information about the selected point of interest.
4. In Paragraph 1, The above processor is, A signal processing device that recognizes the gaze vector information of the occupant based on camera data inside the vehicle.
5. In Paragraph 1, The above processor is, Based on camera data from outside the vehicle, a plurality of objects are detected, and Based on camera data inside the vehicle, the gaze vector information of the occupant is recognized, and A signal processing device that selects an object corresponding to the gaze vector information among the plurality of objects as the point of interest.
6. In Paragraph 1, The above processor is, A signal processing device that, when multiple objects are located in the line of sight of the occupant, selects at least one of the multiple objects as the point of interest based on the characteristic data of the occupant.
7. In Paragraph 1, The above processor is, A signal processing device that, when a moving object is located in the line of sight of the aforementioned passenger, selects a fixed object positioned behind the moving object as the point of interest based on map data.
8. In Paragraph 1, The above processor is, When a moving object is located in the line of sight of the aforementioned occupant, A signal processing device that obtains an inference result or a response result based on multimodal context data including the moving object, and selects the point of interest based on the obtained inference result or response result.
9. In Paragraph 1, The above processor is, When a fixed object is located in the line of sight of the aforementioned occupant, A signal processing device that selects the fixed object as the point of interest, collects information about the point of interest based on map data stored in memory, and outputs the collected information.
10. In Paragraph 1, The above processor is, When a fixed object is located in the line of sight of the aforementioned occupant, A signal processing device that selects the above fixed object as the point of interest, and if there is no information about the point of interest in the map data stored in memory, collects information about the selected point of interest from an external server and outputs the collected information.
11. In Paragraph 1, The above processor is, A signal processing device that, when multiple points of interest are located corresponding to the gaze position of the occupant, selects at least one point of interest among the multiple points of interest based on the characteristic data of the occupant, collects information about the selected point of interest, and outputs the collected information.
12. In Paragraph 1, The above processor is, Voice recognition is performed based on the voice signal of the aforementioned passenger, and A signal processing device that extracts context information corresponding to the confirmation request from a database among the voice recognition content performed above, and extracts context data corresponding to the confirmation request based on the context information.
13. In Paragraph 12, The above processor is, Based on the context data including the above context information, a portion of image data, and map data, a prompt corresponding to occupant characteristics is generated, and A signal processing device that selects a point of interest corresponding to the gaze position of the occupant based on an inference result or response result obtained based on the above prompt, and outputs information about the selected point of interest.
14. In Paragraph 12, The above processor is, Based on the context data including the above context information, a portion of image data, and map data, a prompt corresponding to occupant characteristics is generated, and Controls the transmission of the above-mentioned generated prompt to an external server, and A signal processing device that receives an inference result or a response result from the external server, selects a point of interest corresponding to the passenger's gaze position based on the received inference result or response result, and outputs information about the selected point of interest.
15. In Paragraph 12, It further includes a neural processor; and The above processor is, Based on the context data including the above context information, a portion of image data, and map data, a prompt corresponding to occupant characteristics is generated, and A signal processing device that receives an inference result or response result based on a generated projection from the above neural processor, selects a point of interest corresponding to the gaze position of the occupant based on the received acquired inference result or response result, and outputs information about the selected point of interest.
16. In Paragraph 1, The above processor is, Run a hypervisor, and on the hypervisor, run a display virtualization machine for a display, and The above display virtualization machine is, A signal processing device that selects a point of interest corresponding to the gaze position of the occupant based on an inference result or response result obtained based on the extracted context data, and outputs information about the selected point of interest.
17. In Paragraph 16, The above processor is, On the above hypervisor, a gateway virtualization machine for gateway operation is further run, and The above gateway virtual machine is, Receiving camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of the occupant, A signal processing device that transmits camera data outside the vehicle, camera data inside the vehicle, and voice signals of the occupant to the display virtualization machine through shared memory within the hypervisor.
18. In Paragraph 17, The above processor is, On the above hypervisor, further run a driving control virtual machine, and The above driving control virtualization machine is, Through the shared memory within the hypervisor, camera data from outside the vehicle, camera data from inside the vehicle, and voice signals of the occupant are received, and A signal processing device that executes a driving control service or a driving control application based on camera data outside the vehicle, camera data inside the vehicle, and a voice signal of the occupant.
19. A vehicle control device comprising a signal processing device according to any one of paragraphs 1 through 18.