Image generation system and image generation method

The system addresses the lack of vehicle inclusion in captured images by superimposing vehicle images onto external views, ensuring accurate integration and user satisfaction.

JP2026135742APending Publication Date: 2026-08-25NISSAN MOTOR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025021443
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing image generation systems fail to include the vehicle used by the user in the captured images, failing to meet user requirements.

Method used

The system generates an image by superimposing an image of the vehicle onto an image captured outside the vehicle, focusing on objects of interest to the user, using a virtual viewpoint to maintain orientation and positional accuracy.

Benefits of technology

Enables the generation of images that include the vehicle, allowing users to save and view images with the vehicle integrated, enhancing user experience and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026135742000001_ABST
    Figure 2026135742000001_ABST
Patent Text Reader

Abstract

This invention provides an image generation system and method that can generate images that include vehicles used by the user. [Solution] A new image is generated by superimposing an image of the vehicle 10 used by the user onto an image of the direction D in which an object outside the vehicle of interest to the user exists.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image generation system and an image generation method.

Background Art

[0002] There is known a vehicle system that acquires and synchronizes a video captured inside a vehicle cabin and a video captured around the vehicle, temporarily stores the synchronized video, and stores the synchronized video in a storage device when a trigger event is detected (Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the above prior art, since the saved video does not include the vehicle used by the user, there is a problem that it cannot meet the user's requirement to save a video including the vehicle used by oneself.

[0005] The problem to be solved by the present invention is to provide an image generation system and an image generation method capable of generating an image including the vehicle used by the user.

Means for Solving the Problems

[0006] The present invention solves the above problems by generating a new image by superimposing an image of the vehicle used by the user on an image captured in the direction where there is an object outside the vehicle that the user is interested in.

Effects of the Invention

[0007] According to the present invention, an image including the vehicle used by the user can be generated. [Brief explanation of the drawing]

[0008] [Figure 1] A block diagram showing one embodiment of the image generation system according to the present invention. [Figure 2] Figure 1 is a plan view showing an example of a driving scene in which an image is generated by the image generation system. [Figure 3] This figure shows an example of an image captured using the image capture method shown in Figure 1 during the driving scene shown in Figure 2. [Figure 4] This figure shows an example of an image generated during the driving scene in Figure 2. [Figure 5] This figure shows another example of an image generated in the driving scene shown in Figure 2. [Figure 6] Figure 1 is a plan view showing an example of a virtual viewpoint set by the image generation system. [Figure 7] Figure 1 is a rear view showing an example of a virtual viewpoint set by the image generation system. [Figure 8] Figure 1 is a flowchart showing an example of the processing procedure in the image generation system. [Figure 9] This flowchart shows another example of the processing procedure in the image generation system shown in Figure 1. [Figure 10] This flowchart shows yet another example of the processing procedure in the image generation system shown in Figure 1. [Modes for carrying out the invention]

[0009] Embodiments of the present invention will be described below with reference to the drawings. In the following description, the front side of the drive source (e.g., engine) mounted on the vehicle is defined as the front side, and viewing the vehicle from the front side of the drive source is referred to as a front view of the vehicle. Similarly, viewing the vehicle from the rear is referred to as a rear view of the vehicle, viewing the vehicle from the right or left side is referred to as a side view of the vehicle, viewing the vehicle from above is referred to as a top view of the vehicle, and viewing the vehicle from below is referred to as a bottom view of the vehicle.

[0010] [Configuration of the image generation system] Figure 1 is a block diagram showing one embodiment of the image generation system according to the present invention (hereinafter also referred to as this embodiment). The image generation system generates an image including a vehicle used by a user (hereinafter also simply referred to as the vehicle), and displays the generated image on the vehicle's display device or the user's mobile terminal. The user of the vehicle (hereinafter also simply referred to as the user) is not particularly limited as long as the vehicle can be used appropriately, and includes the occupants of the vehicle.

[0011] As shown in Figure 1, the image generation system 1 comprises a vehicle 10, a server 20, and a terminal 30. These devices constituting the image generation system 1 can exchange information with each other via a network NW. The network NW refers to a telecommunications network for exchanging information between the vehicle 10, the server 20, and the terminal 30, and includes the Internet, Ethernet, LAN (local area network), etc. Known wired or wireless communication can be used for communication on the network NW, and mobile communication such as fifth-generation mobile communication systems (5G) may also be used.

[0012] Vehicle 10 is equipped with an imaging device 11, a map database 12, a navigation device 13, a vehicle position detection device 14, an in-vehicle sensor 15, a communication device 16, and a control device 17. These devices are connected by a CAN (Controller Area Network) or other in-vehicle LAN and can exchange information with each other. Vehicle 10 is driven by autonomous driving control or by manual driving by the driver.

[0013] Autonomous driving control means autonomously controlling the driving operation of the vehicle 10 using the control device 17. The driving operation includes all driving operations such as acceleration, deceleration, starting, stopping, and steering. Autonomously controlling the driving operation means that the control device 17 controls the driving operation using the in-vehicle devices of the vehicle 10. The control device 17 controls these driving operations within a predetermined range, and for driving operations not controlled by the control device 17, manual operation by the driver of the vehicle 10 is performed. On the other hand, manual driving means that the control device 17 does not perform autonomous control of the driving operation, and the driving operation of the vehicle 10 is controlled by the driver's operation.

[0014] The imaging device 11 is a camera equipped with an imaging element such as a CCD (Charge-Coupled Device), images the objects around the vehicle 10, and generates an image including the objects. In order to suppress the occurrence of blind spots where the objects cannot be imaged, a plurality of the imaging devices 11 are provided on the front grille, side mirrors, rear bumper, etc. of the vehicle 10. Further, the vehicle 10 may be provided with a distance measuring device (not shown). The distance measuring device includes a laser radar, a millimeter wave radar, a LiDAR (Light Detection and Ranging) unit, etc.

[0015] The control device 17 acquires image information from the imaging device 11 and recognizes the objects and the driving environment around the vehicle 10. Further, when recognizing the objects and the driving environment, the control device 17 may acquire the position information of the objects from a distance measuring device (not shown). The acquisition of the image information (and position information) by the control device 17 is executed at a predetermined time interval (for example, every 0.1 to 1 millisecond).

[0016] The objects detected by the imaging device 11 are the road and the objects existing around it. The objects detected by the imaging device 11 include the lane boundary lines, center lines, road markings, median strips, road signs, etc. of the road. In addition, the objects detected by the imaging device 11 also include obstacles that can affect the running of the vehicle 10, such as other automobiles (other vehicles), motorcycles, bicycles, pedestrians, etc. Furthermore, the objects detected by the imaging device 11 also include the objects outside the vehicle that the user can recognize (visually recognize) from the position of the vehicle 10. Such objects include landmarks and symbolic buildings (landmarks), monuments (monuments), tourist attractions (including natural objects such as mountains, rivers, and lakesides), historical sites, restaurants, etc.

[0017] The map database 12 is a storage medium in which map information is stored, and is provided inside or outside the vehicle. The control device 17 acquires map information from the map database 12 as necessary. The map information includes information on nodes corresponding to specific points (such as intersections) on the road where the traveling direction of the vehicle 10 changes, and links corresponding to road sections connecting the nodes. The information of the nodes includes position information, information regarding entry and exit of intersections, etc., and the information of the links includes width, radius of curvature of the road, road traffic regulations, etc. The map information may be high-precision map information including information on structures, signs, traffic lights, etc. around the road in addition to the road information for each lane.

[0018] The navigation device 13 refers to the map information acquired from the map database 12 and generates a driving route (hereinafter, also simply referred to as a route) from the current position of the vehicle 10 detected by the own vehicle position detection device 14 to the destination set by the user. The route includes at least information on the road on which the vehicle 10 travels, the driving lane, and the traveling direction of the vehicle 10, and is displayed linearly, for example. The control device 17 acquires the route calculated by the navigation device 13.

[0019] The vehicle position detection device 14 is a positioning system that detects the current position of the vehicle 10. For example, it calculates the current position of the vehicle 10 from radio waves received from GPS (Global Positioning System) satellites. Alternatively, the vehicle position detection device 14 may estimate the current position of the vehicle 10 from information on vehicle speed obtained from the vehicle's speed sensor (not shown) and information on acceleration obtained from the vehicle's acceleration sensor (not shown), and then calculate the current position of the vehicle 10 by comparing the estimated current position with map information. The control device 17 acquires information on the current position of the vehicle 10 (hereinafter also referred to as current position information) from the vehicle position detection device 14 as needed.

[0020] The in-vehicle sensor 15 is a sensor installed in the passenger compartment (hereinafter also simply referred to as the passenger compartment) of the vehicle 10, which detects the user's state and acquires information related to the user's state (hereinafter also referred to as user information). The user's state includes the user's actions, posture, and speech content. The control device 17 acquires user information from the in-vehicle sensor 15 as needed. The in-vehicle sensor 15 includes, for example, an in-vehicle camera and a microphone.

[0021] An in-vehicle camera is a camera equipped with an image sensor such as a CCD, and it captures images of objects inside the vehicle. The objects detected by the in-vehicle camera are objects inside the vehicle, such as the user, the instrument panel, switches, the steering wheel, and the shift lever. The in-vehicle camera is installed in a position where it can detect the user's state, such as on the top of the vehicle's windshield, the rearview mirror, the roof, or the top of the rear window.

[0022] The microphone acquires sounds from inside the vehicle 10 as audio data. These sounds include the voices of the occupants, sounds output from the vehicle 10's speakers (not shown), and the sounds of the vehicle 10 running. Examples of microphones include stand microphones, close-range microphones, and shotgun microphones. The microphone is installed at an appropriate position within a range that allows for proper detection of sounds inside the vehicle.

[0023] The communication device 16 is a communication interface that supports various communication standards such as wired LAN standards and wireless LAN standards. The communication device 16 is a device for exchanging information between the control device 17 and the network NW, and is not particularly limited as long as it is a device that can communicate with other devices via the network NW.

[0024] The control device 17 controls the movement of the vehicle 10 by controlling and coordinating the devices that make up the vehicle 10, and drives the vehicle 10 to its destination. The control device 17 also performs at least a part of the process of generating an image including the vehicle 10 in response to a user's request. The control device 17 is, for example, a computer and comprises a CPU (Central Processing Unit) which is a processor, a ROM (Read Only Memory) which stores programs, and a RAM (Random Access Memory) which functions as an accessible storage device. The CPU of the control device 17 is the operating circuit that executes the programs stored in the ROM of the control device 17 and realizes the functions of the control device 17. In addition to the CPU, an MPU (Micro Processing Unit), ASIC (Application Specific Integrated Circuit), etc. may be used in conjunction with the CPU.

[0025] Server 20 performs at least part of the process of generating an image including the vehicle 10 in response to a user request. Server 20 is, for example, a computer and, like the control unit 17, is equipped with a CPU, ROM, and RAM. The CPU of Server 20 is the operating circuit that executes the program stored in the ROM of Server 20 and realizes the functions of Server 20. As with the control unit 17, an MPU, ASIC, etc. may be used instead of or in conjunction with the CPU.

[0026] The map database 21, like the map database 12, is a storage medium that stores map information. The map database 21 is connected to the server 20 via Ethernet, LAN, or the like. The server 20 retrieves map information from the map database 21 as needed. The map information stored in the map database 21 may be the same as the map information stored in the map database 12, or it may be different.

[0027] Communication device 22, like communication device 16, is a communication interface that supports various communication standards and is a device for exchanging information between server 20 and network NW. Note that communication device 16 and communication device 22 may be the same device or different devices.

[0028] Terminal 30 is a device that displays images generated by the control device 17 or the server 20 to the user. Terminal 30 includes personal computers, smartphones, PDAs (Personal Digital Assistants), etc. Terminal 30 may also be a wearable device such as a smartwatch or a head-mounted display. Terminal 30 is connected to the server 20 via a network NW and a communication device 22 so as to be able to communicate.

[0029] The image generation system 1 of this embodiment includes a recording mode as one of its image generation modes. The recording mode is a mode in which an image including the vehicle 10 and the scenery surrounding the vehicle 10 is generated and the generated image is saved to a server 20 or the like. When the recording mode is activated (ON state), for example, the control device 17 transmits the image captured by the imaging device 11 to the server 20, the server 20 superimposes the image of the vehicle 10 onto the image received from the control device 17 to generate a new image, and saves the generated image. After the vehicle 10 has finished driving, the user can view the image saved on the server 20 or post the saved image to a social networking service (SNS) from the terminal 30. Note that the image generated in the recording mode may be a time-series image (i.e., a video).

[0030] [Functions of the control unit and server] The control device 17 and the server 20 have an image generation function that generates an image including the vehicle 10. The ROMs of the control device 17 and the server 20 store programs for realizing the image generation function, and the CPUs of the control device 17 and the server 20 execute the programs stored in their respective ROMs to realize the image generation function.

[0031] Figure 2 is a plan view showing an example of a driving scene in which an image including the vehicle 10 is generated by the image generation functions of the control device 17 and the server 20. The X-axis, Y-axis, and Z-axis directions shown in Figure 2 correspond to the longitudinal direction, the vehicle width direction, and the height direction of the vehicle 10, respectively. The road shown in Figure 2 is a two-lane road with lanes L1 and L2. The direction of travel for lane L1 is from left to right in the figure (positive X-axis direction), and the direction of travel for lane L2 is from right to left in the figure (negative X-axis direction).

[0032] In the driving scene shown in Figure 2, vehicle 10 is traveling in the positive X-axis direction at position P1 in lane L1. Building A is located to the left front of vehicle 10, and mountain B is located to the right front of vehicle 10. Mountain B is located relatively far from vehicle 10, and it is assumed that the user of vehicle 10 can see mountain B from vehicle 10. In the driving scene shown in Figure 2, if the user activates the recording mode to generate and record an image including vehicle 10, the control device 17 and server 20 will perform the following processing using the image generation function.

[0033] First, the control device 17 determines whether the user has activated the recording mode (i.e., whether the recording mode is ON). For example, the control device 17 determines whether the user has entered an operation to activate the recording mode into an input device (not shown), such as a touch panel. If it is determined that no operation to activate the recording mode has been entered into the input device (i.e., the recording mode is OFF), the control device 17 repeats the determination of whether the user has activated the recording mode. On the other hand, if it is determined that an operation to activate the recording mode has been entered into the input device (i.e., the recording mode is ON), the control device 17 obtains user information from the in-vehicle sensor 15.

[0034] Furthermore, the server 20 may instruct the vehicle 10 (control device 17) to send information regarding input operations to an input device (not shown) to the server 20, and determine whether the user has activated recording mode based on the information regarding input operations received from the vehicle 10. If it is determined that no operation to activate recording mode has been input to the input device, the server 20 may, for example, terminate processing. On the other hand, if it is determined that an operation to activate recording mode has been input to the input device, the server 20 instructs the vehicle 10 (control device 17) to send user information to the server 20.

[0035] Next, the control device 17 determines from the user information whether the user is showing interest in an object outside the vehicle. A user showing interest in an object means that the user is directing their attention to the object. Specifically, the control device 17 determines that the user is showing interest in an object when the user is facing in the direction of the object, when users are talking about the object, or when the user is pointing at the object. An object outside the vehicle is an object outside the vehicle that the user can recognize (see) from the position of the vehicle 10, and in particular, an object outside the vehicle that the user can recognize from the position of the vehicle 10 and which the user can direct their attention to. Objects outside the vehicle include obstacles, landmark buildings, monuments, and tourist attractions.

[0036] For example, the control device 17 acquires an image including the user's body from the in-car camera, performs pattern matching on the image to detect the position of the user's joints, and detects the user's posture from the positional relationship of each joint. Then, based on the detected posture of the user, the control device 17 determines whether or not the user is pointing at an object outside the vehicle. If it is determined that the user is pointing at an object outside the vehicle, the control device 17 determines that there is an object of interest to the user in the direction the user is pointing, and that the user is interested in the object outside the vehicle. On the other hand, if it is determined that the user is not pointing at an object outside the vehicle, the control device 17 determines that there is no object outside the vehicle of interest to the user (hereinafter also referred to as a specific object), and that the user is not interested in the object outside the vehicle.

[0037] If the control device 17 determines that the user is pointing to an object outside the vehicle, it may obtain map information from the map database 12 and current location information from the vehicle position detection device 14, and determine whether or not there is an object that the user could pay attention to in the direction the user is pointing. If it is determined that there is an object that the user could pay attention to in the direction the user is pointing, the control device 17 determines that there is an object that the user is interested in in the direction the user is pointing. On the other hand, if it is determined that there is no object that the user could pay attention to in the direction the user is pointing, the control device 17 determines that there is no specific object.

[0038] Furthermore, the server 20, like the control device 17, may determine from user information whether or not the user is interested in an object outside the vehicle. For example, the server 20 may cause the vehicle 10 (control device 17) to transmit audio data from inside the vehicle acquired from the microphone (in-vehicle sensor 15) to the server 20, perform frequency analysis on the audio data to extract the user's voice, and perform speech recognition processing such as STT (Speech-to-Text) on the extracted audio data to obtain the content of the user's speech.

[0039] The server 20 instructs the vehicle 10 (control device 17) to transmit its current location information to the server 20, retrieves map information from the map database 21, and determines whether the user's utterance includes the name of an object that exists around the vehicle 10 (or is visible to the user from the vehicle 10's position). If the server 20 determines that the user's utterance includes the name of an object that exists around the vehicle 10, the server 20 determines that the user is interested in an object outside the vehicle. On the other hand, if the server 20 determines that the user's utterance does not include the name of an object that exists around the vehicle 10, the server 20 determines that the user is not interested in an object outside the vehicle.

[0040] If the control device 17 determines that the user is interested in an object outside the vehicle, at least one of the control device 17 and the server 20 will identify the direction in which the specific object is located (hereinafter also referred to as the object direction). Similarly, if the server 20 determines that the user is interested in an object outside the vehicle, at least one of the control device 17 and the server 20 will identify the object direction. The object direction is a direction relative to the position of the vehicle 10, and is the direction in which the specific object is located relative to the vehicle 10 (as seen from the vehicle 10). On the other hand, if it is determined that the user is not interested in an object outside the vehicle, the control device 17 and the server 20 will, for example, acquire (or receive) user information again.

[0041] In the driving scene shown in Figure 2, the user activates the recording mode on the input device (not shown), so the control device 17 determines that the recording mode is ON and acquires audio data from inside the vehicle using the microphone (in-vehicle sensor 15). Next, the control device 17 extracts the user's voice from the audio data inside the vehicle, performs speech recognition processing on the extracted audio data, and obtains the content of the user's speech. Then, the control device 17 acquires map information from the map database 12 and determines whether the user's speech content includes the name of at least one of building A and mountain B. For example, if it is determined that the user's speech content includes the name of mountain B, the control device 17 determines that the user is interested in an object outside the vehicle (i.e., mountain B) and identifies the direction in which mountain B is located relative to the vehicle 10.

[0042] If it is determined that the user is interested in an object outside the vehicle, at least one of the control device 17 and the server 20 will identify the direction of the object based on user information, map information, current location information, route information (hereinafter also referred to as route information), etc. For example, if the control device 17 obtains map information from the map database 12, obtains current location information from the vehicle position detection device 14, and determines that the user's utterance, which is user information, includes the name of a tourist attraction visible from the vehicle 10's current location, it will identify the direction from the vehicle 10's current location toward the named tourist attraction as the direction of the object. As another example, the server 20 causes the vehicle 10 (control device 17) to send image information including the user's body and current location information to the server 20, retrieves map information from the map database 21, detects the user's posture from the image, determines whether the user is pointing to an object outside the vehicle based on the detected user's posture, and if it is determined that the user is pointing to an object outside the vehicle, and there is an object in the direction the user is pointing that the user could direct their attention to, the direction the user is pointing is identified as the direction of the object.

[0043] If the direction of the object is determined by at least one of the control device 17 and the server 20, the control device 17 adjusts the field of view (imaging range) of the imaging device 11 as necessary so that an image of the range including the direction of the object can be acquired. For example, if the field of view of the imaging device 11 does not include the direction of the object, the control device 17 changes the field of view to include the range of the image so that an image of the range including the direction of the object is output. The control device 17 then acquires an image of the direction of the object captured from the vehicle 10 (hereinafter also referred to as the first image). The control device 17 may transmit the first image to the server 20 as necessary.

[0044] Similarly, if the object direction is determined by at least one of the control device 17 and the server 20, the server 20 adjusts the field of view of the vehicle's imaging device 11 as necessary so that it can acquire an image that includes the object direction. For example, if the imaging device 11 is a PTZ (Pan Tilt Zoom) camera and the field of view of the imaging device 11 does not include the object direction, the server 20 instructs the control device 17 to move the lens of the PTZ camera so that the field of view includes the object direction. The server 20 then instructs the control device 17 to transmit the first image to the server 20.

[0045] In the driving scene shown in Figure 2, if the user's speech includes the name of mountain B and it is determined that the user is interested in mountain B, the control device 17 obtains current location information from the vehicle position detection device 14 and identifies the direction D from the vehicle's current position P1 toward mountain B as the direction of the object. The control device 17 adjusts the field of view of the imaging device 11 and obtains an image from the imaging device 11 within the range that includes direction D. An example of a first image captured by the imaging device 11 in the driving scene shown in Figure 2 is shown in Figure 3. The first image I1 shown in Figure 3 includes building A on the left side of the image, lanes L1 and L2 in the center of the image, and mountain B in the upper right of the image. The control device 17 transmits the information of the first image I1 shown in Figure 3 to the server 20 via the communication devices 16 and 22 and the network NW.

[0046] The image captured by the vehicle 10's imaging device 11 does not include the vehicle 10, as shown in the first image I1 in Figure 3. Therefore, when generating an image that includes the vehicle 10 and the surrounding scenery in recording mode, a new image (hereinafter also referred to as the superimposed image) is generated by superimposing an image of the vehicle 10 (hereinafter also referred to as the second image) onto the first image I1 acquired from the imaging device 11. The second image may be an image of the vehicle 10 previously captured by an imaging device different from the imaging device 11, or it may be an image of a two-dimensional or three-dimensional model generated by computer graphics.

[0047] When the second image is superimposed on the first image I1, in order to prevent the user from feeling any discomfort regarding the direction of travel or orientation of the vehicle 10 corresponding to the second image, at least one of the control device 17 and the server 20 sets a virtual viewpoint at a position opposite to the position where the specific object exists relative to the vehicle 10, and generates a second image corresponding to the vehicle 10 as seen from the virtual viewpoint. For example, at least one of the control device 17 and the server 20 sets a virtual viewpoint at a position point-symmetric to the position where the specific object exists relative to the vehicle 10. Alternatively, at least one of the control device 17 and the server 20 may set a virtual viewpoint at a position opposite to the position where the specific object exists relative to the vehicle 10, and at a position on a straight line passing through the vehicle 10 and the specific object.

[0048] If multiple images of the vehicle 10 taken from multiple positions around the vehicle 10 are pre-registered, at least one of the control device 17 and the server 20 selects the image taken from the pre-registered multiple images that is closest to the set virtual viewpoint and outputs it as the second image. Also, if at least one of the control device 17 and the server 20 generates the second image using computer graphics, a two-dimensional or three-dimensional model of the vehicle 10 is pre-registered.

[0049] In the driving scene shown in Figure 2, the server 20 instructs the control device 17 to transmit the information of the first image I1 and the current location information to the server 20, and also obtains map information from the map database 21. Based on the current location information and map information, the server 20 sets a virtual viewpoint Q1 at location P2, which is opposite to the location where mountain B is located relative to location P1, the current location of the vehicle 10. Then, the server 20 generates a second image corresponding to the vehicle 10 as seen from the virtual viewpoint Q1 using computer graphics and superimposes it onto the first image I1.

[0050] At least one of the control device 17 and the server 20 sets the position of the second image in the first image I1 based on the positional relationship between the vehicle 10 and the objects surrounding the vehicle 10 when superimposing the second image on the first image I1. For example, at least one of the control device 17 and the server 20 sets the position of the second image in the first image I1 such that the positional relationship between the vehicle 10 and the objects surrounding the vehicle 10 in the superimposed image corresponds to the actual positional relationship between the vehicle 10 and the objects surrounding the vehicle 10. The positional relationship between the vehicle 10 and the objects surrounding the vehicle 10 is, for example, the distance from the vehicle 10 to the objects surrounding the vehicle 10, and the direction in which the objects surrounding the vehicle 10 exist relative to the vehicle 10. Furthermore, if an external object of interest to the user can be identified, at least one of the control device 17 and the server 20 may set the position of the second image in the first image I1 based on the positional relationship between the vehicle 10 and the identified object.

[0051] Figure 4 shows an example of the second image and superimposed image generated by the server 20 in the driving scene shown in Figure 2. The superimposed image I3 shown in Figure 4 is an image in which the second image I2, generated by computer graphics, is superimposed on the first image I1. The second image I2 corresponds to the vehicle 10 as seen from the virtual viewpoint Q1 shown in Figure 2. The position of the second image I2 in the superimposed image I3 (first image I1) is set based on the positional relationship between the vehicle 10 and building A. That is, the second image I2 is superimposed at a position where the positional relationship between the vehicle 10 and building A in the superimposed image I3 corresponds to the actual positional relationship between the vehicle 10 and building A. The server 20 may enlarge or reduce the second image I2, or rotate the second image I2, in order to prevent the orientation or size of the second image I2 in the superimposed image I3 from causing discomfort to the user.

[0052] The server 20 transmits, for example, the information of the superimposed image I3 shown in Figure 4, along with route information, to the terminal 30 via the communication device 22 and the network NW. The terminal 30 displays the superimposed image I3 superimposed on a linear image showing the route for the user. The position where the superimposed image I3 is displayed is, for example, the position in the linear image showing the route that corresponds to the position where the first image I1 used to generate the superimposed image I3 was captured.

[0053] At least one of the control device 17 and the server 20 may determine the direction of the object based on the user's gestures. User gestures include the user's body language, hand movements, and mannerisms, and include the user's posture. For example, in the driving scene shown in Figure 2, if the user's face is facing towards building A, the control device 17 will determine the direction of the object to be the left front of the vehicle 10 where building A is located.

[0054] At least one of the control device 17 and the server 20 may determine the direction of the object based on the user's speech content and map information. For example, in the driving scene shown in Figure 2, if it is determined that the user's speech content includes the name of building A, the control device 17 will determine the direction in which building A is located relative to the vehicle's position P1 (i.e., the left front direction) as the direction of the object.

[0055] At least one of the control device 17 and the server 20 may identify an external object of interest to the user based on user information and map information. For example, in the driving scene shown in Figure 2, if the image captured by the in-vehicle camera detects that the user's face is facing the left front of the vehicle 10, and the map information registers that building A exists in the left front direction of the vehicle 10, the control device 17 identifies building A as the external object of interest to the user.

[0056] Furthermore, at least one of the control device 17 and the server 20 may set the field of view of the vehicle's imaging device 11 so that an object identified as an external object of interest to the user is located at the center of the first image I1. For example, in the driving scene shown in Figure 2, if the external object of interest to the user is identified as building A, the control device 17 adjusts the field of view of the imaging device 11 so that building A is located at the center of the first image I1.

[0057] Furthermore, if multiple objects outside the vehicle are identified as objects of interest to the user, at least one of the control device 17 and the server 20 will set the field of view of the imaging device 11 so that the centers of the multiple objects are located at the center of the first image I1. For example, if three stone monuments lined up in the width direction of the vehicle are identified as objects outside the vehicle of interest to the user, the control device 17 will set the field of view of the imaging device 11 so that the central stone monument in the width direction of the vehicle is located at the center of the first image I1.

[0058] At least one of the control device 17 and the server 20 may set the field of view of the imaging device 11 so that some of the objects are included in the first image I1 if there are multiple objects outside the vehicle of interest to the user and the entirety of all of the objects does not fit in the first image I1. For example, in the driving scene shown in Figure 2, if the objects outside the vehicle of interest to the user are both building A and mountain B, the control device 17 moves (shifts) the field of view of the imaging device 11 to the right in the direction of travel of the vehicle 10 so that the entirety of mountain B is included in the first image I1. Figure 5 shows a superimposed image I3a obtained by superimposing the second image I2 onto the first image I1 which includes the entirety of mountain B. Unlike the superimposed image I3 shown in Figure 4, the superimposed image I3a shown in Figure 5 includes the entirety of mountain B.

[0059] When generating a superimposed image I3, at least one of the control device 17 and the server 20 may calculate the position and direction of the vehicle 10 relative to the specified object, and superimpose a second image I2, which corresponds to the vehicle 10 as viewed from the opposite direction to the direction of the vehicle 10 relative to the specified object, onto the position on the first image I1 corresponding to the position of the vehicle 10 relative to the specified object. For example, in the driving scene shown in Figure 2, if mountain B is the specified object, the server 20 calculates the relative position of the vehicle 10 with respect to the position of mountain B (hereinafter also simply referred to as the relative position), and superimposes the second image I2 onto the position on the first image I1 corresponding to the relative position. The server 20 also calculates the direction in which the vehicle 10 is located relative to mountain B (hereinafter also referred to as the relative direction), and generates a second image I2 corresponding to the vehicle 10 as viewed from the opposite direction to the relative direction (i.e., the direction corresponding to direction D shown in Figure 2).

[0060] At least one of the control device 17 and the server 20 may synthesize images acquired from a plurality of imaging devices 11 mounted on the vehicle 10 to generate an overhead view image of the vehicle 10, associate the height parameter of the vehicle 10 with each pixel of the overhead view image to identify the three-dimensional position corresponding to each pixel, and generate a three-dimensional second image I2 based on the three-dimensional position. For example, the server 20 converts a plurality of images captured by the plurality of imaging devices 11 into an image viewed from above the vehicle 10 using a known image conversion process, and then stitches the converted images together to generate a single overhead view image. When generating the overhead view image, the server 20 refers to a conversion table showing the correspondence between the pixel addresses of each image and the pixel addresses of the overhead view image, and converts the coordinates of each image to the coordinates of the overhead view image. At this time, the server 20 uses a so-called photogrammetry technique to calculate and associate the height position of the vehicle 10 with each pixel of the overhead view image from the plurality of images. Then, the server 20 generates a three-dimensional second image I2 from a planar overhead image using computer graphics, based on the position information of each pixel associated with the height position of the vehicle 10.

[0061] At least one of the control device 17 and the server 20 may, when the recording mode is ON, generate the superimposed image I3 in response to a user instruction (for example, when the user inputs an instruction to generate the superimposed image I3), use one of the first image I1, either the first image I1 captured before the user input the instruction or the first image I1 captured after the user input the instruction, to generate the superimposed image I3. For example, in the driving scene shown in Figure 2, if the external object of interest to the user is identified as building A, and the user inputs an instruction to generate the superimposed image I3, the server 20 will select the first image I1 from the first image I1 captured before the user input the instruction or the first image I1 captured after the user input the instruction, based on the position of building A relative to the vehicle 10, whichever image shows building A more prominently, and generate the superimposed image I3 using the first image I1 in which building A is more prominently visible.

[0062] At least one of the control device 17 and the server 20 may set a virtual viewpoint on the opposite side of the vehicle 10 from the location where the specific object exists, and superimpose a second image I2 corresponding to the vehicle 10 as seen from the virtual viewpoint onto the first image I1. Figure 6 is a plan view showing an example of a virtual viewpoint set for the monument C, which is the specific object. In the example shown in Figure 6, the server 20 sets a virtual viewpoint Q2 at a point-symmetric position P4 with respect to the vehicle 10, on the opposite side of the location P3 where the monument C exists, generates a second image I2 corresponding to the vehicle 10 as seen from the virtual viewpoint Q2, and superimposes it onto the first image I1.

[0063] Furthermore, at least one of the control device 17 and the server 20 may set a virtual viewpoint above the vehicle 10 if the specified object is located below the vehicle 10, and set a virtual viewpoint below the vehicle 10 if the specified object is located above the vehicle 10. Figure 7 is a rear view showing an example of a virtual viewpoint set for the monument C, which is the specified object, similar to Figure 6. In the example shown in Figure 7, since the monument C is located below the center position in the height direction of the vehicle 10, the server 20 sets a virtual viewpoint Q2 at a position P4 that is point-symmetric with respect to the vehicle 10 and above the center position in the height direction of the vehicle 10, generates a second image I2 corresponding to the vehicle 10 as seen from the virtual viewpoint Q2, and superimposes it on the first image I1.

[0064] [Processing in the image generation system] Referring to Figures 8-10, the procedure for information processing by the control device 17 and server 20 will be explained. The processes described below are executed by the processors (CPUs) of the control device 17 and server 20 at predetermined time intervals (for example, every 0.1 to 1 millisecond).

[0065] Figure 8 is a flowchart showing an example of the processing procedure in the image generation system 1, and is a flowchart for when the control device 17 and server 20 generate the superimposed image I3.

[0066] First, in step S1, the control device 17 determines whether the user has activated the recording mode (i.e., whether the recording mode is ON). If it is determined that the recording mode is ON, the process proceeds to step S2. On the other hand, if it is determined that the recording mode is not ON, the control device 17 terminates the process. In step S2, the control device 17 obtains user information from the in-vehicle sensor 15.

[0067] In step S3, the control device 17 determines from the user information whether the user is interested in an object outside the vehicle. If it is determined that the user is interested in an object outside the vehicle, the process proceeds to step S4. On the other hand, if it is determined that the user is not interested in an object outside the vehicle, the control device 17 terminates the process. In step S4, the control device 17 identifies the direction in which the object of interest to the user is located, based on the user information.

[0068] In step S5, the control device 17 acquires the first image I1 from the imaging device 11. In step S6, the control device 17 transmits image information relating to the first image I1 to the server 20 via the communication devices 16, 22 and the network NW. In addition, the control device 17 transmits to the server 20 the current position information and information relating to the direction of the object acquired from the vehicle position detection device 14, along with the image information.

[0069] In step S7, the server 20 sets a virtual viewpoint based on the information received from the control device 17. In step S8, the server 20 generates a second image I2 based on the set virtual viewpoint. In step S9, the server 20 superimposes the second image I2 onto the first image I1 to generate a superimposed image I3. In step S10, the server 20 transmits the route information and the superimposed image I3 to the terminal 30 via the communication device 22 and the network NW, and instructs the terminal 30 to display the route information with the superimposed image I3 superimposed.

[0070] Next, Figure 9 is a flowchart showing another example of the processing procedure in the image generation system 1, and is a flowchart for when the control device 17 generates the superimposed image I3.

[0071] First, in step S31, the control device 17 determines whether the recording mode is ON or OFF. If it is determined that the recording mode is not ON, the control device 17 terminates the process. On the other hand, if it is determined that the recording mode is ON, the process proceeds to step S32. In step S32, the control device 17 obtains user information from the in-vehicle sensor 15.

[0072] In step S33, the control device 17 determines from the user information whether the user is interested in an object outside the vehicle. If it is determined that the user is not interested in an object outside the vehicle, the control device 17 terminates the process. On the other hand, if it is determined that the user is interested in an object outside the vehicle, the process proceeds to step S34. In step S34, the control device 17 obtains map information from the map database 12. In step S35, the control device 17 obtains current location information from the vehicle position detection device 14.

[0073] In step S36, the control device 17 identifies the direction in which an object of interest to the user is located, based on user information, map information, and current location information. In step S37, the control device 17 adjusts the field of view of the imaging device 11 based on the direction identified in step S36. In step S38, the control device 17 acquires the first image I1 from the imaging device 11. In step S39, the control device 17 sets a virtual viewpoint based on the direction identified in step S36. In step S40, the control device 17 generates the second image I2 based on the set virtual viewpoint. In step S41, the control device 17 superimposes the second image I2 onto the first image I1 to generate a superimposed image I3. In step S42, the control device 17 displays the superimposed image I3 on ​​the vehicle 10's display device (not shown).

[0074] Next, Figure 10 is a flowchart showing yet another example of the processing procedure in the image generation system 1, which is a flowchart for when the server 20 generates the superimposed image I3.

[0075] First, in step S51, the server 20 determines whether the recording mode status received from the vehicle 10 is ON. If it is determined that the recording mode status received from the vehicle 10 is not ON, the server 20 terminates processing. On the other hand, if it is determined that the recording mode status received from the vehicle 10 is ON, the process proceeds to step S52. In step S52, the server 20 obtains user information from the vehicle 10.

[0076] In step S53, the server 20 determines from the user information whether the user is interested in an object outside the vehicle. If it is determined that the user is not interested in an object outside the vehicle, the server 20 terminates the process. On the other hand, if it is determined that the user is interested in an object outside the vehicle, the process proceeds to step S54. In step S54, the server 20 obtains map information from the map database 21. In step S55, the server 20 obtains the current location information from the vehicle 10.

[0077] In step S56, the server 20 identifies the direction in which an object of interest to the user is located, based on user information, map information, and current location information. In step S57, the server 20 instructs the vehicle 10's imaging device 11 to adjust the field of view based on the direction identified in step S56. In step S58, the server 20 acquires the first image I1 from the vehicle 10's imaging device 11. In step S59, the server 20 sets a virtual viewpoint based on the direction identified in step S56. In step S60, the server 20 generates the second image I2 based on the set virtual viewpoint. In step S61, the server 20 superimposes the second image I2 onto the first image I1 to generate a superimposed image I3. In step S62, the server 20 transmits the superimposed image I3 to the terminal 30.

[0078] [Embodiments of the present invention] According to this embodiment, an image generation system 1 is provided, comprising a control device 17 and a server 20 for a vehicle 10 used by a user. The system identifies the direction in which an object outside the vehicle of interest to the user exists, acquires a first image I1 taken from the vehicle 10 in the direction in which the object exists, and generates a superimposed image I3 by superimposing a second image I2 of the vehicle 10 onto the first image I1. An image generation method is also provided, which is executed by the image generation system 1. This allows for the generation of an image that includes the vehicle 10 used by the user. Furthermore, the user can acquire a more visually appealing image of the vehicle 10 as a record of their journey.

[0079] In the image generation system 1 and image generation method of this embodiment, the direction in which the object exists is determined based on the user's gesture. This allows the user to specify the field of view of the superimposed image I3 in their preferred direction.

[0080] In the image generation system 1 and image generation method of this embodiment, the direction in which the object exists is determined based on the user's speech content and map information. This allows the user to specify the field of view of the superimposed image I3 in their preferred direction.

[0081] In the image generation system 1 and image generation method of this embodiment, the target object is identified based on information about the user and map information, and the field of view of the imaging device 11 of the vehicle 10 is set so that the target object is located in the center of the first image I1. This allows the user to more accurately image the target object they wish to photograph.

[0082] In the image generation system 1 and image generation method of this embodiment, if there are multiple objects, the field of view is set so that the centers of the multiple objects are located at the center of the first image I1. This allows the user to capture the object they wish to photograph more accurately.

[0083] In the image generation system 1 and image generation method of this embodiment, if there are multiple objects and the entirety of the multiple objects does not fit into the first image I1, the field of view is set so that the entirety of some of the objects is included in the first image I1. This allows the user to capture the object they wish to photograph within the field of view.

[0084] In the image generation system 1 and image generation method of this embodiment, the position and direction of the vehicle 10 relative to the object are calculated, and the second image I2, which corresponds to the vehicle 10 as viewed from the opposite direction to the direction of the vehicle 10 relative to the object, is superimposed on the position on the first image I1 corresponding to the position of the vehicle 10 relative to the object. This allows the second image to be superimposed at a position close to the actual position of the vehicle 10.

[0085] In the image generation system 1 and image generation method of this embodiment, images acquired from a plurality of imaging devices 11 mounted on the vehicle 10 are combined to generate an overhead view image of the vehicle 10, the height parameter of the vehicle 10 is associated with each pixel of the overhead view image to identify the three-dimensional position corresponding to each pixel, and a three-dimensional second image is generated based on the three-dimensional position. This reduces distortion in the second image I2 and makes the joining portions between images smoother.

[0086] In the image generation system 1 and image generation method of this embodiment, when generating the superimposed image I3 in response to the user's instructions, the superimposed image is generated using one of the first images I1, which is either the first image I1 captured before the user inputs the instructions or the first image I1 captured after the user inputs the instructions. This allows for capturing a specific object at a larger size.

[0087] In the image generation system 1 and image generation method of this embodiment, virtual viewpoints Q1 and Q2 are set at positions opposite to the location where the object exists relative to the vehicle 10, and the second image I2 corresponding to the vehicle 10 as seen from the virtual viewpoints Q1 and Q2 is superimposed on the first image I1. This makes it possible to suppress the user from feeling any discomfort regarding the direction of travel or orientation of the vehicle 10 corresponding to the second image.

[0088] In the image generation system 1 and image generation method of this embodiment, if the object is located below the vehicle 10, the virtual viewpoints Q1 and Q2 are set to a position above the vehicle 10, and if the object is located above the vehicle 10, the virtual viewpoints Q1 and Q2 are set to a position below the vehicle 10. This makes it possible to suppress the user from feeling any discomfort with the direction of travel or orientation of the vehicle 10 corresponding to the second image. [Explanation of Symbols]

[0089] 1…Image generation system, 10…Vehicle, 11…Imaging device, 12…Map database, 13…Navigation device, 14…Vehicle position detection device, 15…In-vehicle sensor, 16…Communication device, 17…Control device, 20…Server, 21…Map database, 22…Communication device, 23…Image processing device, 30…Terminal, A…Building, B…Mountain, C…Monument, D…Object direction, NW…Network, P1,P2,P3,P4…Location, Q1,Q2…Virtual viewpoint, L1,L2…Lane, I1…First image, I2…Second image, I3,I3a…Superimposed image

Claims

1. An image generation system comprising a control device and a server for a vehicle used by a user, The control device is The direction in which the object outside the vehicle that the user is interested in is located is identified. A first image is obtained from the vehicle, capturing the direction in which the object is located. The aforementioned server, An image generation system that generates a superimposed image by superimposing a second image of the vehicle onto the first image.

2. The image generation system according to claim 1, wherein the control device identifies the direction in which the object is located based on the user's gesture.

3. The image generation system according to claim 1 or 2, wherein the control device identifies the direction in which the object is located based on the user's speech content and map information.

4. The control device is Based on the user information and map information, the object is identified. The image generation system according to claim 1 or 2, wherein the field of view of the vehicle's imaging device is set so that the object is located at the center of the first image.

5. The image generation system according to claim 4, wherein, if there are multiple objects, the control device sets the field of view so that the centers of the multiple objects are located at the center of the first image.

6. The image generation system according to claim 4, wherein if there are multiple objects and the entirety of the multiple objects does not fit in the first image, the control device sets the field of view so that the entirety of some of the objects is included in the first image.

7. The aforementioned server, The position and direction of the vehicle relative to the object are calculated, The image generation system according to claim 1 or 2, wherein a second image corresponding to the vehicle viewed from a direction opposite to the direction of the vehicle relative to the object is superimposed on a position on the first image corresponding to the position of the vehicle relative to the object.

8. The aforementioned server, Images acquired from multiple imaging devices mounted on the vehicle are combined to generate an overhead view image of the vehicle. By associating a parameter in the height direction of the vehicle with each pixel of the overhead image, the three-dimensional position corresponding to each pixel is identified. The image generation system according to claim 1 or 2, which generates a three-dimensional second image based on the three-dimensional position.

9. The image generation system according to claim 1 or 2, wherein when the server generates the superimposed image in response to the user's instructions, it generates the superimposed image using one of the first images, which is either the first image captured before the user inputs the instructions or the first image captured after the user inputs the instructions.

10. The aforementioned server, A virtual viewpoint is set for the vehicle at a position opposite to the location where the object exists. The image generation system according to claim 1 or 2, wherein the second image corresponding to the vehicle as seen from the virtual viewpoint is superimposed on the first image.

11. The aforementioned server, If the object is located below the vehicle, the virtual viewpoint is set to a position above the vehicle. The image generation system according to claim 10, wherein if the object is located above the vehicle, the virtual viewpoint is set to a position below the vehicle.

12. In an image generation method performed by a control device and server of a vehicle used by a user, The control device is The direction in which the object outside the vehicle that the user is interested in is located is identified. A first image is obtained from the vehicle, capturing the direction in which the object is located. The aforementioned server, An image generation method for generating a superimposed image by superimposing a second image of the vehicle onto the first image.

13. A control device for a vehicle used by a user, The direction in which the object outside the vehicle that the user is interested in is located is identified. A first image is obtained from the vehicle, capturing the direction in which the object is located. A control device that generates a superimposed image by superimposing a second image of the vehicle onto the first image.

14. A server that is connected to the vehicle used by the user in a manner that enables communication, The direction in which the object outside the vehicle that the user is interested in is located is identified. From the vehicle's imaging device, a first image is acquired that captures the direction in which the object exists from the vehicle. A server that generates a superimposed image by superimposing a second image of the vehicle onto the first image.

Citation Information

Patent Citations

  • Multi-media capture system and method

    JP2017194950A