Photograph acquisition and object specification method and apparatus
The method and apparatus enhance vehicle occupant recall of photograph objects and context by capturing images via pointing gestures, processing with AI, and displaying geographic maps with capture direction and location.
Patent Information
- Application Number
- JP2025061132
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-12
- Filing Date
- 2025-04-02
- Publication Date
- 2025-12-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Vehicle occupants often forget the objects of interest in photographs taken during travel and the context of their capture, leading to difficulty in recalling why the photos were taken and their locations.
A method and apparatus that utilize a pointing gesture to capture a photograph, process it with AI to identify objects of interest, and output a graphical user interface with a geographic map, indicating the capture location and direction, enhancing memory recall.
Facilitates recognition of objects of interest in photographs by providing contextual information, aiding memory recall and identification of additional objects in the capture direction.
Smart Images

Figure 2025187002000001_ABST
Abstract
Description
[Background technology]
[0001] Device and vehicle manufacturers are constantly challenged to provide products and services that provide value and benefits to users, such as vehicle occupants, who often take photos from their vehicles, view the photos on a display, and recall from personal memory what was in the photo or why the photo was taken. Summary of the Invention
[0002] An aspect of the present disclosure relates to a method. The method includes capturing a photograph with a camera associated with a vehicle in response to detecting a pointing gesture made by an occupant of the vehicle. The method also includes processing one or more of the photograph to obtain a candidate object of interest within the photograph, a time when the photograph was captured, location information of the vehicle at the time the photograph was captured, a shooting direction extending from the camera at the time the photograph was captured, or map data, and descriptive information associated with the candidate object of interest. The method further includes outputting a graphical user interface by a display. The graphical user interface includes a geographic map, a location icon indicating the location on the geographic map where the photograph was captured, a graphical object representing the shooting direction, and descriptive information associated with the candidate object of interest. The graphical object extends from the location icon in the graphical user interface.
[0003] An aspect of the present disclosure relates to an apparatus. The apparatus includes a processor and a memory having stored thereon instructions that, when executed by the processor, cause the apparatus to perform a step of taking a photograph with a camera associated with a vehicle in response to detecting a pointing gesture made by an occupant of the vehicle. The apparatus also processes one or more of the photograph to obtain a candidate object of interest within the photograph, a time when the photograph was taken, location information of the vehicle at the time the photograph was taken, a shooting direction extending from the camera at the time the photograph was taken, or map data, and descriptive information associated with the candidate object of interest. The apparatus also causes a display to output a graphical user interface. The graphical user interface includes a geographic map, a location icon indicating the location on the geographic map where the photograph was taken, a graphical object representing the shooting direction, and descriptive information associated with the candidate object of interest. The graphical object extends from the location icon in the graphical user interface.
[0004] An aspect of the disclosure relates to a non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor, cause an apparatus to perform a step of taking a photograph with a camera associated with a vehicle in response to detecting a pointing gesture made by an occupant of the vehicle. The apparatus also processes one or more of the photograph to obtain a candidate object of interest within the photograph, the time the photograph was taken, location information of the vehicle at the time the photograph was taken, a shooting direction extending from the camera at the time the photograph was taken, or map data, and descriptive information associated with the candidate object of interest. The apparatus further causes a display to output a graphical user interface. The graphical user interface includes a geographic map, a location icon indicating the location on the geographic map where the photograph was taken, a graphical object representing the shooting direction, and descriptive information associated with the candidate object of interest. The graphical object extends from the location icon in the graphical user interface. [Brief explanation of the drawings]
[0005] [Figure 1]FIG. 1 is a flowchart of a method for acquiring photographs from a vehicle and identifying objects of interest in the photographs according to one or more embodiments. [Figure 2] FIG. 2 is a perspective view of a vehicle according to one or more embodiments. [Figure 3] FIG. 3 is an image of a photograph according to one or more embodiments. [Figure 4] FIG. 4 is a graphical user interface according to one or more embodiments. [Figure 5] FIG. 5 is a graphical user interface according to one or more embodiments. [Figure 6] FIG. 6 is a block diagram of a system for acquiring photographs from a vehicle and identifying objects of interest in the photographs in accordance with one or more embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0006] Aspects of the present disclosure are best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be noted that, in accordance with standard practice in the industry, various features have not been drawn to scale. In fact, the dimensions of various features may be arbitrarily increased or decreased for clarity of discussion.
[0007] The following disclosure provides many different embodiments or examples for implementing various features of the provided subject matter. Specific examples of components, values, operations, materials, arrangements, or equivalents thereof are described below to simplify the disclosure. It should be understood that these are merely examples and are not intended to be limiting. Other components, values, operations, materials, arrangements, or equivalents thereof are contemplated. For example, the formation of a first feature above or on a second feature in the following description includes embodiments in which the first and second features are formed in direct contact with each other, as well as embodiments in which an additional feature is formed between the first and second features such that the first and second features are not in direct contact with each other. Additionally, the present disclosure may repeat reference numerals and / or letters in various examples. This repetition is for purposes of simplicity and clarity and does not, in itself, dictate a relationship between the various embodiments and / or configurations discussed.
[0008] Additionally, spatially relative terms such as "below," "below," "below," "above," "on top," and the like may be used herein to describe the relationship of one element or feature to another element or feature when shown in the figures for ease of description. Spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors may likewise be interpreted accordingly.
[0009] Some vehicles are equipped with infotainment systems that have one or more displays and navigation capabilities and receive user input from vehicle occupants via a touchscreen, microphone, buttons, knobs, joystick, trackpad, motion sensors, cameras, or other suitable controllers. Some infotainment systems are capable of locating addresses based on user input associated with points of interest. Some vehicles are equipped with external cameras that are communicatively coupled to the infotainment systems to take pictures of the scene outside the vehicle in response to user input.
[0010] A vehicle occupant may take many photographs from the vehicle. For example, while traveling, the vehicle occupant may be interested in objects such as buildings, businesses, statues, landscapes, other vehicles, motorcycles, bicycles, people or groups of people, events, or anything else that may be of interest to the vehicle occupant when taking a photograph. However, when the vehicle occupant later reviews the photographs, the vehicle occupant may not remember what objects of interest were present in one or more of the photographs, why one or more of the photographs were taken, or why the vehicle occupant was interested in the objects in the photographs. In addition, the vehicle occupant may not remember information about the location where one or more of the photographs was taken.
[0011] The present specification includes methods and systems for acquiring photographs from a vehicle and identifying objects of interest within the photographs to assist a user in gaining knowledge about the photographs and one or more objects of interest contained within the photographs.
[0012] FIG. 1 is a flowchart of a method 100 for acquiring photographs from a vehicle and identifying objects of interest within the vehicle, according to one or more embodiments.
[0013] In some embodiments, method 100 is implemented using system 600 (FIG. 6). In some embodiments, method 100 is implemented using a system other than system 600 (FIG. 6). In some embodiments, method 100 is implemented in vehicle 200 (FIG. 2). In some embodiments, method 100 is implemented in a vehicle other than vehicle 200 (FIG. 2).
[0014] In some embodiments, method 100 includes detecting a pointing direction and / or a gaze direction of a vehicle occupant, taking a photograph based on a pointing gesture or other suitable user input by the vehicle occupant, and processing information including one or more of location information, line of sight, field of view, direction of travel, vehicle speed, objects in the photograph, historical image data associated with one or more of the location information, line of sight, field of view, direction of travel, or vehicle speed, or other suitable data to identify potential points of interest or objects of interest in the photograph. In some embodiments, the information is input into a multimodal artificial intelligence (AI) to identify potential points of interest and / or objects of interest in the photograph. Information corresponding to the potential points of interest and / or objects of interest in the photograph, such as name, location, and time the photograph was taken, is then obtained, and the information corresponding to the potential points of interest and / or objects of interest is then output on a geographic map viewable via a display of an infotainment system or a mobile device communicatively coupled to the infotainment system. In some embodiments, the mobile device is communicatively coupled to the vehicle via a wired connection or a wireless connection, such as WiFi, Bluetooth, or other suitable wireless connection mode.
[0015] In some embodiments, the geographic map includes at least some of the information overlaid on the geographic map. In some embodiments, the direction of capture of the photo is shown as information overlaid on the geographic map. In some embodiments, the direction of capture of the photo is shown on the geographic map as a cone or triangle or other suitable shape. In some embodiments, the photo is displayed simultaneously with the geographic map. In some embodiments, the geographic map includes two or more icons, the two or more icons indicating that the photo was taken at the location where the icon appears on the geographic map.
[0016] In operation 101, a pointing gesture by a vehicle occupant is detected. In some embodiments, the pointing gesture is replaced or combined with some other suitable user input by the vehicle occupant to interact with the vehicle system and cause the vehicle system to perform a pointing task, such as taking a photo. In some embodiments, an additional user input to the pointing gesture or an alternative user input to the pointing gesture is the gaze direction of the vehicle occupant.
[0017] In some embodiments, the vehicle occupant is the driver of the vehicle. In some embodiments, the vehicle occupant is a passenger of the vehicle. In some embodiments, the driver of the vehicle sits behind the steering wheel of the vehicle. In some embodiments, the vehicle passenger sits in a seat other than the driver's seat, such as a front passenger seat or a rear passenger seat. In some embodiments, the vehicle occupant is any passenger of the vehicle who is sitting or standing in the vehicle.
[0018] In some embodiments, the pointing and / or gaze direction is detected by one or more sensors or cameras in the vehicle. In some embodiments, the vehicle sensor or camera detecting the pointing and / or gaze direction of the vehicle occupant is an interior sensor or camera, such as a sensor or camera facing the interior of the vehicle. In some embodiments, the vehicle sensor or camera detecting the pointing and / or gaze direction of the vehicle occupant is an exterior sensor or camera, such as a sensor or camera facing the exterior of the vehicle. In some embodiments, the vehicle interior sensor or camera and the exterior sensor or camera detecting the pointing and / or gaze direction of the vehicle occupant are physically located inside the vehicle. In some embodiments, the vehicle interior sensor or camera detecting the pointing and / or gaze direction of the vehicle occupant is located on the exterior of the vehicle facing the interior of the vehicle. In some embodiments, the vehicle exterior sensor or camera detecting the pointing and / or gaze direction of the vehicle occupant is located on the interior of the vehicle facing the exterior of the vehicle. In some embodiments, the vehicle's external sensors or cameras detecting the vehicle occupant's pointing motion and / or gaze direction are located on the vehicle's exterior facing the exterior of the vehicle cabin. In some embodiments, the pointing motion and / or gaze direction is detected by a combination of the vehicle's internal sensors or cameras detecting the vehicle occupant's pointing motion and the vehicle's external sensors or cameras detecting the vehicle occupant's pointing motion.
[0019] In operation 103, in response to detecting a pointing gesture or other suitable input made by an occupant of the vehicle, a photograph is taken by one or more cameras to capture a photograph of the view outside the vehicle. In some embodiments, the one or more cameras to capture the photograph of the view outside the vehicle are physically located inside the vehicle. In some embodiments, the one or more cameras to capture the photograph of the view outside the vehicle are mounted on the exterior of the vehicle.
[0020] In operation 105, one or more of the photograph, the time the photograph was taken, the vehicle's position information at the time the photograph was taken, the shooting direction extending from the camera at the time the photograph was taken, or map data are processed to obtain potential objects of interest in the photograph and descriptive information associated with the potential objects of interest.
[0021] In some embodiments, a database associated with the vehicle's infotainment / navigation system is queried based on data obtained from the photograph, one or more of the time the photograph was taken, location information, direction of photograph, or map data to obtain candidate objects of interest within the photograph, and descriptive information associated with the candidate objects of interest. In some embodiments, the database associated with the vehicle's infotainment / navigation system is on-board the vehicle. In some embodiments, the database associated with the vehicle's infotainment / navigation system is remote from the vehicle and accessible via a wired or wireless connection.
[0022] In some embodiments, the capture direction is based on one or more of a detected pointing direction of a pointing gesture of a vehicle occupant, a gaze direction of the vehicle occupant, a speed of the vehicle, or a change in the pointing direction and / or gaze direction of the pointing gesture of the vehicle occupant relative to the side of the vehicle as the vehicle passes the candidate object of interest. The detected pointing direction of the pointing gesture of the vehicle occupant and / or the gaze direction of the vehicle occupant are captured by one or more sensors or cameras in the vehicle that detect the pointing gesture of the vehicle occupant.
[0023] In some embodiments, contextual data including audio data received by a microphone associated with the vehicle within a predetermined time period that includes the time the photograph was taken is processed to obtain candidate objects of interest. In some embodiments, the microphone is an interior microphone configured to capture sounds within the vehicle cabin. In some embodiments, the microphone is an exterior microphone configured to capture sounds outside the vehicle. In some embodiments, the microphone is configured to capture voice commands as user input. In some embodiments, the contextual data includes one or more of a conversation between two or more vehicle occupants, a conversation between a vehicle occupant and a person outside the vehicle, such as via a telephone, video call, or other suitable form of communication, a conversation between a vehicle occupant and a person outside the vehicle, such as through a window or through the vehicle's exterior speakers, a verbal inquiry made by a vehicle occupant, music, other sounds, externally captured sounds, or other suitable interior or exterior sounds that can be captured by one or more microphones associated with the vehicle.
[0024] In some embodiments, processing one or more of the photograph, the time the photograph was taken, the vehicle's location information at the time the photograph was taken, the direction of view extending from the camera at the time the photograph was taken, or the map data and the descriptive information associated with the candidate object of interest to obtain candidate objects of interest includes inputting the photograph, the time the photograph was taken, the vehicle's location information at the time the photograph was taken, the direction of view extending from the camera at the time the photograph was taken, or the map data into a multimodal artificial intelligence (AI) system to generate descriptive information associated with the candidate objects of interest, and processing the descriptive information associated with the candidate objects of interest generated by the multimodal AI. In some embodiments, the multimodal AI is executed locally by one or more processors in the vehicle. In some embodiments, the multimodal AI is executed by one or more processors remote from the vehicle and communicatively coupled to the vehicle via a wired or wireless connection.
[0025] In operation 107, a graphical user interface is output by a display of the vehicle or a portable device communicatively coupled to the vehicle. In some embodiments, the display is associated with the vehicle's infotainment / navigation system. The graphical user interface includes a geographic map, a location icon indicating the location on the geographic map where the photo was taken, a graphical object representing the direction of the photo, and descriptive information associated with the potential object of interest. The graphical object extends from the location icon in the graphical user interface.
[0026] In some embodiments, the graphical object representing the direction of view is a polygon. In some embodiments, the graphical object representing the direction of view is a triangle. In some embodiments, the graphical user interface includes a three-dimensional display and the graphical object representing the direction of view is a cone or other suitable shape. In some embodiments, the graphical object representing the direction of view is a polygon corresponding to the field of view of the camera at the time the picture was taken.
[0027] In some embodiments, the photograph is a composite image of multiple photographs, the quantity of the multiple photographs being based on the speed of the vehicle, and the graphical object representing the direction of photography is shaped based on the field of view of the camera, the speed of the vehicle, and the quantity of the multiple photographs such that the graphical object representing the direction of photography is a composite outer boundary of the field of view of the camera of the composite image of the multiple photographs.
[0028] In some embodiments, the photograph includes two or more candidate objects of interest, and the graphical user interface includes one or more of: a static object icon corresponding to each of the two or more candidate objects; a selectable object icon that, in response to user input, displays descriptive information corresponding to each of the two or more candidate objects; or a highlighted candidate object name corresponding to each of the two or more candidate objects in the photographing direction.
[0029] According to various embodiments, a graphical user interface helps a viewer of a photograph review and / or recognize potential objects of interest in the photograph by providing a graphical object representing the direction of photography and descriptive information related to the potential objects of interest. For example, including a graphical object representing the direction of photography in a graphical user interface with a geographic map may jog the viewer's memory as to why the photograph was taken and / or help them to ascertain whether there are other objects in the photograph and / or near the camera's direction of photography that the vehicle occupant may have intended as objects of interest in the photograph but missed or that were not identified as potential objects of interest.
[0030] Those skilled in the art will recognize that modifications to method 100 are within the scope of the present disclosure. In some embodiments, method 100 includes at least one additional operation. In some embodiments, the order of operations in method 100 is adjusted.
[0031] 2 is a perspective view of a vehicle 200 according to some embodiments. The vehicle 200 is capable of implementing the method 100 (FIG. 1). In some embodiments, the vehicle 200 is capable of implementing the method 100 (FIG. 1) using a system 600 (FIG. 6) mounted on the vehicle. In some embodiments, the vehicle 200 is capable of implementing the method 100 (FIG. 1) based on receiving instructions from a system 600 (FIG. 6) that is remote or separable from the vehicle 200. In some embodiments where the system 600 (FIG. 6) is remote or separable from the vehicle 200, the vehicle 200 is configured to receive instructions to implement the method 100 (FIG. 1) wirelessly or via a wired connection.
[0032] Vehicle 200 includes one or more vehicle systems for implementing the operation of the vehicle. In some embodiments, the one or more vehicle systems include one or more of an infotainment system or a navigation system having one or more displays 201, one or more interior or exterior sensors 203, one or more interior or exterior cameras 205, and at least one camera 207 with a field of view outside the vehicle for taking pictures. In some embodiments, vehicle 200 includes one or more vehicle systems only in the front portion of the cabin. In some embodiments, vehicle 200 includes one or more vehicle systems in both the front portion of the cabin and the rear portion of the cabin.
[0033] FIG. 3 is an image of a photograph 300 according to one or more embodiments.
[0034] Photo 300 is an example of a photograph taken by a camera of vehicle 200 (FIG. 2) according to method 100 (FIG. 1). In this example, photo 300 was taken by an occupant of the vehicle while traveling down a street having several buildings, including a cafe 301, a clothing store 303, and a park 305.
[0035] In response to detecting a pointing gesture made by an occupant of the vehicle, a photograph 300 is taken, which includes a cafe 301 , a clothing store 303 and a park 305 .
[0036] System 600 (FIG. 6) processes photograph 300, the time photograph 300 was taken, vehicle position information at the time photograph 300 was taken, the direction of photography extending from the camera at the time photograph 300 was taken, or map data to obtain candidate objects of interest within photograph 300, and descriptive information associated with the candidate objects of interest.
[0037] For example, when photograph 300 is taken, cafe 301, clothing store 303, and park 305 are all present in photograph 300. Later, when reviewing photograph 300, the vehicle occupant may not remember why photograph 300 was taken or may want information about what objects of interest are in photograph 300. System 600 processes information and data in and associated with photograph 300 to identify potential objects of interest in photograph 300 and provide descriptive information about the potential objects of interest.
[0038] For example, if system 600 determines that cafe 301 is a potential object of interest based on the pointing direction and / or gaze direction of the vehicle occupant's pointing gesture, system 600 obtains descriptive information about cafe 301, referred to in this example as "coffee cafe." System 600 then causes a display, such as display 201 (FIG. 2), to output a graphical user interface including one or more of a geographic map, a location icon indicating the location on the geographic map where the photograph was taken, a graphical object representing the direction of the photograph, or descriptive information associated with the potential object of interest. In some embodiments, the geographic map, the location icon indicating the location on the geographic map where the photograph was taken, the graphical object representing the direction of the photograph, and the descriptive information associated with the potential object of interest are displayed simultaneously. In some embodiments, one or more of the geographic map, the location icon indicating the location on the geographic map where the photograph was taken, the graphical object representing the direction of the photograph, or descriptive information associated with the potential object of interest are displayed via a separate graphical user interface display screen.
[0039] In some embodiments, system 600 obtains candidate objects of interest by inputting photograph 300 and map data into a multimodal AI. In some embodiments, the map data includes information about one or more objects in photograph 300, such as cafe 301, clothing store 303, park 305, etc. The multimodal AI then determines which objects in photograph 300 are candidate objects of interest, for example, based on the positions of the objects in photograph 300. In some embodiments, the multimodal AI is configured to identify which of multiple objects in the photograph are candidate objects of interest in response to determining which object is at the center or closest to the center in the image. For example, in photograph 300, coffee cafe 301 is closest to the center in the photograph. Thus, in this example where the multimodal AI is configured to determine the object at the center of the photograph, the multimodal AI identifies coffee cafe 301 as the candidate object of interest. In some embodiments, in response to determining that the photograph has only one object that is a potential object of interest, the multimodal AI determines that the one object in the image that is a potential object of interest is the candidate object of interest, regardless of the one object's position in the photograph. For example, in FIG. 3, if photo 300 only has clothing store 303 and an open space is shown where cafe 301 and park 305 are located, the multimodal AI will determine that clothing store 303 is a candidate object of interest.
[0040] In some embodiments, contextual data including audio data received by a microphone associated with vehicle 200 is processed to identify potential objects of interest within photograph 300. For example, if, within a predetermined time before, during, or after making a pointing gesture, an occupant of the vehicle says something like, "What's that coffee shop?", "I want to get a coffee at that place," or "Let's get a coffee there and then go to the park next door," system 600 processes the audio information to help identify cafe 301 and / or park 305 as potential objects of interest within photograph 300, and recognizes that clothing store 303 is likely not a potential object of interest.
[0041] FIG. 4 is a graphical user interface 400 according to one or more embodiments.
[0042] Graphical user interface 400 includes a geographic map 401, a location icon 403 indicating the location on the geographic map where a photograph, such as photograph 300 (FIG. 3), was taken, a graphical object 405 representing the direction of the photograph, and descriptive information associated with a candidate object of interest. In this example, the candidate object of interest is a "coffee cafe 407." Also, a clothing store 409 and a park 411 are located in the direction of the photograph. Graphical object 405 extends from location icon 403 in graphical user interface 400. In this example, the descriptive information for "coffee cafe" 407 is underlined. In some embodiments, the descriptive information for the candidate object of interest is highlighted, bolded, rendered in a different color, represented by a thumbnail image of photograph 300, a thumbnail image of a processed and cropped version of photograph 300 focusing on the candidate object of interest, a thumbnail image of an available commercial image of the candidate object of interest, or a thumbnail image based on some other suitable source. In some embodiments, the icon of the candidate object of interest, in this example "Coffee Cafe" 407, is a selectable icon that, when selected via user input, causes photograph 300 to be displayed, descriptive information about the candidate object of interest to be displayed, or some other appropriate action to be triggered. In some embodiments, the descriptive information and photograph 300 are displayed simultaneously with graphical user interface 400. In some embodiments, the descriptive information and photograph 300 are displayed on a separate and distinct graphical user interface from graphical user interface 400.
[0043] In some embodiments, if more than one object of interest candidate is identified in the photograph, the graphical user interface 400 may be highlighted, bolded, rendered in a different color, and represented by a thumbnail image of the photograph 300, by a thumbnail image of a processed and cropped version of the photograph 300 focusing on the object of interest candidate, by a thumbnail image of an available commercial image of the object of interest candidate, or by a thumbnail image based on some other suitable source, and / or optionally include descriptive information for each of the object of interest candidate that is a selectable icon that, when selected via user input, causes the photograph 300 to be displayed, descriptive information about the selected object of interest candidate, or some other appropriate action.
[0044] FIG. 5 is a graphical user interface 500 according to one or more embodiments.
[0045] Graphical user interface 500 includes a cropped image of a candidate object of interest, in this example "Coffee Cafe" 501, and descriptive information 503 about the candidate object of interest. In this example, descriptive information 501 includes the name and address of the candidate object of interest and the date and time the photograph of the candidate object of interest was taken. In some embodiments, the cropped image is generated by cropping out a portion of a source photograph, such as photograph 300 (FIG. 3), that is determined by system 600 to be associated with the candidate object of interest. In some embodiments, graphical user interface 500 includes original photograph 300. In some embodiments, graphical user interface 500 is displayed based on user input received via graphical user interface 400 (FIG. 4). In some embodiments, graphical user interface 500 is displayed simultaneously with graphical user interface 400.
[0046] 6 is a block diagram of a system 600 for acquiring photographs from a vehicle and identifying objects of interest in the photographs in accordance with one or more embodiments. The system 600 includes a hardware processor 602 and a non-transitory computer-readable storage medium 604 that is encoded with, i.e., stores, computer program code 606, i.e., a set of executable instructions. The computer-readable storage medium 604 is also encoded with instructions 607 for interfacing with a manufacturing machine to produce memory arrays. The processor 602 is electrically coupled to the computer-readable storage medium 604 via a bus 608. The processor 602 is also electrically coupled to an input / output (I / O) interface 610 by the bus 608. A network interface 612 is also electrically coupled to the processor 602 via the bus 608. The network interface 612 is connected to a network 614 such that the processor 602 and the computer-readable storage medium 604 can connect to external elements via the network 614. Processor 602 is configured to execute computer program code 606 encoded on computer-readable storage medium 604 to enable system 600 to perform some or all of the operations described in method 100 (FIG. 1) or implemented by vehicle 200 (FIG. 2).
[0047] In some embodiments, processor 602 is a central processing unit (CPU), a multiprocessor, a distributed processing system, an application specific integrated circuit (ASIC), and / or other suitable processing unit.
[0048] In some embodiments, computer-readable storage medium 604 is an electronic, magnetic, optical, electromagnetic, infrared, and / or semiconductor system (or apparatus or device). For example, computer-readable storage medium 604 includes semiconductor or solid-state memory, magnetic tape, removable computer diskette, random access memory (RAM), read-only memory (ROM), rigid magnetic disk, and / or optical disk. In some embodiments using an optical disk, computer-readable storage medium 604 includes a compact disk-read-only memory (CD-ROM), a compact disk-read / write (CD-R / W), and / or a digital video disk (DVD).
[0049] In some embodiments, storage medium 604 stores computer program code 604 configured to cause system 600 to perform some or all of the operations as described in method 100 ( FIG. 1 ) or implemented by vehicle 200 ( FIG. 2 ). In some embodiments, storage medium 604 also stores information used to perform some or all of the operations as described in method 100 ( FIG. 1 ) or implemented by vehicle 200 ( FIG. 2 ), as well as information generated during the performance of some or all of the operations as described in method 100 ( FIG. 1 ) or implemented by vehicle 200 ( FIG. 2 ), such as input data parameters 616, user profile parameters 618, notification data parameters 620, vehicle status parameters 622, and / or a set of executable instructions for performing some or all of the operations as described in method 100 ( FIG. 1 ) or implemented by vehicle 200 ( FIG. 2 ).
[0050] In some embodiments, storage medium 604 stores instructions 607 for interfacing with an external device, such as a mobile device, that enable processor 602 to generate or receive instructions readable by the external device during performance of some or all of the operations described in method 100 (FIG. 1) or implemented by vehicle 200 (FIG. 2).
[0051] System 600 includes an I / O interface 610. I / O interface 610 is coupled to external circuitry. In some embodiments, I / O interface 610 includes a keyboard, keypad, mouse, trackball, trackpad, touchscreen, and / or cursor direction keys for communicating information and commands to processor 602.
[0052] System 600 also includes a network interface 612 coupled to processor 602. Network interface 612 enables system 600 to communicate with a network 614 to which one or more other computer systems are connected. Network interface 612 may include a wireless network interface, such as WiFi, Bluetooth, WiMAX, GPRS, or WCDMA, or a wired network interface, such as a LAN, Ethernet, WAN, USB, IEEE-1394, or other suitable network interface. In some embodiments, some or all of the operations described in method 100 (FIG. 1) or implemented by vehicle 200 (FIG. 2) are implemented in more than one system 600, and information, such as sensor data, window transmittance, forecast information, or vehicle status, is exchanged between the various systems 600 via network 614.
[0053] Supplementary Note 1 An aspect of the present disclosure relates to a method. The method includes capturing a photograph with a camera associated with a vehicle in response to detecting a pointing gesture made by an occupant of the vehicle. The method also includes processing one or more of the photograph to obtain a candidate object of interest within the photograph, a time at which the photograph was captured, location information of the vehicle at the time the photograph was captured, a shooting direction extending from the camera at the time the photograph was captured, or map data, and descriptive information associated with the candidate object of interest. The method further includes outputting a graphical user interface by a display. The graphical user interface includes a geographic map, a location icon indicating a location on the geographic map where the photograph was captured, a graphical object representing the shooting direction, and descriptive information associated with the candidate object of interest. The graphical object extends from the location icon in the graphical user interface.
[0054] Supplementary Note 2 The method described in Supplementary Note 1, wherein the step of processing the photograph, the time the photograph was taken, the vehicle's position information at the time the photograph was taken, the shooting direction extending from the camera at the time the photograph was taken, or map data, and explanatory information associated with the candidate object of interest to obtain candidate objects of interest within the photograph, includes inputting the photograph, the time the photograph was taken, the vehicle's position information at the time the photograph was taken, the shooting direction extending from the camera at the time the photograph was taken, or the map data into a multimodal artificial intelligence (AI) system to generate explanatory information associated with the candidate object of interest, and processing the explanatory information associated with the candidate object of interest generated by the multimodal AI.
[0055] Supplementary Note 3 The method described in Supplementary Note 1 or Supplementary Note 2, wherein the shooting direction is based on one or more of a detected pointing direction of a pointing motion of an occupant of the vehicle, a speed of the vehicle, or a change in the pointing direction of a pointing motion of an occupant of the vehicle relative to the side of the vehicle as the vehicle passes a candidate object of interest.
[0056] Supplementary Note 4 Supplementary Notes 1 to 3: The method of any one of Supplementary Notes 1 to 3, wherein the graphical object representing the shooting direction is a polygon.
[0057] Supplementary Note 5 Supplementary Notes 1 to 3: The method of any one of Supplementary Notes 1 to 3, wherein the graphical object representing the shooting direction is a triangle.
[0058] Supplementary Note 6 Supplementary Notes 1 to 3. The method of any one of Supplementary Notes 1 to 3, wherein the graphical object representing the photographing direction is a polygon corresponding to the field of view of the camera at the time the photograph was taken.
[0059] Supplementary Note 7 4. The method of any one of Supplementary Notes 1 to 3, wherein the photograph is a composite image of a plurality of photographs, the quantity of the plurality of photographs being based on the speed of the vehicle, and the graphical object representing the shooting direction is shaped based on the field of view of the camera, the speed of the vehicle, and the quantity of the plurality of photographs such that the graphical object representing the shooting direction is a composite outer boundary of the field of view of the camera of the composite image of the plurality of photographs.
[0060] Supplementary Note 8 A method according to any one of Supplementary Notes 1 to 7, wherein the names of the candidate objects of interest are highlighted in the graphical user interface.
[0061] Supplementary Note 9 The method of any one of Supplementary Notes 1 to 8, further comprising processing context data including audio data received by a microphone within a predetermined time period including the time the photograph was taken to obtain the candidate object of interest.
[0062] Supplementary Note 10 A method as described in any one of Supplementary Notes 1 to 9, wherein the photograph includes two or more candidate objects of interest, and the graphical user interface includes one or more of a static object icon corresponding to each of the two or more candidate objects of interest, a selectable object icon that displays the explanatory information corresponding to each of the two or more candidate objects in response to user input, or a highlighted candidate object name corresponding to each of the two or more candidate objects in the photographing direction.
[0063] Supplementary Note 11 An aspect of the present disclosure relates to an apparatus. The apparatus includes a processor and a memory having stored thereon instructions that, when executed by the processor, cause the apparatus to perform a step of taking a photograph with a camera associated with a vehicle in response to detecting a pointing gesture made by an occupant of the vehicle. The apparatus also processes one or more of the photograph to obtain a candidate object of interest within the photograph, a time when the photograph was taken, location information of the vehicle at the time the photograph was taken, a shooting direction extending from the camera at the time the photograph was taken, or map data, and descriptive information associated with the candidate object of interest. The apparatus also causes a display to output a graphical user interface. The graphical user interface includes a geographic map, a location icon indicating the location on the geographic map where the photograph was taken, a graphical object representing the shooting direction, and descriptive information associated with the candidate object of interest. The graphical object extends from the location icon in the graphical user interface.
[0064] Supplementary Note 12 The apparatus of Supplementary Note 11, wherein the apparatus is configured to input the photograph, the time the photograph was taken, the vehicle's position information at the time the photograph was taken, the shooting direction extending from the camera at the time the photograph was taken, or map data into a multimodal artificial intelligence (AI) system to process one or more of the photograph, the time the photograph was taken, the vehicle's position information at the time the photograph was taken, the shooting direction extending from the camera at the time the photograph was taken, or map data to obtain candidate objects of interest within the photograph, and explanatory information related to the candidate objects of interest, and process the explanatory information related to the candidate objects of interest generated by the multimodal AI.
[0065] Supplementary Note 13 The apparatus of Supplementary Note 11 or 12, wherein the imaging direction is based on one or more of a detected pointing direction of a pointing motion of an occupant of the vehicle, a speed of the vehicle, or a change in the pointing direction of a pointing motion of an occupant of the vehicle relative to the side of the vehicle as the vehicle passes a candidate object of interest.
[0066] Supplementary Note 14 Supplementary Notes 11 to 13. The apparatus of any one of Supplementary Notes 11 to 13, wherein the graphical object representing the shooting direction is a polygon.
[0067] Supplementary Note 15 Supplementary Notes 11 to 13. The apparatus of any one of Supplementary Notes 11 to 13, wherein the graphical object representing the shooting direction is a triangle.
[0068] Supplementary Note 16 Supplementary Notes 11-13. The apparatus of any one of Supplementary Notes 11-13, wherein the graphical object representing the photographing direction is a polygon corresponding to the field of view of the camera at the time the photograph was taken.
[0069] Supplementary Note 17 14. The apparatus of any one of Supplementary Notes 11 to 13, wherein the photograph is a composite image of a plurality of photographs, the quantity of the plurality of photographs being based on a speed of the vehicle, and the graphical object representing the shooting direction is shaped based on the field of view of the camera, the speed of the vehicle, and the quantity of the plurality of photographs such that the graphical object representing the shooting direction is a composite outer boundary of the field of view of the camera of the composite image of the plurality of photographs.
[0070] Supplementary Note 18 An apparatus described in any one of Supplementary Notes 11 to 17, wherein the names of the candidate objects of interest are highlighted in the graphical user interface.
[0071] Supplementary Note 19 19. The apparatus of any one of Supplementary Notes 11 to 18, further comprising: processing context data including audio data received by a microphone within a predetermined time period including the time the photograph was taken to obtain the candidate object of interest.
[0072] Supplementary Note 20 The apparatus of any one of Supplementary Notes 11 to 19, wherein the photograph includes two or more candidate objects of interest, and the graphical user interface includes one or more of a static object icon corresponding to each of the two or more candidate objects of interest, a selectable object icon that causes the descriptive information corresponding to each of the two or more candidate objects to be displayed in response to user input, or a highlighted candidate object name corresponding to each of the two or more candidate objects in the photographing direction.
[0073] Supplementary Note 21 An aspect of the disclosure relates to a non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor, cause an apparatus to perform a step of taking a photograph with a camera associated with a vehicle in response to detecting a pointing gesture made by an occupant of the vehicle. The apparatus also processes one or more of the photograph to obtain a candidate object of interest within the photograph, the time the photograph was taken, location information of the vehicle at the time the photograph was taken, a shooting direction extending from the camera at the time the photograph was taken, or map data, and descriptive information associated with the candidate object of interest. The apparatus further causes a display to output a graphical user interface. The graphical user interface includes a geographic map, a location icon indicating the location on the geographic map where the photograph was taken, a graphical object representing the shooting direction, and descriptive information associated with the candidate object of interest. The graphical object extends from the location icon in the graphical user interface.
[0074] Supplementary Note 22 22. The non-transitory computer-readable medium of Supplementary Note 21, further comprising: inputting the photograph, the time the photograph was taken, the vehicle's position information at the time the photograph was taken, the shooting direction extending from the camera at the time the photograph was taken, or map data into a multimodal artificial intelligence (AI) system to process one or more of the photograph, the time the photograph was taken, the vehicle's position information at the time the photograph was taken, the shooting direction extending from the camera at the time the photograph was taken, or map data to obtain candidate objects of interest within the photograph, and descriptive information associated with the candidate objects of interest; and processing the descriptive information associated with the candidate objects of interest generated by the multimodal AI.
[0075] Supplementary Note 23 23. The non-transitory computer-readable medium of Supplementary Note 21 or 22, wherein the capture direction is based on one or more of a detected pointing direction of a pointing motion of an occupant of the vehicle, a speed of the vehicle, or a change in the pointing direction of the pointing motion of an occupant of the vehicle relative to the side of the vehicle as the vehicle passes a candidate object of interest.
[0076] Supplementary Note 24 Supplementary Notes 21 to 23. The non-transitory computer-readable medium of any one of Supplementary Notes 21 to 23, wherein the graphical object representing the shooting direction is a polygon.
[0077] Supplementary Note 25 Supplementary Notes 21 to 23. The non-transitory computer-readable medium of any one of Supplementary Notes 21 to 23, wherein the graphical object representing the shooting direction is a triangle.
[0078] Supplementary Note 26 Supplementary Notes 21 to 23. The non-transitory computer-readable medium of any one of Supplementary Notes 21 to 23, wherein the graphical object representing the photographing direction is a polygon corresponding to the field of view of the camera at the time the photograph was taken.
[0079] Supplementary Note 27 24. The non-transitory computer-readable medium of any one of Supplementary Notes 21 to 23, wherein the photograph is a composite image of a plurality of photographs, the quantity of the plurality of photographs being based on the speed of the vehicle, and the graphical object representing the shooting direction is shaped based on the field of view of the camera, the speed of the vehicle, and the quantity of the plurality of photographs such that the graphical object representing the shooting direction is a composite outer boundary of the field of view of the camera of the composite image of the plurality of photographs.
[0080] Supplementary Note 28 28. The non-transitory computer-readable medium of any one of Supplementary Notes 21 to 27, wherein names of the candidate objects of interest are highlighted in the graphical user interface.
[0081] Supplementary Note 29 29. The non-transitory computer-readable medium of any one of Supplementary Notes 21 to 28, further comprising processing context data including audio data received by a microphone within a predetermined time period that includes the time the photograph was taken to obtain the candidate object of interest.
[0082] Supplementary Note 30 A non-transitory computer-readable medium described in any one of Supplementary Notes 21 to 29, wherein the photograph includes two or more candidate objects of interest, and the graphical user interface includes one or more of: a static object icon corresponding to each of the two or more candidate objects of interest; a selectable object icon that causes the descriptive information corresponding to each of the two or more candidate objects to be displayed in response to user input; or a highlighted candidate object name corresponding to each of the two or more candidate objects in the photographing direction.
[0083] The foregoing outlines features of several embodiments so that those skilled in the art may better understand aspects of the present disclosure. Those skilled in the art should appreciate that this disclosure may readily be used as a basis for designing or modifying other processes and structures to carry out the same purposes and / or achieve the same advantages as the embodiments introduced herein. Those skilled in the art should also appreciate that such equivalent structures do not depart from the spirit and scope of the present disclosure, and that various changes, substitutions, and alterations can be made herein without departing from the spirit and scope of the present disclosure.
Claims
1. 1. A processor-implemented method comprising: taking a photograph with a camera associated with the vehicle in response to detecting a pointing gesture made by an occupant of the vehicle; processing the photograph to obtain candidate objects of interest within the photograph, one or more of the time the photograph was taken, the vehicle's position at the time the photograph was taken, the direction of a photograph extending from the camera at the time the photograph was taken, or map data, and descriptive information associated with the candidate objects of interest; outputting a graphical user interface by a display, said graphical user interface comprising: Geographical maps and a location icon indicating the location on the geographic map where the photograph was taken; a graphical object representing the shooting direction and extending from the location icon in the graphical user interface; descriptive information associated with the candidate object of interest; Including steps and A method comprising:
2. processing the photograph to obtain candidate objects of interest within the photograph, one or more of the time the photograph was taken, the position information of the vehicle at the time the photograph was taken, the direction of photography extending from the camera at the time the photograph was taken, or map data, and descriptive information associated with the candidate objects of interest, inputting one or more of the photograph, the time the photograph was taken, the vehicle's location at the time the photograph was taken, the direction of the photograph extending from the camera at the time the photograph was taken, or the map data into a multimodal artificial intelligence (AI) system to generate descriptive information associated with the candidate object of interest; processing the description information generated by the multimodal AI and associated with the candidate object of interest; The method of claim 1 , comprising:
3. 2. The method of claim 1, wherein the capture direction is based on one or more of a detected pointing direction of a pointing gesture of an occupant of the vehicle, a speed of the vehicle, or a change in the pointing direction of a pointing gesture of an occupant of the vehicle relative to a side of the vehicle as the vehicle passes a candidate object of interest.
4. The method of claim 1 , wherein the graphical object representing the view direction is a polygon.
5. The method of claim 1 , wherein the graphical object representing the shooting direction is a triangle.
6. The method of claim 1 , wherein the graphical object representing the photographing direction is a polygon corresponding to the field of view of the camera at the time the photograph was taken.
7. 4. The method of claim 1, wherein the photograph is a composite image of a plurality of photographs, the quantity of the plurality of photographs being based on the speed of the vehicle, and the graphical object representing the shooting direction is shaped based on the field of view of the camera, the speed of the vehicle, and the quantity of the plurality of photographs such that the graphical object representing the shooting direction is a composite outer boundary of the field of view of the camera of the composite image of the plurality of photographs.
8. The method of claim 1 , wherein the names of the candidate objects of interest are highlighted in the graphical user interface.
9. 4. The method of claim 1, further comprising: processing context data including audio data received by a microphone within a predetermined time period including the time the photograph was taken to obtain the candidate object of interest.
10. the photograph includes two or more candidate objects of interest; 4. The method of claim 1, wherein the graphical user interface includes one or more of: a static object icon corresponding to each of the two or more object candidates; a selectable object icon that displays the explanatory information corresponding to each of the two or more object candidates in response to user input; or a highlighted object candidate name corresponding to each of the two or more object candidates in the shooting direction.
11. 1. An apparatus comprising: a processor; A memory having instructions recorded thereon that, when executed by the processor, cause the device to: taking a photograph with a camera associated with the vehicle in response to detecting a pointing gesture made by an occupant of the vehicle; processing the photograph to obtain candidate objects of interest within the photograph, one or more of the time the photograph was taken, the vehicle's position at the time the photograph was taken, the direction of a photograph extending from the camera at the time the photograph was taken, or map data, and descriptive information associated with the candidate objects of interest; outputting a graphical user interface by a display, said graphical user interface comprising: Geographical maps and a location icon indicating the location on the geographic map where the photograph was taken; a graphical object representing the shooting direction and extending from the location icon in the graphical user interface; descriptive information associated with the candidate object of interest; Including steps and Execute the memory and An apparatus comprising:
12. the apparatus to process the photograph to obtain candidate objects of interest within the photograph, one or more of the time the photograph was taken, the position information of the vehicle at the time the photograph was taken, the direction of photography extending from the camera at the time the photograph was taken, or map data, and descriptive information associated with the candidate objects of interest; inputting one or more of the photograph, the time the photograph was taken, the vehicle's location at the time the photograph was taken, the direction of the photograph extending from the camera at the time the photograph was taken, or the map data into a multimodal artificial intelligence (AI) system to generate descriptive information associated with the candidate object of interest; processing the description information generated by the multimodal AI and associated with the candidate object of interest; The apparatus of claim 11 , wherein the apparatus causes the following to be executed:
13. 12. The device of claim 11, wherein the imaging direction is based on one or more of a detected pointing direction of a pointing motion of an occupant of the vehicle, a speed of the vehicle, or a change in the pointing direction of a pointing motion of an occupant of the vehicle relative to a side of the vehicle as the vehicle passes a candidate object of interest.
14. The apparatus of claim 11 , wherein the graphical object representing the imaging direction is a polygon.
15. The apparatus of claim 11 , wherein the graphical object representing the imaging direction is a triangle.
16. 14. The apparatus of claim 11, wherein the graphical object representing the photographing direction is a polygon corresponding to the field of view of the camera at the time the photograph was taken.
17. 14. The apparatus of claim 11, wherein the photograph is a composite image of a plurality of photographs, the quantity of the plurality of photographs being based on the speed of the vehicle, and the graphical object representing the direction of photography is shaped based on the field of view of the camera, the speed of the vehicle, and the quantity of the plurality of photographs such that the graphical object representing the direction of photography is a composite outer boundary of the field of view of the camera of the composite image of the plurality of photographs.
18. 14. The apparatus of claim 11, wherein names of the candidate objects of interest are highlighted in the graphical user interface.
19. 14. The apparatus of claim 11, further comprising: processing context data including audio data received by a microphone within a preset time period including the time the photograph was taken to obtain the candidate object of interest.
20. taking a photograph with a camera associated with the vehicle in response to detecting a pointing gesture made by an occupant of the vehicle; processing the photograph to obtain candidate objects of interest within the photograph, one or more of the time the photograph was taken, the vehicle's position at the time the photograph was taken, the direction of a photograph extending from the camera at the time the photograph was taken, or map data, and descriptive information associated with the candidate objects of interest; outputting a graphical user interface by a display, said graphical user interface comprising: Geographical maps and a location icon indicating the location on the geographic map where the photograph was taken; a graphical object representing the shooting direction and extending from the location icon in the graphical user interface; descriptive information associated with the candidate object of interest; Including steps and A computer program that causes a processor to execute the following.
Citation Information
Patent Citations
Object specification device
JP2007080060A
Object management image generation device and object management image generation program
JP2011243076A
Gesture input apparatus
JP2020052875A
Monitoring device, monitoring system, vehicle and monitoring method
JP2020164003A
Display control device, display control method and display control program
JP2021189823A
Cited By
Surge protective device modules and assemblies
US12506334B2