Information processing device, information processing method, and information processing program
The apparatus integrates image and audio inputs with AI for real-time context recognition, addressing the inefficiencies in existing systems by providing relevant information through audio output, leveraging edge and cloud computing for enhanced user assistance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2026-04-02
AI Technical Summary
Existing systems fail to accurately and efficiently provide relevant information to users based on their real-time visual and audio inputs, particularly in indoor environments, lacking seamless integration and real-time processing capabilities.
An information processing apparatus and method that integrates image and audio input units with a control unit to acquire and analyze user's line of sight images and audio, utilizing AI models for context recognition and providing relevant information through audio output, while leveraging edge and cloud computing for efficient data processing.
Enables accurate and timely provision of information to users by matching visual inputs with associated data, enhancing context recognition and reducing latency through localized processing and cloud support, thus improving user assistance and decision-making.
Smart Images

Figure 0007839520000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] Patent Document 1 discloses a control system including an input unit that receives user input, a processing unit that performs specific processing using a text generation model that generates text according to the input text, and an output unit that controls an electronic device to output the result of the specific processing. When the processing unit receives the position information of a user in an indoor facility, the processing unit performs specific processing to provide the user with specific information related to the indoor facility based on the output of the text generation model when the information including the position information is used as the input text.
Prior Art Document
Patent Document
[0003]
Patent Document 1
Summary of the Invention
[0004] One aspect of an embodiment of the present invention is an apparatus including a control unit, a storage unit, a communication unit, and an image input unit. The control unit includes a first image acquisition unit that acquires a first image, which is an image in the visual field direction of the user's line of sight, from the image input unit and stores it in the storage unit. The control unit includes a position information acquisition unit that acquires the position information of the user and stores it in the storage unit. The control unit includes a second image acquisition unit that transmits the position information to a related device via the communication unit and receives, via the communication unit from the related device, a second image, which is an image associated with the position information, in the related device. The control unit includes a determination unit that compares the first image and the second image and determines whether they match. When it is determined that they match, the control unit includes a related information acquisition unit that receives, via the communication unit from the related device, related information, which is information associated with the second image, and stores it in the storage unit.
[0005] In one aspect of the above, the apparatus further comprises an audio output unit. The control unit includes a generation processing unit that transmits the related information to a generation processing unit including an artificial intelligence model via the communication unit, and receives from the generation processing unit via the communication unit a related sentence generated by the artificial intelligence model based on the related information in the generation processing unit. The control unit also includes an audio output processing unit that outputs the related sentence as audio from the audio output unit.
[0006] In one aspect described above, the apparatus further comprises an audio input unit. The control unit includes an audio acquisition unit that acquires audio from the audio input unit. The control unit includes an audio analysis unit that analyzes the audio. The second image acquisition unit receives the second image from the associated device via the communication unit according to the user's intent based on the analyzed audio.
[0007] In the above aspect, the location information acquisition unit corrects the location information according to the determination result of the determination unit.
[0008] In one aspect of the above, the control unit includes a route information processing unit that receives route information to a destination corresponding to the location information from the related device via the communication unit, and transmits a request regarding the route information to the related device via the communication unit. If the determination unit determines that there is no match, the route information processing unit transmits a request to change the route information to the related device via the communication unit.
[0009] Another aspect of one embodiment of the present invention is a method performed by a computer comprising a control unit, a storage unit, a communication unit, and an image input unit. The method includes the steps of the control unit acquiring a first image from the image input unit which is an image of the user's line of sight and storing it in the storage unit. The control unit acquiring the user's location information and storing it in the storage unit. The control unit transmitting the location information to an associated device via the communication unit and receiving a second image from the associated device via the communication unit which is an image associated with the location information in the associated device. The control unit comparing the first image and the second image and determining whether they match. If the control unit determines that they match, it includes the steps of receiving associated information from the associated device via the communication unit which is information associated with the second image and storing it in the storage unit.
[0010] Another aspect of one embodiment of this invention is a program that causes a computer to perform the above method. [Brief explanation of the drawing]
[0011] [Figure 1] Figure 1 is a schematic diagram of system S in one embodiment of the present invention. [Figure 2] Figure 2 is a schematic diagram of the functional configuration of device 1 in one embodiment of the present invention. [Figure 3] Figure 3 is a schematic diagram of an example of the use of device 1 in one embodiment of the present invention. [Figure 4] Figure 4 is a schematic diagram of the functional configuration of the related device 2 in one embodiment of the present invention. [Figure 5] Figure 5 is a schematic diagram of an example of the use of the related device 2 in one embodiment of the present invention. [Figure 6] Figure 6 is a schematic diagram of the functional configuration of the production processing apparatus 3 in one embodiment of the present invention. [Figure 7] Figure 7 is a schematic diagram of the operation flow in one embodiment of the present invention. [Figure 8] Figure 8 is a schematic diagram of the operation flow in one embodiment of the present invention. [Figure 9] Figure 9 is a schematic diagram of the operation flow in one embodiment of the present invention. [Figure 10] Figure 10 is a schematic diagram of the operation flow in one embodiment of the present invention. [Modes for carrying out the invention]
[0012] The present invention will be described in detail below with reference to the drawings illustrating its embodiments. Note that the following embodiments are not intended to limit the invention as described in the claims. Furthermore, not all combinations of features described in the embodiments are necessarily essential as means of solving the invention.
[0013] Figure 1 is a schematic diagram of a system S in one embodiment of the present invention. System S consists of a device 1, an associated device 2, a generation processing device 3, and a network N. Device 1, the associated device 2, and the generation processing device 3 are connected to each other via the network N so that they can communicate with one another.
[0014] Device 1 is worn by the user and performs input, processing, and output. In Figure 1, Device 1 is shown as an earphone-type device worn on the user's left ear. Such a device may have a separate camera and earphone, or may be implemented as a glasses-type wearable, but a preferred embodiment is one in which the camera is built into the earphone shape and simultaneously acquires images (video) and sound in high quality. In particular, an earphone-type device equipped with a camera structure that can capture images from a viewpoint close to the user's line of sight while utilizing the limited space around the ear is desirable. In addition, although Device 1 is worn on the user's left ear in Figure 1, it may be worn on the user's right ear, or as two Device 1s, each containing a camera module optimized for the left and right ears, and worn on both ears. Furthermore, Device 1 may have features such as a wide field of view for the lens, miniaturization of the camera module, ventilation and heat dissipation design, and non-ear-blocking methods such as bone conduction or open-ear type. Also, unlike glasses-type wearables (e.g., smart glasses), Device 1 does not need to have an image display unit. In this case, the processing and memory capacity required for image display become unnecessary, resulting in power savings and reduced latency.
[0015] Figure 2 is a schematic diagram of the functional configuration of device 1 in one embodiment of the present invention. Device 1 comprises a control unit 100, a storage unit 150, a communication unit 160, an image input unit 170, an audio input unit 180, and an audio output unit 190. In this embodiment, by including the image input unit 170, the audio input unit 180, and the audio output unit 190, device 1 achieves effects such as energy saving through proximity sensing, highly efficient information processing in a small device, and close integration with physical context.
[0016] The control unit 100 is implemented by a control circuit including a central processing unit (CPU). Specifically, it can be implemented by electronic circuits such as a microprocessor, microcontroller, digital signal processor (DSP), or application-specific integrated circuit (ASIC). The CPU may be a multi-core processor with a clock frequency of 1 GHz or higher and multiple cores. The control unit 100 may also include a graphics accelerator mechanism including one or more such components, such as a 2D accelerator, a 3D accelerator, or a video accelerator.
[0017] The memory unit 150 includes a main memory unit and an auxiliary memory unit. The main memory unit is a volatile or non-volatile semiconductor memory medium that the CPU directly reads and writes to. Specifically, it is implemented by dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), resistive random access memory (ReRAM), etc. The storage capacity of the main memory unit may be, for example, in the range of 1GB to 128GB. The auxiliary memory unit is a storage medium other than the main memory unit and is suitable for long-term storage of large amounts of data. Specifically, it is implemented by hard disk drives (HDDs), solid-state drives (SSDs), optical discs, magnetic tapes, non-volatile memory, etc. The storage capacity of the auxiliary memory unit may be, for example, in the range of 1TB to 10TB. The memory unit 150 may also store data using blockchain technology.
[0018] The memory unit 150 stores data such as images (videos) and voices input to the device 1 in a secure area (encrypted storage) dedicated to the user himself / herself as a principle. However, when the data is transmitted and stored in another device (such as cloud storage) from the perspective of optimal utilization of the capacity, the control unit 100 of the device 1 performs processes such as anonymization and tokenization during communication with the other device, and realizes the use of the data (for example, training of an artificial intelligence model, etc.) in a form where an individual cannot be identified. Thereby, sensitive data and analysis results of the user individual are managed by the device 1 (or a dedicated gateway), and strict handling of user information becomes possible.
[0019] Note that the control unit 100 may have a function of performing biometric authentication (such as voiceprint authentication and face authentication) at the timing of attachment / detachment to the user. Further, the control unit 100 may adopt an encryption chip at the hardware level so that data cannot be restored even if a malicious third party physically obtains the device. Thereby, unauthorized use of the device 1 can be prevented.
[0020] The communication unit 160 transmits and receives information using wireless communication technology. Specifically, communication protocols such as cellular communication (5G, 4G LTE), WiFi (IEEE 802.11a / b / g / n / ac), Bluetooth (registered trademark), near-field wireless communication (NFC), and infrared communication can be used. Further, the communication unit 160 may transmit and receive information using wired communication technology. As the wired communication technology, Ethernet, optical fiber communication, serial communication (RS-232, RS-485), USB, or technologies similar thereto can be applied. The communication unit 160 includes an interface for transmitting and receiving data, and can transmit a control signal, a data packet, or an audio / video signal based on wireless communication technology or wired communication technology. Further, the communication unit 160 may have a protocol conversion function and include a configuration for ensuring compatibility between different communication methods.
[0021] The image input unit 170 may be a camera module provided in the device 1 worn on the user's ear. For example, the image input unit 170 is realized by a camera module having an image sensor that optically acquires images. Specifically, it may include an image sensor with a charge-coupled device (CCD) or complementary metal-oxide-semiconductor (CMOS) image sensor, with a pixel count of 5 to 100 megapixels and a sensor size of 1 / 2.3 inch to 1 inch. In this embodiment, the image input unit 170 may be a digital camera (including a digital video camera) incorporated into the device 1 worn on the user's ear. The image input unit 170 can be made easier for the user to wear by being provided as a small camera on the outside of the earphone-type housing, or by being built integrally into the earphone-type housing.
[0022] The image input unit 170 is located at the user's eye level and can acquire images (including video, which is a series of images) in the direction of the user's line of sight. This has the advantage of acquiring images that the user perceives visually with greater accuracy, thereby increasing their usefulness as personal data. In this specification, "image" includes not only still images but also a series of images (video, real-time video, etc.).
[0023] The image input unit 170 may include a camera module that captures the area in front of the user, as well as modules that capture the area behind the user, modules that capture a different direction from the front and rear, and modules that are additionally installed depending on the purpose of improving accuracy in each direction. This allows for the inclusion of multi-angle images as the first image, enabling the acquisition of image information for rearward safety checks and images of the user's blind spots.
[0024] The audio input unit 180 may be implemented by a microphone having a transducer that converts sound waves into electrical signals. Specifically, various microphone types such as electrostatic, piezoelectric, and electromagnetic types can be used. A microphone with a frequency response of 20Hz to 20kHz and a sensitivity of -40dB to 0dB can be used. In this embodiment, the audio input unit 180 may be a microphone incorporated into the device 1 worn on the user's ear. The audio input unit 180 may consist of multiple beamforming microphones positioned to avoid mutual interference with the image input unit 170 (camera, etc.) as microphones for acquiring sound.
[0025] The voice input unit 180 may continuously input voice spoken by the user. Alternatively, the voice input unit 180 may input voice according to specific conditions. For example, the voice input unit 180 may input voice in response to user instructions, or it may automatically input voice when voice is produced, and not input voice if no voice is produced or if the voice is weak.
[0026] One of the features of this embodiment is that the device 1 is equipped with an image input unit 170 and an audio input unit 180, and the control unit 100 integrates and analyzes the images (video) and audio acquired from the image input unit 170 and the audio input unit 180, thereby improving context recognition accuracy and enabling more accurate advice and assistance. Specifically, real-time acquisition is possible, where the image input unit 170 acquires image data and the audio input unit 180 acquires it simultaneously. Then, by analyzing the acquired images and audio together, multimodal analysis becomes possible, which simultaneously recognizes the user's speech content and gaze direction, such as the situation the user is in and what they are talking about. Furthermore, by combining surrounding object recognition, speech recognition, speech synthesis, etc., it is possible to provide additional information about products that come into the user's field of vision via audio, or to suggest the most suitable reply candidate to the user based on the surrounding conversation content, thereby enabling understanding of the user's situation and dialogue, and assisting the user's actions and decision-making.
[0027] The audio output unit 190 is implemented by an earphone or speaker, which is an audio playback device that converts electrical signals into sound waves. Specifically, various driver types such as dynamic, balanced armature, and hybrid can be employed, and as an example, an audio output device with a playback frequency response of 20Hz to 20kHz and an impedance in the range of 16 ohms to 120 ohms can be used. In this embodiment, the audio output unit 180 may be an earphone incorporated into the device 1 that is worn in the user's ear.
[0028] Although not shown in Figure 1, in this embodiment, device 1 may be equipped with other sensors such as an acceleration sensor, a gyroscope sensor, and a geomagnetic sensor. This allows for the detection of user movements and changes in posture via device 1. Furthermore, device 1 may be equipped with sensors such as a heart rate sensor, pulse wave sensor, blood pressure sensor, blood oxygen concentration sensor, skin potential sensor (EDA sensor), and electromyography sensor (EMG sensor). This makes it possible to acquire the user's biological information. In addition, device 1 may be equipped with a temperature sensor, humidity sensor, atmospheric pressure sensor, light sensor, ultraviolet sensor, infrared sensor, CO2 sensor, PM2.5 sensor, VOC (volatile organic compound) sensor, odor sensor (gas sensor), pH sensor, electrochemical sensor, and ion sensor. This enables monitoring of the surrounding weather and air environment, and allows for the detection of chemical substances and analysis of gaseous components.
[0029] The control unit 100 includes a first image acquisition unit 102, a position information acquisition unit 104, a second image acquisition unit 106, a determination unit 108, a related information acquisition unit 110, a generation processing unit 112, an audio output processing unit 114, an audio acquisition unit 116, an audio analysis unit 118, and a route information processing unit 120.
[0030] The first image acquisition unit 102 acquires a first image from the image input unit 170, which is an image in the direction of the user's line of sight, and stores it in the storage unit 150.
[0031] Here, the first image acquisition unit 102 will be described with reference to Figure 3. Figure 3 is a schematic diagram of an example of the use of the apparatus 1 in one embodiment of the present invention.
[0032] In the front view of Figure 3, device 1 is worn on the user's left ear. The side view is a view from the user's right side, and although device 1 is on the opposite side, it is schematically superimposed on the user. In this side view, a dashed line 11 is shown indicating the direction of the user's line of sight from device 1. The direction of the line of sight may also be expressed as the line of sight direction. The first image acquisition unit 102 acquires a first image, which is an image of the direction of the user's line of sight, input from the image input unit 170. The first image acquisition unit 102 then stores the acquired first image in the storage unit 150.
[0033] As described above, the first image is an image (including video) input from the image input unit 170 in the direction of the user's line of sight, and is, for example, an image of a building, object, landscape, etc. that is within the user's field of view. The input of the first image may be performed based on instructions from the user via the voice input unit 180, or it may be performed continuously or at predetermined timings according to predetermined conditions.
[0034] The first image acquisition unit 102 may identify an object (e.g., a building or sign) in the user's line of sight in the first image. For example, the first image acquisition unit 102 may use computer vision technology, object detection algorithms, and contextual understanding models to identify the object using user gaze tracking (eye tracking via the image input unit 170 or gaze detection algorithm), gaze point analysis, and image recognition technology. Specifically, it may utilize gaze detection technology, image analysis technology, convolutional neural networks (CNNs), and transfer learning to identify the user's line of sight and the object of attention in the image with high accuracy. Furthermore, the first image acquisition unit 102 may use sensors to sense head movements and fine-tune image input, or use augmented reality (AR) glasses, which overlay virtual information onto the real field of view connected as another device, to acquire the first image by supplementing the location the user actually wants to see, which cannot be captured by head orientation alone.
[0035] The location information acquisition unit 104 acquires the user's location information and stores it in the storage unit 150. The location information may be acquired by GPS (Global Positioning System). In addition to GPS, the location information may be acquired by one or more combinations of technologies such as technologies that use Wi-Fi radio waves, technologies that use RFID (Radio Frequency Identifier) tags, IMES (Indoor Measurement and Estimation System) positioning technology, IPS (Indoor Positioning System) systems, GNSS (Global Navigation Satellite System), and RTK (Real Time Kinematic). The location information acquisition unit 104 stores the acquired location information in the storage unit 150. Note that the location information may include not only point information such as latitude, longitude, coordinates, and PoI (Point of Interest) data of the location, but also surrounding information such as direction, orientation, slope, terrain, season, weather, and day / night information at that location.
[0036] Furthermore, the location information acquisition unit 104 corrects the location information according to the determination result of the determination unit 108, which will be described later.
[0037] The second image acquisition unit 106 transmits location information to the associated device 2 via the communication unit 160, and the associated device 2 receives the second image, which is an image associated with the location information, from the associated device 2 via the communication unit 160.
[0038] Here, the second image acquisition unit 102 will be described with reference to Figures 4 and 5. Figure 4 is a schematic diagram of the functional configuration of the related device 2 in one embodiment of the present invention. Figure 5 is a schematic diagram of an example of the use of the related device 2 in one embodiment of the present invention.
[0039] In Figure 4, the related device 2 comprises a control unit 200, a storage unit 250, and a communication unit 260. The related device 2 may be, for example, a cloud server, a database server, etc. In this embodiment, device 1 functions as an edge terminal and communicates with the related device 2, which is a cloud server, to optimize data processing and the entire system. Device 1 is equipped with a camera (image input unit 170), a microphone (voice input unit 180), a sensor, or other data acquisition means, and transmits the acquired data to the related device 2 in real time or at a predetermined timing. As for the communication method, wired communication (Ethernet, optical fiber communication, serial communication, etc.) or wireless communication (WiFi, cellular communication, LPWA, etc.) can be applied.
[0040] Related device 2 may store and analyze data received from device 1 and return processing results and control signals to device 1 as needed. For example, a navigation system can be built on related device 2, route search and traffic analysis can be performed on the cloud side, and optimal route information based on the results can be sent to device 1. In this case, device 1 sends location information such as GPS information and progress data to related device 2, and related device 2 calculates the optimal route considering the latest map information, traffic conditions, weather information, past driving data, etc., and returns the result to device 1. This reduces the processing load on device 1 while realizing appropriate route guidance based on the latest traffic conditions. Device 1 and related device 2 may operate while communicating with each other, realizing efficient data processing through the cooperation of edge computing and cloud computing. Device 1 can improve the overall processing efficiency of the system by locally performing tasks that require low-latency processing, such as rendering the navigation screen (for example, rendering the navigation screen on an image output device such as a display connected to device 1) and voice guidance via the voice output unit 190, and offloading computationally intensive route search and traffic analysis to related device 2.
[0041] The control unit 200 may be implemented by a control circuit including a CPU, similar to the control unit 100.
[0042] The control unit 200 includes a navigation processing unit 202. The navigation processing unit 202 provides a navigation system or map service that guides the user to a destination by registering destination and route information to the map data. Although not shown, data and information required for the navigation system, etc., may be stored in the storage unit 250, or they may be stored in other storage devices and transmitted and received via the communication unit 260.
[0043] The storage unit 250 may include a main storage unit and an auxiliary storage unit, similar to the storage unit 150. The storage unit 250 includes a database 252 and location information 254.
[0044] Database 252 includes a second image associated with location information 254. For example, database 252 may associate location information with an image taken at the location indicated by the location information and store the image as the second image. In addition to location information and the second image, database 252 may also store one or more pieces of information related to the location and the second image as related information. For example, related information may include the latitude and longitude of the location, coordinates, altitude, geology, disaster-related information, etc. Related information may also include the name, address, telephone number, etc. of the object shown in the second image.
[0045] The second image may be a combination of multiple panoramic images, such as Google Street View (registered trademark), which virtually represents the real world on a map. The second image may also be any image related to a specific location, including, for example, images of the interior or basement of a building at a specific location, images of natural landscapes or urban scenery, images of buildings, etc., and images taken by ordinary people and associated with a specific location.
[0046] Location information 254 is the user's location information transmitted from device 1 and stored in associated device 2. The second image and related information are identified by the location information 254 and database 252.
[0047] The communication unit 260 may transmit and receive information using one or more wireless communication technologies and wired communication technologies, similar to the communication unit 160.
[0048] In this embodiment, the related device 2 is a different device from device 1, but in other embodiments, the functions of related device 2 may be implemented in device 1. For example, the functions of the control unit 200 of related device 2 may be implemented in the control unit 100 of device 1, and some or all of the data stored in the storage unit 250 may be stored in the storage unit 150 of device 1, thereby enabling the functions of related device 2 described in this specification to be implemented in device 1. Alternatively, device 1 and related device 2 may cooperate, and some of the functions of related device 2 may be implemented in device 1. For example, only the necessary data from the data stored in the storage unit 250 of related device 2 may be temporarily stored in the storage unit 150 of device 1. This allows for faster processing without going through the communication unit, and enables operation via offline or low-bandwidth communication. In these embodiments, when some or all of the functions of related device 2 are implemented in device 1, the related device 2 in this specification may be read as device 1, the storage unit 250 as storage unit 150, and so on, and the configuration described as related device 2 may be read as the corresponding configuration of device 1.
[0049] Figure 5 schematically shows an example of use in which device 1 uses related device 2 to perform navigation 501 and acquire a second image 502 and related information 503.
[0050] Navigation 501 may be implemented in the navigation processing unit 202 of the related device 2. In Figure 5, navigation 501 displays route information from the current location to the destination. The current location corresponds to the location identified by the user's location information 254. Although the route information from the current location to the destination is shown visually in Figure 5 for convenience, the route information may be shown by voice in this embodiment. For example, the route information may be shown by voice from the current location, indicating the direction, distance, landmarks, etc., and the subsequent route to the destination may be shown in the same way by voice.
[0051] The second image 502 is stored in the database 252 of the related device 2 as an image associated with the current location, i.e., the location identified by the user's location information 254. In Figure 5, the second image 502 shows an image of a building located at the current location. The second image 502 may be an image that was previously taken of that location. Alternatively, the second image 502 may be an image that was generated as an image associated with that location. For example, the second image 502 may be a projected image of the completed building generated from the building plan at that location.
[0052] Related information 503 refers to information held in the database 252 of the related device 2 as being associated with the second image 502. In Figure 5, related information 503 shows the name, address, and telephone number of the building image shown in the second image 502. Related information 503 may be information that has been stored in advance as being associated with the second image 502. In addition, related information 503 may include detailed metadata such as the date and time the image was taken, day and night, season, weather, update date and time, frequency, position and orientation of the target object, and position information correction information.
[0053] In this embodiment, the related information 503 is shown as information associated with the second image 502 in the database 252 of the related device 2, but in other embodiments, the related information 503 may be identified without such association. For example, the related information 503 may be information associated with the location information 254 in the database 252. Alternatively, the related information 503 may be information identified by the control unit 200 of the related device 2 through web scraping of one or more of the location information 254 and the second image 502.
[0054] The second image acquisition unit 106 transmits location information to the associated device 2 via the communication unit 160, so that the associated device 2 stores it as location information 254 in the storage unit 250. The second image acquisition unit 106 then receives the second image 502, which is the image associated with the location information 254, from the associated device 2 via the communication unit 160.
[0055] Furthermore, in this embodiment, the second image acquisition unit 106 may acquire a second image of an object projected onto the image acquired by the first image acquisition unit 102, according to certain conditions. For example, if the first image acquisition unit 102 continuously acquires an image on which a specific object is projected for a certain period of time, the second image acquisition unit 106 may acquire a second image, assuming that the user wants to obtain related information about that object. Here, the conditions such as the aforementioned certain period of time may be set in advance by the user via the voice input unit 180 or other input device connected to the device 1 (for example, a keyboard connected by communication technology such as wireless communication technology), or they may be set as initial settings.
[0056] The determination unit 108 compares the first image and the second image and determines whether they match. As described above, the first image is an image input from the image input unit 170 in the direction of the user's line of sight, and the second image is an image associated with the current location, i.e., the location identified by the user's location information, in the database 252 of the related device 2. The determination unit 108 uses the first image as the target image and the second image as the reference image, and determines the similarity using matching techniques such as template matching, shape matching, and feature matching for the whole or a part of the image, and may determine that they match if the similarity exceeds a certain threshold. For example, the determination unit 108 uses an image processing algorithm to compare the features (including local features), pixel information, feature points, structural similarity, etc., of the two images in detail. For feature extraction, algorithms such as SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), and ORB (Oriented FAST and Rotated BRIEF) may be used. Comparison methods may utilize techniques such as image matching, feature extraction, and image similarity evaluation using convolutional neural networks (CNNs).
[0057] The comparison between the first and second images by the determination unit 108 may be performed on only a portion of each image. For example, among the multiple objects or spaces captured in the first image, the comparison with the second image may be performed only on the objects or spaces that are particularly in focus based on the user's eye movements. This allows for more accurate acquisition of information about the objects or spaces that the user is paying attention to.
[0058] Furthermore, the determination unit 108 may perform a comparison with the first image using only the second image whose creation and update timings meet predetermined conditions. The predetermined conditions may be, for example, that the second image was created or updated within the last six months, or the conditions may be set for the most recent period. The setting of the conditions and the determination of whether the second image meets the predetermined conditions may be performed by the second image acquisition unit of device 1, or by the control unit 200 of related device 2. This allows the user to obtain relevant information according to the current situation.
[0059] Furthermore, the determination unit 108 may determine whether each piece of information included in the location information and each piece of information included in the related information satisfy predetermined conditions. For example, the determination unit 108 may determine whether the information on the season, weather, and day / night cycle at the location, included in the location information acquired by the location information acquisition unit 104, and the information on the season, weather, and day / night cycle included in the related information associated with the location information satisfy predetermined conditions (for example, two or more pieces of information match). Based on the determination result, the determination unit 108 may perform one or more comparisons and matching checks between the first image and the second image, and may store the determination result in the storage unit 150 and output it as sound to the audio output unit 190.
[0060] In this embodiment, the determination unit 108 is implemented in the control unit 100 of device 1, but in other embodiments, the determination unit 108 may be implemented in related device 2 or other devices. In this case, the processing of the determination unit 108 is performed in a device different from device 1, and device 1 may receive the result via the communication unit 160.
[0061] If the determination unit 108 determines that the first image and the second image match, the related information acquisition unit 110 receives related information 503, which is information associated with the second image, from the related device 2 via the communication unit 160 and stores it in the storage unit 150. In Figures 4 and 5, the user's location information is stored as location information 254 in the storage unit 250 of the related device 2 as the current location of the navigation 501. Here, if the determination unit 108 determines that the first image, which was input by the image input unit 170 at the user's current location indicated by the location information 254 and acquired by the first image acquisition unit 102, matches the second image 502 associated with the location information 254 in the database 252, then related information 503, which is information associated with the second image 502, is identified. The related information acquisition unit 110 receives related information 503 from the related device 2 via the communication unit 160 and stores it in the storage unit 150.
[0062] As described above, in Figure 5, the related information 503 shows the name, address, and telephone number of the building image shown in the second image 502. The related information acquisition unit 110 may receive one or more of the building image name, address, and electric field number from the related device 2 and store them in the storage unit 150. This allows the device 1 to utilize the related information of the building captured in the user's line of sight.
[0063] Furthermore, the related information 503 may be information stored in advance as information related to the second image 502, and may include detailed metadata such as the date and time the image was taken, day or night, season, weather, update date and time, frequency, position and orientation of the target object, and position information correction information. The related information acquisition unit 110 may receive one or more of such detailed metadata and correction information from the related device 2 and store them in the storage unit 150. This enables seamless multimodal position estimation that accurately estimates the current location and orientation of an image (including video) being captured by the user in real time.
[0064] In this embodiment, the related information 503 is shown as information associated with the second image 502 in the database 252 of the related device 2, but in other embodiments, the related information 503 may be identified without such association. For example, the related information 503 may be information associated with the location information 254 in the database 252. Alternatively, the related information 503 may be information identified by the control unit 200 of the related device 2 through web scraping of one or more of the location information 254 and the second image 502. This allows the user to acquire and utilize a wide range of related information regarding objects, spaces, etc., perceived in their field of view.
[0065] The generation processing unit 112 transmits related information 503 to the generation processing unit 3, which includes the artificial intelligence model 352, via the communication unit 160, and receives the related sentences generated by the artificial intelligence model 352 based on the related information 503 from the generation processing unit 3 via the communication unit 160.
[0066] Here, the generation processing unit 112 will be described with reference to Figure 6. Figure 6 is a schematic diagram of the functional configuration of the generation processing device 3 in one embodiment of the present invention.
[0067] The generation processing device 3 comprises a control unit 300, a storage unit 350, and a communication unit 360. The generation processing device 3 may be, for example, a cloud server, a database server, etc. In this embodiment, device 1 functions as an edge terminal and communicates with the generation processing device 3 of an external device, such as a cloud server, to optimize data processing and the entire system. Device 1 is equipped with a camera (image input unit 170), a microphone (voice input unit 180), a sensor, or other data acquisition means, and transmits the acquired data to the generation processing device 3 in real time or at a predetermined timing. As for the communication method, wired communication (Ethernet, optical fiber communication, serial communication, etc.) or wireless communication (WiFi, cellular communication, LPWA, etc.) can be applied.
[0068] The generation and processing unit 3 stores and analyzes data received from device 1, and sends processing results and control signals back to device 1 as needed. For example, it can perform machine learning or AI processing on the cloud side and send optimal control parameters to device 1. The generation and processing unit 3 may also play a role in integrating and managing data from multiple devices 1, optimizing the entire system, and detecting anomalies. Furthermore, device 1 and the generation and processing unit 3 may operate while communicating with each other, realizing efficient data processing through the collaboration of edge computing and cloud computing. Device 1 can improve the overall processing efficiency of the system by performing calculations requiring low latency locally and offloading computationally intensive processing to the generation and processing unit 3.
[0069] The control unit 300 may be implemented by a control circuit including a CPU, similar to the control unit 100.
[0070] The memory unit 350 may include a main memory unit and an auxiliary memory unit, similar to the memory unit 150. It stores an artificial intelligence model 352. The artificial intelligence model 352 may be a large-scale language model. The process performed by the artificial intelligence model 352 may utilize natural language processing (NLP) techniques, transformer models, or other generative artificial intelligence algorithms.
[0071] The communication unit 360 may transmit and receive information using one or more wireless communication technology and wired communication technology, similar to the communication unit 160.
[0072] The generation processing unit 112 transmits related information 503 to the generation processing unit 3, which includes the artificial intelligence model 352, via the communication unit 160. The generation processing unit 3 then generates related sentences using the artificial intelligence model 352 based on the related information 503. Here, related sentences are not limited to sentences but may also include words, signals, etc. The generation of related sentences may be performed according to one or more conditions stored in the storage unit 150 of the device 1 or the storage unit 350 of the generation processing unit 3 and user instructions input by the voice input unit 180. For example, the generation processing unit 112 may generate related sentences containing only the name of the building that the user is interested in from the related information 503 shown in Figure 5. Also, for example, if the user asks for the building's telephone number, the generation processing unit 112 may generate related sentences containing only the telephone number from the related information 503 shown in Figure 5. Here, the generation processing unit 112 converts the instructions and questions input by the user into prompts for the artificial intelligence model 352. The generation processing unit 112 then receives the related sentences generated by the artificial intelligence model 352 from the generation processing unit 3 via the communication unit 160. The generation processing unit 112 may also be configured to convert instructions or questions input by the user into search queries and send them to a database (including relational databases, vector databases, etc.) in the storage unit 150 or the related device 2, the generation processing unit 3, or other devices, and to receive the search results.
[0073] The voice output processing unit 114 outputs the related sentence as voice from the voice output unit 190. Specifically, the voice output processing unit 114 may convert the received text data of the related sentence into natural and clear speech using, for example, deep neural network-based speech synthesis (text-to-speech: TTS) technology. The converted speech is provided to the user through the voice output unit 190, such as earphones or speakers. During speech synthesis, emotional expression, speaking speed, naturalness of pronunciation, etc., may be taken into consideration, and additional information according to the user's condition (for example, barrier-free information such as necessary landmarks for users with physical difficulties) may be added. This allows the user to appropriately recognize the related sentence as voice.
[0074] In this embodiment, the generation processing unit 112 may generate a related sentence in a language desired by the user based on related information obtained from the matching determination of the first image and the second image via the generation processing device 3. For example, if the related information is written in Japanese and the user desires audio output in English, the generation processing unit 112 may generate an English related sentence from the Japanese related information based on conditions set in advance by the user or instructions given through the audio input unit 180.
[0075] In this embodiment, the generation processing device 3 is a different device from device 1, but in other embodiments, the functions of the generation processing device 3 may be realized in device 1, so-called on-device AI. For example, the functions of the control unit 300 of the generation processing device 3 may be realized in the control unit 100 of device 1, and some or all of the data stored in the storage unit 350 may be stored in the storage unit 150 of device 1, thereby realizing some or all of the functions of the generation processing device 3 described in this specification in device 1. In this case, device 1 may incorporate a dedicated accelerator such as a DSP (Digital Signal Processing) or NPU (Neural Processing Unit), and the control unit 100 of device 1 may perform simple image (video) analysis (face detection, object recognition, etc.) and speech analysis (ASR: Automatic Speech Recognition, NLP: Natural Language Processing, etc.) as a SoC (System on a Chip) with low power consumption. Specifically, object recognition and person recognition on a video frame basis, and speech recognition and dialogue models may be operated in real time and synchronously. This eliminates the need for a communication unit, resulting in faster processing, reduced power consumption and latency, and enabling operation offline or via low-bandwidth communication.
[0076] Furthermore, device 1 and the generation processing unit 3 may cooperate, and some functions of the generation processing unit 3 may be implemented in device 1. For example, some of the data stored in the memory unit 350 of the generation processing unit 3 may be stored in the memory unit 150 of device 1, and the artificial intelligence model (first artificial intelligence model) included in the memory unit 150 of device 1 and the artificial intelligence model 352 (second artificial intelligence model) included in the memory unit 350 of the generation processing unit 3 may be used in combination. For example, the generation processing unit may perform simple analysis of the first image (face detection, object detection, reading of character information (Optical Character Recognition / Reader: OCR), recording, searching, etc.) and speech analysis (ASR, NLP, etc.) using the first artificial intelligence model, and perform more computationally intensive processing using the second artificial intelligence model. Also, device 1 may minimize communication with the generation processing unit 3 (cloud), and confidential data may be encrypted at device 1 (edge). Furthermore, device 1 may perform initial data processing on-device, blocking faces and personally identifiable information before sending it to the generation processing unit 3 (cloud). This allows for the streamlining of the entire system's processing, such as handling some high-load tasks via cloud integration, while still considering battery capacity.
[0077] In these embodiments, if the device 1 implements some or all of the functions of the generation processing device 3, the configuration described as the generation processing device 3 in this specification may be read as the corresponding configuration of the device 1, such as the generation processing device 3 being read as the generation processing unit 112 of the device 1, and the storage unit 350 being read as the storage unit 150.
[0078] The audio acquisition unit 116 acquires audio from the audio input unit 180. During audio acquisition, preprocessing such as noise reduction, volume normalization, and sampling frequency conversion can be performed using digital signal processing technology.
[0079] The speech analysis unit 118 analyzes the speech acquired by the speech acquisition unit 116. The speech analysis unit 118 converts the speech into text data and performs advanced speech analysis such as sentiment analysis, speaker identification, and language identification. Specifically, it may use a Hidden Markov Model (HMM), Deep Neural Network (DNN), or a transformer-based speech recognition model to perform highly accurate speech analysis.
[0080] The second image acquisition unit 106 receives the second image 502 from the related device 2 via the communication unit 160 according to the user's intent based on the analyzed voice. For example, if the user inputs the voice "What is that?" to the voice input unit 180, and the voice analysis unit 118 analyzes the voice as a question about the first image, the second image may be acquired. Here, the user's intent may not only be the content of the text indicated by the voice spoken by the user (including instructions, questions, etc.), but also the user's emotions, thoughts, context, etc., inferred from the pitch, length, tone, etc., of the voice spoken by the user. Furthermore, the user's intent is not limited to the voice spoken by the user, but may also be an intent that the user "must want to do," inferred by AI, etc., from sounds other than the user's voice, such as ambient sounds. With this configuration, the amount of processing and memory capacity required to acquire the second image can be reduced, and the second image can be acquired and stored efficiently.
[0081] The second image acquisition unit 106 analyzes keywords, emotions, context, etc., extracted from the speech analysis results and receives the second image that best matches this information from the related device 2. The image selection algorithm may include a machine learning-based recommendation system, semantic image search technology, etc.
[0082] The route information processing unit 120 receives route information to the destination corresponding to the location information 254 from the related device 2 via the communication unit 160, and transmits a request for route information to the related device 2 via the communication unit 160. The processing of the route information processing unit 120 will be explained with reference to Figure 5. In Figure 5, the navigation 501 is shown. The navigation 501 includes the current location, the destination, and route information between the two. In Figure 5, the current location, destination, and route information between the two are schematically shown, but in this embodiment, each element shown in the navigation 501 may be coordinate system data.
[0083] The route information processing unit 120 transmits the user's location information 254, acquired by the location information acquisition unit 104, to the related device 2 via the communication unit 160. The navigation processing unit 202 of the related device 2 acquires route information, including the optimal route to the destination, recommended route, and traffic information, in accordance with the location information 254. The acquired route information may include information such as the distance of the route, estimated travel time, recommended means of transport, and waypoints.
[0084] The route information processing unit 120 may output the route information as an AR (Augmented Reality) overlay from the audio output unit 190. This allows the route information processing unit 120 to provide the user with dynamic route guidance.
[0085] If the determination unit 108 determines that the two images do not match, the route information processing unit 120 may send a request to change the route information to the related device 2 via the communication unit. For example, if the first image in the user's line of sight does not match the second image associated with the location information at the location where the first image was acquired, there may be obstacles, road construction, temporary road closures, etc., at that location. When such an image mismatch is detected, the route information processing unit 120 sends a request to the related device 2 to reacquire more appropriate route information. This request may include the current location information, the latest information (such as the period in which the current time is included) such as traffic information related to that location, detailed information about the image mismatch, and the user's movement status. Based on this information, the related device 2 can dynamically update and optimize the route information.
[0086] Furthermore, the route information processing unit 120 may determine the error with the route information by inferring changes in the user's direction of travel or estimating the user's posture or movement from the first image (real-time video) continuously acquired by the first image acquisition unit 102. If the route information processing unit 120 determines that there is an error with the route information, it may send a request for correction of the error to the related device 2.
[0087] Figure 7 is a schematic diagram of the operation flow in one embodiment of the present invention. Each step will be described in detail below.
[0088] In step S701, the first image, which is an image of the field of view from the user's line of sight, is acquired and stored in the storage unit 150. Specifically, the image input unit 170 captures the first image from the user's viewpoint, and the first image acquisition unit 102 stores this image as the first image in the storage unit 150. The first image represents the initial state of the field of view that the user is currently focusing on.
[0089] In step S702, the system acquires the user's location information and stores it in the storage unit 150. Specifically, the location information acquisition unit 104 acquires location information to identify the user's current location. The acquired location information is stored in the storage unit 150 in association with the first image. This location information functions as context information in subsequent image processing.
[0090] In step S703, the acquired location information is transmitted to the related device 2, and the second image associated with the location information is received. Specifically, the previously acquired location information is transmitted to the related device 2 via the communication unit 160. Based on the transmitted location information, the related device 2 searches for or generates the corresponding second image and transmits the second image to device 1.
[0091] In step S704, the first image and the second image are compared to determine whether the two images match. Specifically, the determination unit 108 compares the feature quantities, pixel information, or feature points of the two images. This comparison process determines whether the two images represent the same place, object, field of view, etc.
[0092] In step S704, if it is determined that the images do not match (No in step S704), the process may be terminated, or the process may be repeated by returning to the start according to predetermined conditions. On the other hand, if it is determined that the images do match (Yes in step S704), the process proceeds to step S705. In step S705, related information associated with the second image is received from the related device 2 and stored in the storage unit 150. This related information may include, for example, the name, address, shooting time, detailed metadata of the image, and position information correction information of the object projected onto the second image.
[0093] According to the above embodiment, it is possible to accurately associate the user's gaze direction with image information of their current location, and to efficiently acquire information that is visible in the user's field of view based on the location information, and to provide information in a contextual manner.
[0094] Figure 8 is a schematic diagram of the operation flow in one embodiment of the present invention. Each step will be described in detail below.
[0095] In step S801, the relevant information is transmitted to the generation processing unit 3, and the relevant sentences generated by the artificial intelligence model 352 are received. Specifically, the generation processing unit 112 transmits the previously acquired relevant information (location information, image information, etc.) to the generation processing unit 3 via the communication unit 160. Based on the received information, the generation processing unit 3 automatically generates the relevant sentences using the artificial intelligence model 352, which employs machine learning, deep learning, and other techniques.
[0096] In step S802, the generated related sentence is output as audio from the audio output unit 190. Specifically, the audio output processing unit 114 converts the received text data of the related sentence into natural and clear audio using, for example, speech synthesis technology using a deep neural network. The converted audio is output from the audio output unit 190.
[0097] According to the above embodiment, by automatically generating relevant documents using artificial intelligence technology based on acquired information, and further outputting those documents as audio, it becomes possible to provide users with more interactive and contextually relevant information.
[0098] Figure 9 is a schematic diagram of the operation flow in one embodiment of the present invention. Each step will be described in detail below.
[0099] In step S901, the process of acquiring sound from the audio input unit 180 is performed. Specifically, the audio acquisition unit 116 acquires the user's voice or ambient sounds, etc., via the audio input unit 180, such as a microphone or audio input device. When acquiring sound, pre-processing such as noise reduction, volume normalization, and sampling frequency conversion may be performed using digital signal processing technology.
[0100] In step S902, the acquired audio is analyzed. Specifically, the audio analysis unit 118 converts the audio into text data and performs advanced audio analysis such as sentiment analysis, speaker identification, and language identification. Specifically, a Hidden Markov Model (HMM), Deep Neural Network (DNN), or a transformer-based audio recognition model may be used to perform highly accurate audio analysis.
[0101] In step S903, the system processes the reception of a second image based on the analyzed speech, according to the user's intent. The second image acquisition unit 106 analyzes keywords, emotions, context, etc., extracted from the speech analysis results, and receives the second image that best matches this information from the associated device 2. The image selection algorithm may include a machine learning-based recommendation system, semantic image search technology, etc.
[0102] According to the above embodiment, it becomes possible to utilize the diverse information obtained from the user's voice input and its analysis to provide information that is highly matched to the user's intentions and emotions.
[0103] Figure 10 is a schematic diagram of the operation flow in one embodiment of the present invention. Each step will be described in detail below.
[0104] In step S1001, the system receives route information to the destination from the related device 2 based on the acquired location information and sends a request for route information to the related device 2. Specifically, the route information processing unit 120 transmits the location information to the related device 2 via the communication unit 160. The navigation processing unit 202 of the related device 2 acquires route information, including the optimal route to the destination, recommended route, and traffic information, based on the location information. The acquired route information may include information such as the distance of the route, estimated travel time, recommended means of transport, and waypoints.
[0105] In step S1002, the first and second images are compared to determine whether they match. The control unit uses an image processing algorithm to compare the features, pixel information, feature points, and structural similarity of the two images in detail. Techniques such as image matching, feature extraction, and image similarity evaluation using a convolutional neural network (CNN) can be used for the comparison method.
[0106] In step S1002, if it is determined that the images match (Yes in step S1002), the process may be terminated, or the process may be returned to the start and repeated according to predetermined conditions. On the other hand, if it is determined that the images do not match (No in step S1002), the process proceeds to step S1003. In step S1003, a request to change the route information is sent to the related device 2. Specifically, if an image mismatch is detected, a change request is sent to reacquire more appropriate route information. This request may include the current location information, detailed information about the image mismatch, and the user's movement status. Based on this information, the related device 2 can dynamically update and optimize the route information.
[0107] According to the above embodiments, it becomes possible to achieve more dynamic and adaptive route navigation by combining location information and real-time image information.
[0108] As described above, according to one embodiment of the present invention, more accurate and contextual information can be provided in real time according to the user's current location, field of view, intentions conveyed through voice, etc.
[0109] In addition to the embodiments and their modifications and other embodiments described above, the present invention can also be realized in the following modifications or other embodiments.
[0110] In this embodiment, the device 1 is described as constituting the present invention, but the present invention may be configured as a system (for example, system S shown in Figure 1) that includes one or more related devices 2 and generation processing devices 3 in addition to the device 1. Specifically, the system according to the present invention is a system in which a device, a related device, and a generation processing device are connected to communicate via a network, wherein the device comprises a control unit, a storage unit, a communication unit, and an image input unit, and the control unit may include a first image acquisition unit that acquires a first image which is an image in the direction of the user's line of sight from the image input unit and stores it in the storage unit, a location information acquisition unit that acquires the user's location information and stores it in the storage unit, a second image acquisition unit that transmits the location information to the related device via the communication unit and receives a second image which is an image associated with the location information in the related device from the related device via the communication unit, a determination unit that compares the first image and the second image and determines whether they match, and if they match, a related information acquisition unit that receives related information which is information associated with the second image from the related device via the communication unit and stores it in the storage unit. In addition, the system according to the present invention may consist of one or more of the matters described herein.
[0111] In this embodiment, Device 1 is described as an earphone-type device worn in the user's ear, but Device 1 is not limited to this and may be implemented using other computers. Other computers may include portable computers such as PDAs, microcomputers, smartphones, and wearable computers. Furthermore, other computers may include personal computers, general-purpose computers, workstations, and supercomputers. In this case, some functions of Device 1 may be implemented on the device carried by the user, while other functions may be implemented on the computer that is connected to the device in a communicative manner.
[0112] In this embodiment, the device 1 is described as comprising an image input unit 170 and an audio input unit 180 as the main input devices, and an audio output unit 190 as the main output device. However, in other embodiments, the device 1 may be connected to other input devices such as a keyboard, mouse, touch panel, or scanner. The device 1 may also be connected to other output devices such as a display or printer. Data can be transmitted and received between the device 1 and other input devices and other external devices using wireless communication technologies including WiFi, 5G, and Bluetooth, or other appropriate wireless communication protocols. Furthermore, if a wired connection is required, it is possible to connect via a compatible physical interface such as USB, Ethernet, HDMI, or serial communication (RS-232, RS-485, etc.) and perform bidirectional communication.
[0113] The control unit 100 may further include an encryption processing unit. If the encryption processing unit determines that the data acquired from the image input unit 170 and the audio input unit 180 contains confidential data, it may encrypt the confidential data. Furthermore, if the encryption processing unit determines that the data acquired from the image input unit 170 and the audio input unit 180 contains information that can identify a face or an individual, it may process the data to block such information before transmitting it to another device. These encryption and access control functions make it possible to create a system that prevents data that is constantly being recorded from being leaked to the outside, resulting in a design that takes into consideration the privacy of not only the user but also those around them, and enables the appropriate handling of highly confidential personal information and other sensitive data.
[0114] In this embodiment, the control unit 100 of the device 1 stores data and information such as the first image, location information, and related information in the storage unit 150. However, in other embodiments, the control unit 100 may store some or all of the data and information in the storage unit of the related device 2, the generation and processing device 3, or other external devices. Other external devices may include, for example, other devices owned by the user, such as a smartphone, necklace-type terminal, pendant-type terminal, smartwatch, earring-type terminal, bracelet-type terminal, hairband-type terminal, or personal computer, and may be implemented in the form of one or more magnetic tapes, magnetic disks, optical disks, semiconductor disks, etc. Device 1 and other external devices can transmit and receive data using wireless communication technologies including WiFi, 5G, Bluetooth, or other appropriate wireless communication protocols. Furthermore, if a wired connection is required, it is also possible to connect via a compatible physical interface such as USB, Ethernet, HDMI, or serial communication (RS-232, RS-485, etc.) and perform bidirectional communication. The control unit 100 stores one or more of the images and information in one or more storage units of the external device described above, thereby minimizing the storage capacity of the device 1 and reducing its weight, while also enabling the appropriate storage and use of necessary data and information.
[0115] In this embodiment, the determination unit 108 performs a matching determination between the first image and the second image, and if a matching is determined, the related information acquisition unit 110 acquires related information. However, in other embodiments, the control unit 100 may estimate the user's current position and orientation from the result of the matching determination between the first image and the second image, and output the estimated result as audio from the audio output unit 190. This allows the user to know their current position and orientation more accurately based on the objects they are seeing.
[0116] In this embodiment, the determination unit 108 performs a matching determination between the first image and the second image, and if a matching is determined, the related information acquisition unit 110 acquires related information, and the generation processing unit 112 transmits the related information to the generation processing unit 3 to generate a related sentence. However, in other embodiments, the generation processing unit 108 may generate a related sentence via the generation processing unit 3 based on input in the device 1 or predetermined conditions, regardless of the processing of the determination unit 108. For example, the generation processing unit 112 may perform the following inference or generation via the generation processing unit 3.
[0117] (a) The generation processing unit 108 first performs inference or generation on the data input to the device 1 using the artificial intelligence model 352 of the generation processing device 3, reflects the result in the first image, performs a match determination by the determination unit 108, and then infers or generates the determination result using the artificial intelligence model 352 before outputting the result. This allows the determination unit 108 to perform determination after appropriate processing of the target data, and further optimizes the output of the determination result in a format that is easy for the user to understand.
[0118] (b) The generation processing unit 108 first performs inference or generation on the data input to the device 1 using the artificial intelligence model 352 of the generation processing device 3, reflects the result in the first image, performs a matching determination by the determination unit 108, and then outputs the result. This allows the determination unit 108 to perform a determination after appropriate processing of the target data, and the processing speed of the subsequent output can be increased.
[0119] (c) The generation processing unit 108 performs a matching determination on the data input to the device 1, and then infers or generates the determination result using the artificial intelligence model 352 before outputting the result. This increases the processing speed up to the determination and also allows for optimization such as outputting the determination result in a format that is easy for the user to understand.
[0120] (d) The generation processing unit 108 outputs the result of inference or generation by the artificial intelligence model 352 to the data input to the device 1 without judgment by the judgment unit 108. This makes it possible to perform processing by the artificial intelligence model only when image judgment is not required to produce an appropriate output for the input.
[0121] The processing performed by the generation processing unit 108 may be carried out in response to user instructions via the voice input unit 180, or it may be carried out automatically according to predetermined conditions even without user instructions.
[0122] In this embodiment, the first image acquisition unit 102 and the storage unit 150 may function as a drive recorder for security and recording purposes. In this case, the first image input unit 170 may include camera modules for multiple directions, and may also record inputs from the audio input unit 180 and other sensors mentioned above (such as biosignals like acceleration, electric potential, smell, temperature, and pulse). The control unit 100 may also perform a selective information extraction algorithm, a context-dependent behavior analysis function, a dynamic information importance evaluation function, etc. The storage unit 150 may perform efficient data storage by compression or tokenization when recording continuously, semantic labeling, automatic tagging of abnormal or specific behaviors (e.g., determination of harassment behavior according to behavior patterns, detection of suspicious behavior, detection of anomalies based on environmental changes, etc.), dynamic data retention and deletion based on user requests, and data storage based on the determination results of the determination unit 108. The audio output unit 190 or other output device may perform selective output of analysis results. These configurations enable selective information recording with consideration for privacy, efficient searching and management of real-time video, and flexible information processing based on user requests.
[0123] The artificial intelligence model 352 in this embodiment may be an artificial intelligence model trained on the user's personal data (hereinafter also referred to as "personal AI agent"). Personal data includes, for example, the user's voice obtained with the user's consent, images in which the user is the subject or images specified by the user, and text created or specified by the user. Specifically, personal data may be the user's medical information. Medical information may include the user's medical examination results at a medical institution, electronic medical records, voice and image data from daily life, and inferred data such as emotions and medical condition based on pulse data. Medical information may also include medical statistics obtained from a population similar to the user's attributes (age, gender, body type, height, weight, etc.). Further examples of personal data are also described in the following processing example.
[0124] Furthermore, an artificial intelligence model trained on personal data includes, for example, an artificial intelligence model trained using the above-mentioned personal data as training data, and such an artificial intelligence model includes neural networks, random forests, large-scale language models, etc. In the following embodiments, the artificial intelligence model is described as a large-scale language model, but embodiments using other artificial intelligence models are also included in the disclosure of this specification.
[0125] A personal AI agent is formed by acquiring and storing personal data, and training an artificial intelligence model using that personal data. Personal data may be acquired not only from the device 1 described in this embodiment, but also from other devices used by the user. Furthermore, personal data may be acquired from a data storage service (e.g., an information fund, information trust, information bank, etc.) with the user's consent.
[0126] Furthermore, personal data may be stored in the storage unit 150 of device 1. For example, profile data such as the user's interests, preferences, and conversation style may be stored in the storage unit 150 of device 1. The personal data stored in the storage unit 150 of device 1 may be periodically synchronized with personal data stored in related device 2, generation and processing device 3, or other devices.
[0127] User consent to personal data may be given each time data is collected, or it may be given selectively, comprehensively, continuously, or conditionally for a specific period, subject, use, etc.
[0128] Here, we will explain an example of the operation flow in which the generation processing unit 3 performs training, inference, or generation of a personal AI agent.
[0129] In step 1, the generation and processing device 3 receives personal data of users who have agreed to use their personal data as training data from the image input unit 170, voice input unit 180, various sensors, and other input devices of the device 1, or acquires it from a data storage service, and stores it in the storage unit 350.
[0130] Furthermore, in the data storage service, data may be identified as belonging to a specific user by a unique user number, and data classification may be improved by storing it with tags. In addition, each data item may be stored in vector form to enable similarity searches.
[0131] In step 2, the generation processing device 3 trains the artificial intelligence model 352 using personal data as training data. Here, the personal data used as training data may include, for example, the model of the home appliance purchased by the user, the date of purchase, the region of purchase, and the store of purchase. In this case, the trained artificial intelligence model 352 (personal AI agent) can perform inferences such as recommendations and generate answers to questions based on the home appliance owned by the user during inference and generation processes. Furthermore, the personal data used as training data may also include, for example, the user's way of speaking, catchphrases, and expressions. As a result of long-term learning, the personal AI agent can express the user's conversational style and values.
[0132] In step 3, the generation processing device 3 performs processing such as inference or generation using the trained artificial intelligence model 352 (personal AI agent). The generation processing device 3 inputs the acquired data into the personal AI agent to perform inference or generation based on the user's personal data. For example, the generation processing device 3 may use the personal AI agent to infer the user's emotions, mental state, preferred information and responses, objects and places of interest, etc., or to infer the situation the user is in (for example, whether they are under stress or facing an emergency based on data such as heart rate). It may also generate relevant sentences and other content according to the user's emotions, etc., and situation. For example, it may estimate and record the user's food preferences, work style, schedule management characteristics, etc., and propose appropriate scheduling. Furthermore, the personal AI agent may generate multifaceted information related to the image subject, such as dynamically generating relevant information, background, and contextual information about the image subject using a multimodal language model. Furthermore, the personal AI agent may assist the user's understanding by referring to cloud-based databases and internet information based on the object identified in the first image and providing appropriate information in real time (for example, outputting it as voice from the voice output unit 190 of device 1). In addition, from the perspective of privacy protection, the personal AI agent may include anonymization and authentication mechanisms to ensure user consent and information security. With these configurations, it becomes possible to perform processing such as inference or generation according to the user's preferences, and furthermore, it becomes possible to easily obtain relevant information and perform appropriate processing from actions such as the user focusing on an object.
[0133] Here, we describe an example of processing such as inference or generation using the personal AI agent in this embodiment. In the following, the operating entity is described as the generation processing unit 3, but the generation processing unit 112 of device 1 may also perform the following inference or generation via the generation processing unit 3.
[0134] (Reminders, Task Management): The generation processing device 3 may use a personal AI agent to add tasks based on the user's spoken language input from the voice input unit 180 of the device 1, and automatically notify the user of tasks with approaching deadlines (for example, by outputting voice from the voice output unit 190 of the device 1). In this case, the personal AI agent has the function of automatically identifying or extracting the content, priority, and deadline of tasks from the user's speech using speech recognition technology, natural language processing technology, etc., and registering them in a schedule management system (not shown). Furthermore, a machine learning algorithm is used to learn the user's past task management patterns and realize more accurate reminder notifications. For example, the deadlines of registered tasks can be dynamically managed, and reminders according to priority can be presented at the appropriate time. In addition, it is possible to link with calendar applications and cloud task management systems to improve the user's work efficiency.
[0135] (Suggestion, Advice): The generation processing device 3 may use a personal AI agent to analyze the user's values and past behavioral history based on personal data input from the image input unit 170, voice input unit 180, various sensors, and other input devices of the device 1, as well as personal data obtained from data storage services, and present options that contribute to decision-making. In this case, the personal AI agent analyzes or inputs the user's attribute information, past selection history, speech patterns, and contextual information into a machine learning model, and uses an advanced inference algorithm to generate personalized suggestions and advice based on individual preferences and values. For example, in purchasing decisions, when the user inputs "Which one should I buy?" into the voice input unit 180 of the device 1, a first image including the item to be purchased is input from the image input unit 170, the generation processing unit 112 identifies the item to be purchased through product image recognition via the generation processing device 3, and uses a recommendation algorithm to comprehensively evaluate price, quality, past purchase history, etc., and proposes similar products or optimal options in real time (for example, outputting them as voice from the voice output unit 190 of the device 1). Furthermore, in a dining scenario, for example, the personal AI agent may utilize nutritional analysis technology and recommendation algorithms to analyze ingredient and calorie information from the menu acquired from the image input unit 170 of the device 1, and provide personalized recommended menus, healthy options, and nutritional information, taking into account past meal history. In addition, it may have functions to support user decision-making, such as quantitatively evaluating the advantages and disadvantages of multiple alternatives and including them in the above suggestions. This configuration makes it possible to present the optimal option that contributes to the user's decision-making.
[0136] (Automatic Summarization and Review): The generation processing device 3 may use a personal AI agent to automatically summarize recordings of meetings and conversations, such as online meetings and face-to-face meetings, input from the voice input unit 180 of the device 1, and automatically generate logs for reviewing daily activities. In this case, the personal AI agent utilizes speech recognition technology, natural language processing technology, and a summary generation algorithm to accurately extract the context, importance, and key points of the utterances and generate a concise summary. Furthermore, by performing multimodal analysis on data input from the image input unit 170, voice input unit 180, various sensors, and other input devices of the device 1, the device may comprehensively analyze audio including sounds made by participants and ambient sounds in addition to utterances during meetings, text including charts and graphs analyzed from the first image input from the image input unit 170 of the device 1, such as whiteboards and presentation materials during meetings, and non-verbal information input from various sensors, etc. A machine learning model may automatically determine high-priority topics and tasks based on the context of the meeting to create more accurate summaries and follow-up documents. In addition, the device may record and analyze the user's behavior history and utterances to create logs for reviewing daily activities. The generated summaries and logs are stored in the storage unit 150 of device 1 and may be output as voice from the voice output unit 190 in response to user instructions via the voice input unit 180, or as text from another output device. This configuration allows for the easy creation of highly accurate summaries, improving work efficiency and contributing to improved lifestyles through log review.
[0137] (Information Retrieval and Integration): The generation processing device 3 may use a personal AI agent to integrate with the user's devices (e.g., smartphones) and cloud services, etc., and realize comprehensive information retrieval and sharing functions solely through the voice interface provided by the voice input unit 180 of the device 1. In this case, the personal AI agent integrates with the user's devices and databases or internet search engines on the cloud, etc., and uses cross-data source search technology, semantic search algorithms, and natural language understanding technology to identify accurate information from the user's spoken content, ambiguous voice input, etc., and present it in the most optimal format (for example, output as voice from the voice output unit 190 of the device 1). For example, the generation processing device 3 may retrieve and generate an operating method for an identified electrical product (the electrical product may be identified by image search and matching via network N, comparison with a pre-stored list of electrical products used by the user, estimation from the user's location information (placement of electrical products), or other methods), and output the result as voice from the device 1. Here, if the image of an electrical product in the user's line of sight has features such as an illuminated indicator light, the output of the operation method may estimate and output the necessary operation method from an image of a manual or other document related to those features. Furthermore, for privacy protection, a personal information access control function may be implemented. The acquired information is stored in the storage unit 150 of device 1 and may be output as voice from the voice output unit 190 in response to user instructions via the voice input unit 180, or as text from another output device, or may be shared via email or chat application on a linked device. This configuration can improve the efficiency of information utilization.
[0138] (Support for everyday conversation): The generation processing device 3 may use a personal AI agent to anticipate the user's thoughts and propose supplementary information and appropriate expressions in real time (for example, output as voice from the voice output unit 190 of device 1) in the user's everyday conversation input from the voice input unit 180 of device 1 (for example, output as voice from the voice output unit 190 of device 1). In this case, the personal AI agent uses speech recognition, natural language processing, dialogue context analysis, emotion recognition technology, language models, etc. to analyze the context of the conversation and predictively understand the flow. The personal AI agent detects ambiguity in the topic and unclear language expressions and provides advanced dialogue support functions that complement the user's intentions. With this configuration, appropriate information and expressions can be presented immediately. In addition, if the user has difficulty putting their thoughts into words, the device can support smooth dialogue by organizing them into concise phrases.
[0139] (Healthcare, Mental Care): The generation processing device 3 may use a personal AI agent to estimate the user's fatigue level and stress level by analyzing images (video), audio, and other signals input from the image input unit 170, audio input unit 180, various sensors, and other input devices of the device 1, and provide appropriate health management advice (for example, output as audio from the audio output unit 190 of the device 1). In this case, the personal AI agent uses facial recognition, voice tone analysis, biosignal monitoring technology, etc., to analyze changes in facial expressions, voice, posture, etc., and continuously evaluate and monitor the physical and mental health status. If necessary, it will suggest rest, stretching, and relaxation techniques to support preventive healthcare and mental care. Furthermore, it may automatically record morning health checks and nighttime sleep reviews to support the user's health management and perform mental care dialogue at appropriate times. This configuration can support the user's daily health management.
[0140] (Learning, Skill Improvement): The generation processing device 3 may use a personal AI agent to automatically generate a learning program optimized for the individual, such as for language learning or presentation practice, from everyday conversations and actions acquired through input from the image input unit 170, voice input unit 180, various sensors, and other input devices of the device 1. In this case, the personal AI agent uses natural language processing, speech analysis, and gesture recognition technology to analyze the user's speech content, learning style, and communication characteristics. The personal AI agent then detects areas for improvement in the individual, such as language learning, presentation skills, and nonverbal communication, and provides customized feedback and learning content for improvement, such as suggesting more natural expressions in language learning (for example, outputting them as voice from the voice output unit 190 of the device 1), and evaluating the rhythm of speech and the appropriateness of gestures in presentation practice. This configuration can support the user in improving their effective speaking and presentation skills.
[0141] (Estimation of health status based on user voice): The generation processing device 3 may use a personal AI agent to estimate the user's medical status by analyzing the user's voice input from the voice input unit 180. In this case, the personal AI agent uses voice biometric authentication technology, voice parameter analysis, and multivariate biosignal processing algorithms to estimate the health status from the subtle characteristics of the voice. If the health status claimed by the user differs from the health status inferred from the analysis of the voice and other data input from device 1, and comparison and matching with past personal data, the personal AI agent may output the difference (for example, outputting it as voice from the voice output unit 190 of device 1). Specifically, the personal AI agent may include a contradiction detection algorithm (statistical significance analysis of subjective declarations and objective data, anomaly detection models using machine learning, context-dependent inference engines, etc.) and a notification mechanism (natural language explanation via the voice output unit 190, reliability scoring by data source, stepwise information disclosure considering privacy, etc.) as difference detection. For example, if a user says "I feel fine," but the personal AI agent infers that the user is unwell based on the tone of voice, visual images, pulse rate, and other biometric data, the personal AI agent may generate a related sentence that highlights the difference between the user's stated medical condition and the inferred medical condition, and output it as voice from the voice output unit 190.
[0142] According to the personal AI agent described above, it can deeply learn the user's thoughts, behaviors, interests, etc., based on the user's conversations and audio / image data obtained from their field of vision. As a result, the personal AI agent can behave like a surrogate of the user, providing natural conversation and advanced support.
[0143] Although the above description assumes that the generation processing device 3 includes a personal AI agent, the personal AI agent may be stored in the storage unit 150 of the device 1 and the above processes may be executed.
[0144] Furthermore, the processing according to the present invention can be realized not only by hardware configuration but also by software or programs. Specifically, the various functions of the present invention are controlled by a program executed by the processor, thereby controlling the operation of the device.
[0145] The program according to the present invention may be stored in the memory or storage device (HDD, SSD, flash memory, etc.) of a computer and read and executed by a processor. Removable storage such as CD-ROMs, DVD-ROMs, USB memory sticks, and SD cards, as well as network-based storage methods such as cloud storage, can be used as storage media for the program. Furthermore, the program according to the present invention can also be provided in a form that is downloaded from a server via a network and installed on a terminal device. This allows users to apply the latest version of the program through online updates.
[0146] The program of the present invention can be implemented as an application program running on an operating system (OS), or as firmware for embedded systems. It can also be provided as code that runs in a scripting language or on a virtual machine. In addition to running independently, the program of the present invention is also envisioned to function in conjunction with other software modules and external systems. Examples include data exchange with other systems via APIs and operation in a distributed processing environment based on a microservices architecture.
[0147] The program's processing includes, based on arithmetic processing by a processor, writing data to a storage device, data communication over a network, and displaying and accepting information input through a user interface. It may also perform data analysis, recognition, and inference processing using machine learning models and AI algorithms.
[0148] The present invention is not limited to the program form described above, but can also be applied to various software components such as firmware, middleware, and driver software, and can operate integrally as part of the entire system.
[0149] It should be noted that the above embodiments are merely representative examples of the present invention, and the technical scope of the present invention is not limited to the above specific descriptions. It will be clear to those skilled in the art that various modifications and changes are possible without departing from the spirit of the present invention. [Explanation of Symbols]
[0150] S System N Network 1 device 2 Related devices 3. Generating Processing Unit 11. Direction of the user's line of sight 100, 200, 300 Control Unit 102 First Image Acquisition Unit 104 Location information acquisition unit 106 Second Image Acquisition Unit 108 Judgment section 110 Related Information Acquisition Unit 112 Generation Processing Unit 114 Audio output processing unit 116 Voice acquisition unit 118 Voice Analysis Unit 120 Route Information Processing Unit 150, 250, 350 storage section 160, 260, 360 Communications Department 170 Image Input Section 180 Voice Input Section 190 Audio output section 202 Navigation Processing Unit 252 databases 254 Location information 352 Artificial Intelligence Models 501 Navigation 502 Image 2 503 Related Information
Claims
1. A device comprising a control unit, a storage unit, a communication unit, and an image input unit, The control unit, A first image acquisition unit acquires a first image from the image input unit, which is an image in the direction of the user's line of sight, and stores it in the storage unit. A location information acquisition unit that acquires the user's location information and stores it in the storage unit, A second image acquisition unit transmits the location information to an associated device via the communication unit, and receives a second image, which is an image associated with the location information, from the associated device via the communication unit in the associated device. A determination unit that compares the first image and the second image and determines whether they match, If a match is determined, the system includes a related information acquisition unit that receives related information, which is information associated with the second image, from the related device via the communication unit and stores it in the storage unit, The control unit, The route information processing unit further includes receiving route information to a destination corresponding to the location information from the related device via the communication unit, and transmitting a request regarding the route information to the related device via the communication unit. If the determination unit determines that they do not match, the route information processing unit transmits a request to change the route information to the related device via the communication unit.
2. The apparatus according to claim 1, comprising an audio output unit and not comprising an image display unit, The apparatus according to claim 1, wherein the determination unit causes the result of the determination to be output as sound to the sound output unit.
3. The control unit, A generation processing unit transmits the related information via the communication unit to a generation processing unit including an artificial intelligence model, and receives the related sentences generated by the artificial intelligence model based on the related information from the generation processing unit via the communication unit. The apparatus according to claim 2, further comprising: an audio output processing unit that outputs the aforementioned related sentence as audio from the audio output unit.
4. It also features an audio input section, The control unit, A voice acquisition unit that acquires sound from the aforementioned voice input unit, It further includes a voice analysis unit that analyzes the aforementioned voice, The apparatus according to claim 3, wherein the second image acquisition unit receives the second image from the associated device via the communication unit in accordance with the user's intent based on the analyzed voice.
5. The apparatus according to claim 1 or 2, wherein the position information acquisition unit corrects the position information according to the determination result of the determination unit.
6. A method performed by a computer comprising a control unit, a storage unit, a communication unit, and an image input unit, The control unit, The steps include acquiring a first image from the image input unit, which is an image of the user's line of sight, and storing it in the storage unit, The steps include: acquiring the user's location information and storing it in the storage unit; The steps include transmitting the location information to the associated device via the communication unit, and receiving a second image, which is an image associated with the location information, from the associated device via the communication unit in the associated device, A step of comparing the first image and the second image and determining whether they match, If a match is determined, the process includes the step of receiving related information, which is information associated with the second image, from the related device via the communication unit and storing it in the storage unit, The control unit, The further step includes receiving route information to a destination corresponding to the location information from the associated device via the communication unit, and transmitting a request regarding the route information to the associated device via the communication unit, The method, wherein if it is determined in the determination step that they do not match, a request to change the route information is transmitted to the related device via the communication unit.
7. The method according to claim 6, which is performed by a computer equipped with an audio output unit and not equipped with an image display unit, The method according to claim 6, wherein the step of making the determination includes the step of causing the result of the determination to be output as sound to the sound output unit.
8. A program that causes a computer to perform the method described in claim 6 or 7.
Citation Information
Patent Citations
Cross-reality system for map processing using multi-resolution frame descriptors
CN115398314A
Information providing system and mobile temrinal
JP2006119797A
Knowledge information processing server system with image recognition system
JP2013088906A
Design article recognition information providing system, design article recognition information providing method, and design article recognition information providing program
JP2022083256A
Visual citations for information provided in response to multimodal queries
JP2024163063A