Image-to-speech system, device, method, and response output system

The image reading system with multiple cameras and AI analysis addresses the limitations of existing technologies by automatically capturing and vocalizing surrounding objects, enhancing safety and convenience for visually impaired users while walking.

WO2025206343A1PCT designated stage Publication Date: 2025-10-02MAXELL LTD
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/012880
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-08
Filing Date
2025-03-28
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing image reading technologies, such as the 'Seeing AI' app, are not optimized for use while a user is walking, requiring frequent manual input operations and struggling to effectively capture complex surroundings with a single camera.

Method used

An image reading system with multiple cameras that automatically captures images periodically, analyzes them using AI to detect objects, and reads out relevant information at predetermined times without user input, supporting visually impaired individuals by providing continuous surroundings awareness.

Benefits of technology

Enables hands-free, continuous detection and vocalization of surrounding objects, enhancing safety and convenience for visually impaired users by reducing the need for manual operation and improving the capture of complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025012880_02102025_PF_FP_ABST
    Figure JP2025012880_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides image-to-speech technology that can be used more suitably. An image-to-speech system according to the present invention includes a device that is carried by or worn by a user and performs speaking of objects from images captured by a camera. The device repeatedly automatically captures images from video of the camera at predetermined timings (step S2). The captured images are analyzed to acquire information including text representing objects in the images (step S3). An object and text to be spoken are decided in a predetermined determination, on the basis of the acquired information (step S4). Speaking of the text representing the object that is decided, is automatically repeated from the device at predetermined timings (step S5).
Need to check novelty before this filing date? Find Prior Art

Description

Image reading system, device, and method, and response output system

[0001] The present disclosure relates to a technology for converting images and videos into text, character strings, etc., and reading them aloud.

[0002] There is a technology (sometimes referred to as image reading or voice reading) that uses AI (artificial intelligence) to recognize and analyze images and videos (including still images and videos) captured by a camera, convert them into text, character strings, etc. (in other words, create sentences), and read them aloud. For example, one such technology is the "Seeing AI" app for iPhones provided by Microsoft, as shown in Non-Patent Document 1. The "Seeing AI" app uses AI to analyze photographic images captured by a camera, convert them into text, and read them aloud. Such technology is effective, for example, for supporting the visually impaired.

[0003] Further, examples of prior art include Japanese Patent Application Laid-Open No. 2000-325389 (Patent Document 1), Japanese Patent Application Laid-Open No. 2006-251596 (Patent Document 2), and Japanese Patent Application Laid-Open No. 2017-77446 (Patent Document 3).

[0004] JP 2000-325389 A JP 2006-251596 A JP 2017-77446 A

[0005] <URL: https: / / www.microsoft.com / ja-jp / ai / seeing-ai>, “Seeing AI”

[0006] For example, the "Seeing AI" app converts images captured by a user into text. The "Seeing AI" app also has a "Scene" function that allows users to take photos and describe the captured scene.

[0007] However, such technology does not sufficiently consider how to make it more suitable for use while the user is walking, and there is room for improvement. For example, the user's input operations, such as a capture operation, are time-consuming. As the user's surroundings change over time as the user walks, the user must input operations multiple times in order to grasp the surroundings over time through audible reading of the camera image. Furthermore, as the user walks, the surroundings may be complex, with various objects in front, behind, left, right, and so on. Therefore, grasping the complex surroundings through audible reading of the camera image using a single camera is difficult.

[0008] Prior art examples such as Patent Document 1 disclose that, for example, a single camera on a smartphone performs image reading or the like in response to a user's capture operation. However, the prior art examples do not provide specific embodiments on how to best support visually impaired persons using images captured by the camera. Furthermore, the prior art examples do not consider preferred embodiments for when a person wears the camera. Furthermore, the prior art examples do not sufficiently consider preferred functions and processes tailored to usage scenarios such as supporting visually impaired persons, or user convenience, with regard to those embodiments.

[0009] An object of the present disclosure is to provide a technology that can be more suitably used with respect to the image reading-out technology.

[0010] A representative embodiment of the present disclosure has the following configuration: One embodiment is an image reading system including a device carried or worn by a user, which reads out aloud an object from an image captured by a camera, wherein the device automatically and repeatedly captures images from the camera's video at predetermined times, analyzes the captured images to obtain information including text representing the object in the image, determines the object and text to be read out based on the obtained information at a predetermined judgment, and automatically and repeatedly reads out the text representing the determined object from the device at predetermined times.

[0011] According to a representative embodiment of the present disclosure, it is possible to provide a technology that can be used more preferably with respect to the image reading technology. Problems, configurations, effects, etc. other than those described above will be described in the description of the embodiment.

[0012] 1 is a diagram showing the configuration of an image reading-out system according to the first embodiment. 2 is a diagram showing another configuration of the image reading-out system according to the first embodiment. 3 is a diagram showing another configuration of the image reading-out system according to the first embodiment. 4 is a diagram showing the processing flow of the image reading-out system and method according to the first embodiment. 5 is a diagram showing an example of the configuration of an image reading-out device according to the first embodiment. 6 is a diagram showing an example of the configuration of an image reading-out device according to the first embodiment. 7 is a diagram showing an example of the configuration of an image reading-out device according to the first embodiment. 8 is a diagram showing an example of the configuration of an image reading-out device according to the first embodiment. 9 is a diagram showing an example of the configuration of an image reading-out device (device) according to the first embodiment. 10 is a diagram showing an example of the configuration of a server device according to the first embodiment. 11 is a diagram showing an example of the configuration of predetermined timing for image capture, voice reading, etc. according to the first embodiment. 12 is a diagram showing an example of the configuration of multiple cameras according to the first embodiment. 13 is a diagram showing an example of ON / OFF of multiple cameras according to the first embodiment. 14 is a diagram showing an example of control of multiple cameras according to the first embodiment. 15 is a diagram showing an example of an object according to the first embodiment. 16 is a diagram showing an example of the relationship between a user and an object according to the first embodiment. 17 is a diagram for supplementary explanation regarding priority of movement direction, etc. according to the first embodiment. 18 is a diagram showing control modes according to the first embodiment. 1 is a diagram showing an example of a setting information table in the first embodiment. 2 is a diagram showing an example of image recognition and voice reading in the first embodiment. 3 is a diagram showing an example of image recognition and voice reading in the first embodiment. 4 is a diagram showing an example of image recognition and voice reading in the first embodiment. 5 is a diagram showing an example of image recognition and voice reading in the first embodiment. 6 is a diagram showing an example of a processing sequence between a device and a server apparatus in the first embodiment. 7 is a diagram showing basic control related to periodic control in the first embodiment. 8 is a diagram showing control of not performing voice reading when there is no change in the first embodiment. 9 is a diagram showing control of performing voice reading to indicate that there is no change when there is no change in the first embodiment. 10 is a diagram showing an example of control of a re-read instruction in the first embodiment. 11 is a diagram showing an example of voice reading using different expressions in the first embodiment. 12 is a diagram showing an example of control of the front and rear cameras in the first embodiment. 13 is a diagram showing an example of control of using different thresholds for judgment between the front and rear cameras in the first embodiment. 14 is a diagram showing an example of support for the visually impaired in the first embodiment. 15 is a diagram showing an example of support for walking on a road at night in the first embodiment. 16 is a diagram showing an example of support for "walking while using a smartphone" in the first embodiment. Fig. 1 is a diagram showing an example of vehicle assistance in embodiment 1. Fig. 2 is a diagram showing an example of control depending on whether server communication is possible or not in embodiment 1. Fig. 3 is a diagram showing an example of voice reading depending on the priority of an object in embodiment 1.FIG. 1 is a diagram showing examples of risk factors and a safe space guide in the first embodiment. FIG. 2 is a diagram showing an example of voice reading that prioritizes a specific object in the first embodiment. FIG. 3 is a diagram showing an example of an image from a 360-degree camera in the first embodiment. FIG. 4 is a diagram showing an example of control for outputting voice in different modes depending on the camera shooting direction in the first embodiment. FIG. 5 is a diagram showing an example of control for 3D voice output in the first embodiment. FIG. 6 is a diagram showing a system configuration in the second embodiment. FIG. 7 is a diagram showing a device configuration in the second embodiment. FIG. 8 is a diagram showing a server configuration in the second embodiment. FIG. 9 is a diagram showing a processing flow in the second embodiment. FIG. 10 is a diagram showing an example of a system configuration in the second embodiment. FIG. 11 is a diagram showing an example of a system configuration in the second embodiment. FIG. 12 is a diagram showing an example of a system configuration in the second embodiment. FIG. 13 is a diagram showing an example of a system configuration in the second embodiment. FIG. 14 is a diagram showing an example of a system configuration in the second embodiment. FIG. 15 is a diagram showing an example of a system configuration in the second embodiment. Fig. 10 is a diagram showing part 2 of the conversation flow in the first specific example in embodiment 2. Fig. 11 is a diagram showing the conversation flow in the second specific example in embodiment 2. Fig. 12 is a diagram showing an example of a screen display of a device in embodiment 2.

[0013] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the drawings, identical parts are generally designated by the same reference numerals, and repeated explanations will be omitted. In the drawings, the representation of components may not represent their actual positions, sizes, shapes, ranges, etc., in order to facilitate understanding of the invention.

[0014] For the purpose of explanation, when describing processing by a program, the program, functions, processing units, etc. may be described as the main components, but the main hardware components are the processor, or a controller, device, computer, system, etc. that is configured with the processor, etc. The computer executes processing according to the program read into memory using resources such as memory and communication interfaces as appropriate through the processor. This realizes predetermined functions, processing units, etc. The processor is configured, for example, with semiconductor devices such as a CPU / MPU or GPU. Processing is not limited to software program processing, and can also be implemented using dedicated circuits. Dedicated circuits such as FPGAs, ASICs, and CPLDs can be used.

[0015] The program may be pre-installed as data on the target computer, or may be distributed as data from a program source to the target computer. The program source may be a program distribution server on a communication network, or a non-transitory computer-readable storage medium, such as a memory card or disk. The program may be composed of multiple modules. The computer system may be composed of multiple devices. The computer system may be composed of a client-server system, a cloud computing system, an IoT system, etc. Various data and information may be composed of structures such as tables and lists, for example, but are not limited to these. Expressions such as identification information, identifiers, IDs, names, and numbers are interchangeable.

[0016] First Embodiment An image reading system and the like according to a first embodiment will be described with reference to FIG. 1 and subsequent figures.

[0017] The image reading system and method of the first embodiment use a device (also referred to as an image reading device) equipped with a camera and an audio output device as a device carried or worn by a user. This device is equipped with at least a camera (i.e., a photographing function) and an audio output device (i.e., an audio output function), as well as other necessary software and hardware. This device includes a wearable device, and may be, for example, a smartphone, a smartwatch, a tablet terminal, smart glasses, a head-up display (HMD), a shoulder device, or the like. This device may also be a hat, clothing, bag, shoes, or the like equipped with a camera, etc. The device may be equipped on a hat, clothing, etc.

[0018] The image reading system of the first embodiment utilizes multiple cameras mounted on a device that is an image reading apparatus. The multiple cameras may be, for example, cameras in two directions (front and back) or cameras in four directions (front, back, left and right). Each camera may be, in particular, a stereo camera or a sensor with a distance measurement function. The number and orientation of cameras mounted on a device are not limited and various configurations are possible.

[0019] The device captures an image (or takes a picture) from the input video of each camera. This image (also referred to as a captured image) is a photographic image of the surroundings captured in a specific direction, i.e., the shooting direction of the camera (e.g., forward or backward), relative to the user and the device. The image captures various objects, such as pedestrians and cars, as the surroundings.

[0020] The device automatically captures images periodically (periodically, referred to as a first time interval) from the input video of the camera at a predetermined timing. As a result, the device obtains a group of images capturing the user's surroundings continuously / intermittently in time series. The image reading system detects objects in the surroundings from the group of images through analysis, including recognition by AI (e.g., a machine learning model). For example, the image reading system extracts features associated with the objects from the group of images through AI recognition, and detects, as objects, portions of the features that change significantly between images over time, in other words, portions where there is a significant change.

[0021] The image reading system determines whether a detected object is to be read aloud based on a predetermined judgment. The predetermined judgment may be, for example, whether the object is close to the user and device. In addition, if there are multiple objects, the image reading system also determines the priority of the objects to be read aloud, in other words, the output order.

[0022] The image reading system converts the object determined as the target into text representing the object based on the AI ​​recognition result. The AI ​​(e.g., a machine learning model) responds with, for example, information on the position and size of features associated with the object and information on the text representing the object as the recognition result.

[0023] The image reading system device reads out text representing an object (e.g., a pedestrian) in order of priority for voice reading from a voice output device. The text and voice read out may be, for example, "pedestrian" or "there is a pedestrian ahead."

[0024] The image reading system also automatically and periodically (periodically, referred to as a second time interval) reads out the object aloud at a predetermined timing. The period of the read-out aloud (second time interval) may be different from the period of the image capture (first time interval).

[0025] As described above, the device that is an image reading apparatus automatically and intermittently captures images at predetermined timings, repeatedly converts the captured images into text using AI recognition, and automatically and repeatedly reads aloud based on the captured images at predetermined timings. This image capture and aloud reading operation basically does not require any operational input by the user, such as a capture instruction operation.

[0026] The image reading system stores data and information on the history of the continuous image capture and voice reading in a storage resource, for example, for at least a predetermined period of time, so that it can be referenced later.

[0027] [System Configuration (1)] Fig. 1 shows the configuration of an image reading system according to embodiment 1. The image reading system of Fig. 1 is a system including a device 1, which is an image reading apparatus. The system of Fig. 1 includes the device 1 and a server device 2 on a cloud computing system 3, which are appropriately connected by communication via a communication network 9.

[0028] In other words, the device 1 is a user terminal (such as a mobile terminal, a wearable terminal, an information processing terminal, or a computer) that is carried or worn by the user U1. The user U1 may be, for example, a visually impaired person or a non-visually impaired person. Specific examples of the device 1 include a smartphone 1A, a smart watch 1B, smart glasses (or HMD) 1C, a hat 1D, clothing 1E, a shoulder device 1F, and other wearable devices.

[0029] Note that one user U1 may use multiple devices 1. The multiple devices 1 may communicate with each other.

[0030] The device 1 includes a control function 10, a communication function 11, an image capture function 12, an audio output function 13, etc. The control function 10 controls the image reading function. The communication function 11 performs communication processing with the server device 2, which is the AI ​​server 2, via a communication network 9. The image capture function 12 is a function of capturing an image using a camera. The audio output function 13 is a function of outputting audio from an audio output device such as a speaker or an earphone jack. In addition, although not shown, the device 1 may also include a display function, an audio input function, etc. The device 1 includes at least the image capture function 12 and the audio output function 13. Details of the device 1 will be described later using FIG. 7A.

[0031] The server device 2 has the functions of AI 20, or in other words, is an AI server 2. In other words, AI 20 has an analysis function and a recognition function. This AI 20 is software and hardware that recognizes an object (in other words, a feature) from an input image, converts the object (feature) into text representing the object (feature), and outputs information about the object including the text. Details of the server device 2 will be described later using FIG. 7B.

[0032] The device 1 and the server device 2 may function as a client-server system. The device 1 transmits a request (e.g., an image) for recognition by the AI ​​20 to the server device 2 as needed. In response to the request, the server device 2 performs recognition using the AI ​​20 and transmits the recognition result (e.g., text) to the device 1 as a response.

[0033] In this embodiment, the AI ​​20 of the server device 2 performs conversion to text. The conversion from text to speech is performed by the control function 10 of the device 1 (step S5 in FIG. 4 ). As a modified example, the AI ​​20 of the server device 2 may also perform conversion from text to speech after conversion to text, and output the converted speech data to the device 1.

[0034] In addition, in the embodiment described below, it is possible to switch between performing recognition by the AI ​​20 of the server device 2 or within the device 1, depending on the communication status between the device 1 and the server device 2, etc. In this case, an AI function 20b may be provided within the device 1 in Fig. 1 as a corresponding component. This AI function 20b is a copy of the function of the AI ​​20 of the server device 2, or software / hardware that realizes a simple analysis function / recognition function with lower performance than the function of the AI ​​20.

[0035] [System Configuration (2)] Fig. 2 shows a system configuration of a modified example. The system in Fig. 2 is a system in which a device 1a (in other words, an input / output device) having an image capturing function 12 and an audio output function 13, a device 1b having a control function 10 and a communication function 11 as a main body, and a server device 2 which is an AI server 2 are connected. The device 1a may be, for example, smart glasses 1C. The device 1b may be, for example, a smartphone 1A. The combination of the device 1a and the device 1b corresponds to the device 1 in Fig. 1.

[0036] [System Configuration (3)] Figure 3 shows a system configuration of a modified example. The system of Figure 3 is a system in which the functions of the AI ​​20 of the AI ​​server 2 of Figure 1 are integrated into a single device 1 such as a smartphone 1A. In this case, the AI ​​20 (in other words, analysis function and recognition function) is provided within the device 1, and there is basically no need to access the server device 2 (Figure 1) of the communication network 9 when recognizing camera images. The AI ​​20 within the device 1 may be downloaded and installed from a server device of a business operator on the communication network 9, or may be updated or upgraded via communication as appropriate.

[0037] [AI] The AI ​​20 of the AI ​​server 2 in FIG. 1 is, in other words, a large-scale DB (database) for image analysis and text conversion. A multimodal large-scale language model (MLLM) can be applied to this AI 20. The MLLM is an LLM that can process multiple types of data and information, such as images, text, and audio. The MLLM applied to the AI ​​20 in this embodiment has the function of describing and expressing characteristic parts / areas / object shapes and positions in images using words / text / character strings / natural language based on learning of a group of images.

[0038] As MLLM technology that can be applied to AI20, for example, publicly known technologies such as those shown in the following reference information 1 and 2 can be applied.

[0039] [Reference 1] <https: / / arxiv.org / abs / 2306.13549>, "A Survey on Multimodal Large Language Models"

[0040] [Reference 2] <https: / / browse.arxiv.org / pdf / 2306.13549.pdf>, "A Survey on Multimodal Large Language Models"

[0041] [Usage Scenarios] The image reading system according to the first embodiment can be used in a wide range of usage scenarios and applications, including assisting a user when walking. This system can be used in a variety of usage scenarios. More specific usage scenarios include the following:

[0042] 1. Support for the visually impaired: In this application, objects around the user are detected and read aloud to assist the visual acuity of a visually impaired user. FIG. 25A, which will be described later, shows an example of support for the visually impaired. For example, in support for the visually impaired, a visually impaired person with normal hearing may find it difficult or impossible to recognize the surroundings while walking with their own eyes. However, an image reading system converts the surroundings from camera images into text and reads it aloud. This allows the visually impaired person to recognize objects in the surroundings with their hearing, contributing to traffic safety and convenience.

[0043] 2. Nighttime road assistance: In this application, for example, in dark conditions such as roads at night or in unsafe areas, it is possible to detect whether a suspicious person is following the user from behind, and a voice readout is performed to warn the user. An example of nighttime road assistance is shown in Fig. 25B, which will be described later.

[0044] 3. Support for "walking while using a smartphone": In this application, when the user uses the smartphone 1A or the like while walking or standing still after ensuring the safety of the user's surroundings, surrounding objects are detected and audible reading is performed as support. Figure 25C described below shows an example of support for "walking while using a smartphone." Note that this application does not encourage so-called "walking while using a smartphone" or "using a smartphone while doing other things," which involves using a smartphone or the like in a situation where the safety of the user's surroundings is not ensured, but provides support that contributes to the safety of the user and those around them.

[0045] 4. Vehicle driving assistance: In this application, when a user is riding a vehicle such as an electric bicycle or an electric kick scooter, surrounding objects are detected and audible reading is provided as assistance. This application also does not encourage the use of smartphones or other devices while driving, which may lead to accidents, but provides assistance that contributes to the safety of the user and those around them. An example of vehicle driving assistance is shown in Figure 25D, which will be described later.

[0046] Even for users who are not visually impaired, the above-mentioned nighttime road assistance, support for dealing with "walking while using a smartphone," and vehicle driving assistance are effective.

[0047] [Image Reading Method and Processing Flow] FIG. 4 shows a processing flow of the image reading method according to the first embodiment. The system of FIG. 1 executes the processing flow of FIG. 4. In step S1, the device 1 of the user U1 turns on the image reading function. In step S2, the device 1 captures an image from the camera at a predetermined timing. In other words, the device 1 takes an image. In step S3, the system analyzes the captured image obtained in step S2. The analysis includes object recognition by the AI ​​20. As a result of the analysis, the system obtains information including text representing the object (in other words, words, character strings, sentences, descriptions, etc.). In step S4, the system determines the object and text to be read aloud based on the result of step S3 through a predetermined judgment (e.g., a judgment of the degree of change, as described below). In step S5, the system performs a voice reading (in other words, voice output) of the text representing the object determined in step S4 from the device 1 at a predetermined timing. In step S6, device 1 checks whether the image reading function should be turned OFF. If it is turned OFF (YES), the flow ends. If it is to remain ON (NO), the flow returns to step S1 and the same process is repeated.

[0048] [Device Example] Figure 5A shows an example of the implementation of cameras and the like in smart glasses 1C as an example of device 1 in Figure 1. In this example, smart glasses 1C are equipped with multiple cameras with four shooting directions: front, back, left, and right. Note that the coordinate system WU in Figure 2 indicates a coordinate system (X, Y, Z) based on user U1 and device 1. The X-axis / X direction is the left-right direction as seen from user U1, the Y-axis / Y direction is the front-back direction as seen from user U1, and the Z-axis / Z direction is the up-down direction as seen from user U1.

[0049] The smart glasses 1C are equipped with cameras C11 and C12 located on the left and right sides of the front surface, which take images facing forward, and cameras C21 and C22 located on the left and right sides of the rear surface, which take images facing backward. The smart glasses 1C also have cameras C31 and C32 located at the front and rear of the right side, which take images facing right, and cameras C41 and C42 located at the front and rear of the left side, which take images facing left. The cameras C11 and C12 take images in the forward direction (+Y direction), in other words, they are front cameras. The cameras C21 and C22 take images in the backward direction (-Y direction), in other words, they are rear cameras. The cameras C31 and C32 take images in the right direction (+X direction), in other words, they are right cameras. The cameras C41 and C42 take images in the left direction (-X direction), in other words, they are left cameras.

[0050] In this example, for example, the front camera functions as a stereo camera (in other words, a sensor with a distance measurement function) using two cameras C11 and C12, but is not limited to this. A single camera may be used as the front camera. The determination of relative distance, which will be described later, can be made by analyzing images using the AI ​​20, but the distance measurement function of such a stereo camera may also be used.

[0051] A device 1 such as smart glasses 1C may be provided with a remote controller 1c or the like. Alternatively, another smartphone 1A or the like may function as a remote controller for the smart glasses 1C or the like. When a user U1 operates a button or the like provided on the remote controller 1c, a signal such as an instruction is transmitted from the remote controller 1c to the smart glasses 1C. The smart glasses 1C operates in accordance with the signal.

[0052] Although not shown, the smart glasses 1C are also equipped with a display, a microphone, a speaker, a vibrator, a battery, etc. The audio output device is not limited to a speaker, and bone conduction earphones or the like may also be used.

[0053] If the user U1 is not visually impaired, he or she may use the video information displayed on the display surface 1Cd of the smart glasses 1C by looking at it.

[0054] FIG. 5B shows a shoulder-mounted device 1F as another example of the device 1. The shoulder-mounted device 1F is a type of device that is worn around the shoulder or neck. The shoulder-mounted device 1F in FIG. 5B has two rod-shaped housings on the left and right, and a semi-ring-shaped housing on the back side that connects the two housings. This shoulder-mounted device 1F is equipped with multiple cameras that can capture images in four directions: front, back, left, and right. The multiple cameras are the same as the multiple cameras in FIG. 5A.

[0055] The above example is an implementation example in which cameras are provided in four directions (front, rear, left, and right), but this is not limiting; cameras that capture images in at least one direction may be used. The left and right cameras (C31, C32, C41, C42) may be omitted, and only front and rear cameras (C11, C12, C21, C22) may be provided. If cameras (C21, C22) that capture images in the rearward direction are provided, the situation behind the user can be appropriately captured. If cameras (C31, C32, C41, C42) that capture images in the left and right directions are provided, the situation in the left and right directions can be appropriately captured. Note that, for example, if the angle of view of the front camera is sufficiently large (e.g., close to 180 degrees), the situation in the left and right directions can also be captured.

[0056] [Device Usage: Smartphone] FIG. 6A shows an example of a usage mode in which the device 1 is a smartphone 1A. FIG. 6A shows a state in which a user U1 holds the smartphone 1 in his / her hand and captures images of the user U1 in the front-to-back directions (±Y directions) using cameras on the front and back (front and rear) of the smartphone 1A. In FIG. 6A, the user U1 holds the smartphone 1A in his / her hand so that the housing is nearly vertical. Camera C1 is a front camera (outer camera) mounted on the back side (the side without a display screen) of the flat housing of the smartphone 1A, with its optical axis for capturing images in the forward direction (+Y). Camera C2 is a rear camera (inner camera) mounted on the front side (the side with a display screen) of the flat housing of the smartphone 1A, with its optical axis for capturing images in the rearward direction (-Y).

[0057] In addition, if the user U1 is not visually impaired, the user U1 may use the smartphone 1A by viewing the video information on the display screen 1Ad of the smartphone 1A as needed.

[0058] The user U1 may hold the smartphone 1A in a direction other than the forward direction shown in the figure, in which case the front camera C1 can capture an image in that direction. If the smartphone 1A is equipped with a camera that captures an image in another direction, that camera can also be used.

[0059] [Device Usage: Hat and Clothing] FIG. 6B shows another example of a usage of device 1, in which cameras are mounted on hat 1D and clothing 1E to capture images of the user U1 in the front and back directions. For example, hat 1D is equipped with a camera C3 on the front side that captures images in the forward direction (+Y) and a camera C4 on the back side that captures images in the backward direction (-Y). Alternatively, clothing 1E is equipped with a camera C5 on the front side that captures images in the forward direction (+Y) and a rear camera C6 on the back side that captures images in the backward direction (-Y). These cameras may be connected via communication with device 1 (e.g., smartphone 1A), which serves as the main body. Similarly, left and right cameras may be mounted.

[0060] The device 1 may be attached to the body, clothes, bag, shoes, etc. of the user U1. For example, a camera or a device 1 with a camera may be attached to a pocket / holder of clothes 1E. The user U1 may wear / carry multiple devices 1 such as smartphones 1A. Even in the case shown in FIG. 6B , the device 1 can capture the front and rear directions of the user U1 using front and rear cameras.

[0061] [Device Configuration Example] Fig. 7A shows a configuration example of device 1. Device 1 includes a processor 301, a memory 302, a non-volatile memory (i.e., storage) 303, a camera (i.e., a photographing unit) 304, an image processing unit 305, an image storage unit 306, a position detection unit 311, an orientation detection unit 312, a biometric information detection unit 313, a button / operation unit 314, a display / display unit 315, a microphone / audio input unit 316, a speaker / audio output unit 317, a connector / input / output interface 318, a battery / power supply unit 319, a vibration unit 320, and a communication unit 330. The communication unit 330 is equipped with a communication interface, and includes, for example, a first wireless communication unit 331, a second wireless communication unit 332, a third wireless communication unit 333, and a wired communication unit 334. Each wireless communication unit is equipped with a corresponding wireless communication interface, for example, a mobile network, a wireless LAN, Bluetooth (registered trademark), or the like. The camera 304 includes a plurality of cameras, for example, cameras #1 to #8, which correspond to the eight cameras on the front, back, left and right sides in FIG.

[0062] The processor 301 is composed of a CPU or the like, and controls the entire device 1 and each unit thereof. The memory 302 stores data and information to be processed by the processor 301. The processor 301 reads a program from the non-volatile memory 303 into the memory 302 and executes processing in accordance with the program. This allows various functions to be realized as execution modules.

[0063] The non-volatile memory 303 stores various data and information, such as programs such as an OS 341, image reading software 342, AI data 343, user information 344, setting information 345, and access destination information 346. The image reading software 342 is data such as a program for realizing the image reading function in embodiment 1. The AI ​​data 343 is data such as recognition results exchanged with the AI ​​20 of the server device 2. Note that when recognition processing is performed on the device 1 side, the AI ​​data 343 includes data corresponding to the AI ​​20. The user information 344 is information about the user U1. The setting information 345 is system setting information and user setting information. The access destination information 346 is information for communicating with the server device 2, etc.

[0064] The camera 304 includes a circuit for controlling image capture by the multiple cameras, etc. The image processing unit 305 processes images captured by the camera 304. The image storage unit 306 stores data such as images captured by the camera 304 and images processed by the image processing unit 305.

[0065] The position detection unit 311 detects the position of the device 1 using, for example, GNSS or the like and obtains position information. The orientation detection unit 312 includes an orientation sensor and detects the orientation of the device 1. The biometric information detection unit 313 detects, for example, the fingerprint of the user U1 to perform fingerprint authentication, or detects the face of the user U1 to perform face authentication. The button / operation unit 314 includes a power button, volume button, etc. and accepts operation input by the user U1. The display / display unit 315 performs processing to display video / images on the display. The microphone / audio input unit 316 inputs and recognizes audio through a microphone. The speaker / audio output unit 317 outputs audio through a speaker, earphone jack, etc. The connector / input / output interface 318 is equipped with an input / output interface for connecting input / output devices. The battery / power supply unit 319 supplies the necessary power to each unit based on the battery. The vibration unit 320 generates vibrations as needed.

[0066] 7B shows an example of the configuration of the server 2. The server 2 has a processor 401, a memory 402, a non-volatile memory (in other words, storage) 403, an operation input unit 414, a connector / input / output interface 418, a battery / power supply unit 419, and a communication unit 430. The communication unit 430 is equipped with a communication interface, and includes, for example, a first wireless communication unit 431, a second wireless communication unit 432, a third wireless communication unit 433, and a wired communication unit 434.

[0067] The processor 401 is composed of a CPU / GPU, etc., and controls the entire server device 2 and each unit. The memory 402 stores data and information to be processed by the processor 401. The processor 401 reads a program from the non-volatile memory 403 into the memory 402 and executes processing in accordance with the program. This allows various functions to be realized as execution modules.

[0068] The non-volatile memory 403 stores various data and information, such as programs such as an OS 441, AI software 442, AI data 443, user information 444, setting information 445, and access destination information 446. The AI ​​software 442 is data such as a program for implementing the AI ​​20 in FIG. 1 that performs analysis processing including image recognition processing. The AI ​​data 443 is data such as a model and image data for the recognition processing of the AI ​​20, and the recognition results exchanged between the AI ​​20 and the device 1. The user information 444 is information about the user U1. The AI ​​setting information 445 is setting information for the AI ​​20. The access destination information 446 is information for communicating with the device 1, etc.

[0069] The operation input unit 414 accepts operation input. The connector / input / output interface 418 is equipped with an input / output interface for connecting an input / output device. The battery / power supply unit 419 supplies the necessary power to each unit based on a battery.

[0070] [Specified timing for image capture, etc.] Figure 8 shows an example of the relationship between specified timing, periodicity, etc. between input image 801 (time point or image frame) from camera 304 of device 1, image capture 802, recognition by AI 20 803, and voice reading 804.

[0071] The input video 801 has image frames at each time point in the time series. The image frames are captured at a predetermined frame cycle of the camera 304. For example, frame f1 is captured at time point 1. Image capture 802 is automatically performed continuously / intermittently at a preset cycle (first time interval). Images are captured from the image frames of the input video 801 at that cycle. For example, frame 1 at time point 1 is captured as image 1, and frame 3 at time point 3 is captured as image 2. In this example, image capture 802 is performed at a cycle of two frames. Note that in practice, capture may be performed once every eight frames, for example. The captured images are stored in memory (image storage unit 306). The images may be captured as still images by each camera, or as a frame of a video being captured.

[0072] Recognition 803 by AI20 is performed for each captured image. In this example, recognition 803 by AI20 is performed at the same cycle as image capture 802. Voice reading 804 is automatically performed at a preset cycle (second time interval) for objects detected from the results of recognition 803 and determined as targets for voice reading. For example, voice reading 1 is performed for an object detected from image 1 at time point 1, and voice reading 2 is performed for an object detected from image 3 at time point 3. In this example, voice reading 804 is performed at a cycle of four frames. Note that if no object is detected as a result of recognition 803 of the captured image, or if the detected object is not a target for voice reading, voice reading 803 will not be performed even at the timing corresponding to the cycle.

[0073] The device 1 stores a history of the processes such as the image capture 802, recognition 803, and voice reading 804 (including data such as captured images and recognition results) for at least a predetermined period of time in a storage resource (e.g., the non-volatile memory 403). The device 1 may transmit the history data to the server device 2 or the like for storage.

[0074] If the device 1 determines that the image captured this time has a different movement or a significant change in the object compared to the image captured last time, the device 1 controls the device 1 so that the object is given priority in the voice reading. A specific example will be described later.

[0075] [Number of Cameras and Shooting Direction] FIG. 9A is a schematic diagram viewed vertically from above, showing an example of the number and shooting direction of the cameras 304 of the device 1 of user U1. State A is a first example. User U1 is walking, for example, on a sidewalk 901. User U1 carries or wears device 1 (for example, smartphone 1A). The first example is an example in which two cameras (C1, C2) on the front and back of smartphone 1A, as shown in FIG. 6A, are used as front and rear cameras. For ease of understanding, the front and rear cameras (C1, C2) are illustrated separately in the drawing. DC1 is the shooting direction of the front camera (C1), which is the +Y direction corresponding to the front direction of user U1. DC2 is the shooting direction of the rear camera (C2), which is the -Y direction corresponding to the rear direction of user U1. In this example, the front camera (C1) and rear camera (C2) are equipped with wide-angle lenses, and the shooting range / angle of view is close to 180 degrees. The semicircle represents an overview of the shooting range / angle of view. The same applies when applying FIG. 6B.

[0076] State B is a second example. The second example is an example in which four cameras (front, back, left, and right) are used on the smart glasses 1C as shown in FIG. 5A. A user U1 is wearing a device 1 (e.g., smart glasses 1C). Here, the front camera is C10, the rear camera is C20, the right camera is C30, and the left camera is C40. DC10 is the shooting direction (+Y) of the front camera C10. DC20 is the shooting direction (-Y) of the rear camera C20. DC30 is the shooting direction (+X) of the right camera C30. DC40 is the shooting direction (-X) of the left camera C40. The dashed triangles are an overview of the shooting range / angle of view of each camera, and have a field of view of at least 45 degrees or more, and may be close to 180 degrees. The same applies when FIG. 5B is applied.

[0077] The camera 304 is not limited to being integrated into the device 1, and may be configured such that an external camera device is connected to the device 1 (e.g., smartphone 1A) that serves as the main body. The camera device may also be attached to a predetermined position on the body (e.g., hat 1D or clothing 1E).

[0078] [Camera On / Off] In another embodiment, only some of the multiple cameras of the device 1 may be set to be used in an on state. FIG. 9B shows an example of turning multiple cameras on / off. State A is an example in which only the front camera C10 of the four cameras in the front, back, left, and right directions of the smart glasses 1C is turned on. State B is an example in which the front camera C10 and the rear camera C20 are turned on. State C is an example in which the front camera C10, the right camera C30, and the left camera C40 are turned on. State D is an example in which the right camera C30 and the left camera C40 are turned on.

[0079] [Example of Controlling Multiple Cameras] Furthermore, when using multiple cameras, such as front and rear cameras, on device 1, it is also possible, as an example, to simultaneously and in parallel perform the processes from image capture to voice reading (e.g., FIG. 8 ) for each camera. However, this can result in a high processing load and potential delays. Therefore, as an example, the processing of images from each of the multiple cameras may be controlled to have a predetermined timing or frequency, in other words, to be time-separated. For example, the overall schedule may be controlled to separate the predetermined timing and frequency for the capture, recognition, and voice reading processes from the front camera (C1) and the capture, recognition, and voice reading processes from the rear camera (C2). For example, the following control example can be given.

[0080] First, processing parts that require a low load even in parallel processing, such as image capture and image storage, may be performed repeatedly in the same manner for both the front and rear cameras. If the parallel processing load for image capture and image storage is high, the timing may be divided and the processing may be performed sequentially. Then, when the device 1 reads the captured images from each camera from memory and performs subsequent processing, the timing is divided for each image from each camera. For example, the images from the front and rear cameras may be read and processed alternately.

[0081] Control 1 in Fig. 9A shows an example of control of the front and rear cameras corresponding to state A. For example, at time points 1 and 3, the image from the front camera (C1) is processed (recognition and voice reading), and at the next time points 2 and 4, the image from the rear camera (C2) is processed (recognition and voice reading), and this is repeated alternately. In this example, the cycle (second time interval) of voice reading in processing the image from the front camera (C1) and the cycle (second time interval) of voice reading in processing the image from the rear camera (C2) alternate, and the processing frequency ratio is 1:1.

[0082] Control 2 in Figure 9A shows an example of control of the front, rear, left, and right cameras corresponding to state B. For example, at time point 1, an image from the front camera (C10) is processed, at the next time point 2, an image from the rear camera (C20) is processed, at the next time point 3, an image from the right camera (C30) is processed, and at the next time point 4, an image from the left camera (C40) is processed, and so on. These four directions of processing are treated as one set and are repeated in the same manner thereafter. In this example, the voice read-out period (second time interval) is the same for processing images from each direction, but the timing is shifted by one, resulting in a processing frequency ratio of 1:1:1:1.

[0083] As another example of control, it is possible to set the system to a state where only images from the front camera (C1) are repeatedly processed under normal circumstances, and when a predetermined condition is met, change to a state where only images from the rear camera (C2) are repeatedly processed. Alternatively, it is possible to set the system to a state where only images from the front camera (C1) are repeatedly processed under normal circumstances, and when a predetermined condition is met, change to a state where images from the front camera (C1) and images from the rear camera (C2) are processed alternately.

[0084] FIG. 9C shows an example of control for changing the processing ratio of each camera in the case of a device 1 equipped with at least front and rear cameras (C1, C2) similar to state A in FIG. 9A . Control A repeats processing of only images from the front camera (C1). Control B mainly processes images from the front camera (C1) while also inserting processing of images from the rear camera (C2) at a frequency of 3:1. Control C alternates between processing images from the front camera (C1) and processing images from the rear camera (C2) at a frequency of 1:1 (similar to control 1 in FIG. 9A ). Control D mainly processes images from the rear camera (C2) while also inserting processing of images from the front camera (C1) at a frequency of 3:1. Control E repeats processing of only images from the rear camera (C2)

[0085] The image reading system may be configured to switch between such multiple controls at appropriate times based on a user's operation input or a predetermined decision.

[0086] From the perspective of computational power, it may be difficult to achieve real-time and highly accurate voice reading for all surrounding situations in front, behind, left, and right of the user. Furthermore, even if all surrounding situations in each direction are read out simultaneously, the user may become confused. In such cases, the camera to be used, processing ratio, control mode, etc., as described above, can be set / selected depending on which direction recognition and reading is prioritized for the user and usage scenario.

[0087] For example, if rear support is not required depending on the user, control A may be used. If only rear support is required depending on the user, control E may be used. Also, it is possible to apply control C under normal circumstances to recognize the front and rear in a balanced manner, switch to control B when focusing on an object in front, and switch to control D when focusing on an object in the rear (see FIG. 23 described below).

[0088] [Surrounding Situation of User] FIG. 10 is an explanatory diagram of the surrounding situation when a user U1 wearing a device 1 (e.g., smart glasses 1C) is standing still or walking, for example, on a sidewalk 901, and is a schematic diagram viewed vertically from above. The surrounding situation has front, back, left, and right directions as seen from the user U1. In this example, it is assumed that the user U1 is walking forward (+Y direction) on the sidewalk 901. The white arrow indicates the direction of walking / movement. In this example, a bicycle lane 902 and a roadway 903 are located on the right side of the sidewalk 901.

[0089] In the illustrated example, examples of objects to be detected include a pedestrian 1001, a bicycle 1002, and a car (automobile) 1003, all of whom are other people. The pedestrian 1001 is walking on the sidewalk 901 in the -Y direction shown in the figure, and is approaching the front of the user U1. From the perspective of the user U1, the pedestrian 1001 is in front. The bicycle 1002 is traveling in the -Y direction in the bicycle lane 902, and is approaching the user U1 diagonally from the front. From the perspective of the user U1, the bicycle 1002 is diagonally in front and to the right. The car 1003 is traveling on the roadway 903 in the +Y direction. From the perspective of the user U1, the car 1003 is on the right side. Such objects are the targets of voice reading.

[0090] [Positional Relationship Between User and Object] Figure 11 shows an example of the positional relationship between the user U1 and an object. An explanation will be given using an example of a pedestrian 1001 (Figure 10) facing in front of the user U1. State A is a case where the user U1 is stationary and the pedestrian 1001 is stationary. State B is a case where the user U1 is stationary and the pedestrian 1001 is walking. State C is a case where the user U1 is walking and the pedestrian 1001 is stationary. State D is a case where the user U1 is walking and the pedestrian 1001 is walking.

[0091] For each state, the relative positional relationship, speed, and other relationships between user U1 and pedestrian 1001 can be considered. The dashed circle represents a predetermined distance range 1100 centered on the positions of user U1 and device 1. A radius 1101 of the range 1100 corresponds to the threshold value of the relative distance between user U1 and the object. In the example of FIG. 11 , the threshold value of the relative distance is the same in each direction. In each of states A to D, there is a relative distance DA between user U1 and pedestrian 1001. In state A, the relative speed between user U1 and pedestrian 1001 is 0. In state B, when the speed of pedestrian 1001 is v2 (-v2), the relative speed is v2 (-v2). In state C, when the speed of user U1 is v1 (+v1), the relative speed is v1 (+v1). In state D, if the speed of the pedestrian 1001 is v2 (-v2) and the speed of the user U1 is v1 (+v1), the relative speed is +v1-(-v2)=v1+v2.

[0092] In this embodiment, the above-described relative distance and relative speed can be determined basically based on the analysis of the camera image. However, without being limited to this, the position, orientation, speed, etc. of the user U1, the position, orientation, speed, etc. of the object, and the relative distance and speed, etc. between the user U1 and the object may be determined using various sensor functions provided in the device 1.

[0093] Furthermore, the above-described relative distance and speed can be used for predetermined control. For example, when the relative distance between the user U1 and the object is within a threshold, the object may be subject to voice reading or an alert. For example, when the relative distance DA is within a predetermined distance range 1100, the object may be subject to voice reading. Furthermore, the predetermined distance range 1100 (corresponding threshold) may be set in multiple stages. For example, when the relative distance DA is within a first range, the object may be subject to voice reading, and when it is within a smaller second range, the object may be subject to voice reading with an additional alert. Similarly, when the relative speed between the user U1 and the object is equal to or greater than a threshold, the object may be subject to voice reading or an alert.

[0094] In addition, in this embodiment, the relative relationship between the user and the object is grasped by handling the amount of change in the object between captured images (described later). This makes it possible to similarly process various situations depending on the combination of the user being stationary or moving and the object being stationary or moving, as shown in FIG.

[0095] 12 is a supplementary diagram illustrating the relationship between the movement direction of the user U1 and the shooting direction of the camera 403. The movement direction of the user U1 may or may not match the shooting direction of the camera.

[0096] State A is a case where the device 1 of user U1 has front and rear cameras. User U1 is moving forward (+Y direction). User U1's head is facing forward (+Y). The direction of user U1's movement, the direction of his head, and the shooting direction (DC1) of the front camera (e.g., C1) are the same. In state A, if recognition is set to prioritize the direction of user U1's movement, control can be performed to prioritize the capture and recognition of images from the front camera (C1). Note that, for example, if it is desired to recognize an object to the left (-X) of user U1, images from front and rear cameras with a sufficiently large angle of view may be used (combined). Alternatively, images from the left and right cameras of the smart glasses 1C may be used.

[0097] As another example, in state B, user U1 is moving to the left (-X direction). User U1's head is facing forward (+Y). The direction of user U1's movement and the shooting direction of the front camera (C1) do not match. The direction of user U1's head and the shooting direction of the front camera (C1) match. In state B, when setting to prioritize recognition in the direction of user U1's movement (for example, leftward), images from front and rear cameras with a sufficiently large angle of view may be used (or used together). Alternatively, images from the left and right cameras of the smart glasses 1C may be used. Furthermore, when it is desired to prioritize recognition in state B in terms of the head direction, images from the front camera (C1) may be used.

[0098] As another example, state C is a case where the smart glasses 1C of user U1 have cameras on the front, back, left, and right. User U1 is moving, for example, in the forward direction (+Y), but his head is facing, for example, diagonally forward to the left. The direction of user U1's head (diagonally forward left) and the shooting direction of the front camera (C10) are consistent. The direction of user U1's movement (+Y) and the shooting direction of the front camera (C10) (diagonally forward left) do not match. In addition, the shooting direction of the right camera (C30) is diagonally forward to the right. The direction of user U1's movement (+Y) and the shooting direction of the right camera (C30) (diagonally forward right) also do not match. There is no camera that directly captures the forward direction (+Y), which is the direction of movement, but there are cameras with adjacent shooting directions as close as possible to the forward direction, namely the front camera (C10) and the right camera (C30). In state C, when the setting is to prioritize recognition in the direction of movement of user U1 (for example, forward), it is sufficient to use both the image from the front camera (c10) and the image from the right camera (C30), which have a sufficiently large angle of view. For example, it is also possible to use an image that combines the right part of the image from the front camera (C10) with the left part of the image from the right camera (C30). Also, in state C, when it is desired to prioritize recognition in the direction of the head, it is sufficient to use the image from the front camera (C10).

[0099] As another example, state D is a case where the front and rear cameras C5, C6 are located on the torso of user U1, for example, on clothing 1E (FIG. 6B). User U1 is moving, for example, in the forward direction (+Y). User U1's head is facing, for example, diagonally forward to the left, while the torso and clothing 1E are facing forward (+Y). The direction of user U1's movement (+Y) and the shooting direction of the front camera (C5) are the same. The direction of user U1's head (diagonally forward to the left) and the shooting direction of the front camera (C5) are not the same. The same applies when the smartphone 1C held by user U1 is facing forward. In state D, if the setting is such that recognition is prioritized in the direction of user U1's movement (for example, forward), the image from the front camera (C5) can be used.

[0100] As in the example of state C above, when it is desired to recognize a direction different from the camera's shooting direction (for example, the direction of movement), this can be achieved by using an image from a camera with a sufficiently large angle of view or by combining images from multiple cameras with different shooting directions. Even in the case of the setting that prioritizes the direction of movement, which will be described later, this can be achieved by using images from each camera, as in the example above.

[0101] The device 1 may use various sensors to detect the movement direction of the user U1 in three-dimensional space, the direction of the head, the direction of the trunk, and the direction of each camera. For example, the movement direction of the user U1 can be detected by a sensor of the posture detection unit 312 ( FIG. 7A ). The device 1 may use each detected direction for control. Furthermore, if the device 1 (e.g., smart glasses 1C) has a gaze detection function, the gaze direction detected by the gaze detection function may be used for control.

[0102] [Control Mode] FIG. 13 is a table summarizing several control modes as an explanatory diagram of the control modes. The image reading system having the device 1 and server device 2 of FIG. 1 may implement only one specific mode as a control mode, or may implement multiple modes and be able to switch between modes as needed. For example, table 1301 shows control modes related to camera shooting direction. As an example of a control mode, the first mode (M1) is a mode in which shooting and recognition are performed only in the forward direction based on the user U1 and device 1. The second mode (M2) is a mode in which shooting and recognition are performed in the forward and backward directions. The third mode (M3) is a mode in which shooting and recognition are performed in the forward, backward, left, and right directions. For example, for the four cameras (front, back, left, and right) of the smart glasses 1C of FIG. 5A, which camera to use is set and controlled depending on the control mode.

[0103] Table 1302 also shows control modes related to the use and cooperation of the server device 2. As examples of the control modes, the first mode (MA) is a mode in which the device 1 performs analysis and recognition processing on its own without communicating with the server device 2. The second mode (MB) is a mode in which communication is performed with the server device 2, and analysis and recognition processing is performed on the server device 2. The third mode (MC) is a mode in which communication is performed with the server device 2, and analysis and recognition processing is performed on the server device 2 when communication with the server device 2 is possible (or good), and analysis and recognition processing (for example, simple analysis and recognition, which will be described later) is performed on the device 1 itself when communication with the server device 2 is impossible (or has deteriorated).

[0104] [Setting Information] Fig. 14 shows an example of a table of setting information for this system. This setting can be made for each user. It is also possible for each user to use multiple settings. User settings for the image reading function can be made on the screen of a graphical user interface (GUI) provided by the device 1 or server device 2 of this system.

[0105] This table has the following setting items: #1 "Camera used," #2 "Reading camera when multiple cameras are used," #3 "Camera image capture interval," #4 "Reading interval," #5 "Determination criteria for reading from the main camera," #6 "Determination criteria for reading from the sub-camera," #7 "Reading target priority," #8 "Unchanged reading function," #9 "Reading function with different expressions," #10 "Safe space guide function," #11 "History retention period," #12 "Other detailed settings," and the like.

[0106] In the #1 "Camera in Use" item, the camera 403 provided in the device 1 can be set to be used for shooting by turning it on or off. For example, the front, back, left, and right cameras of the smart glasses 1C can be set to be on or off. Alternatively, when using a 360-degree camera, which will be described later, it can be set to be on or off. On is the enabled state, and off is the disabled state.

[0107] In the item #2 "Reading camera when using multiple cameras," when using multiple cameras for voice reading, it is possible to set which camera's image to use for voice reading. A voice reading priority / priority order may be set for each of the multiple cameras.

[0108] The setting value "Even" is a control that evenly executes multiple voice readings corresponding to the multiple cameras used. An example of equal control is to set an equal processing ratio for the front, rear, left, and right cameras as described above (FIG. 9A), and perform voice readings sequentially in a predetermined order.

[0109] The setting value "Forward direction basic (main)" basically reads aloud from the image from the front camera. Reads aloud from other directions are performed by interrupting only when necessary (for example, when certain conditions are met). For example, in control B of Figure 9C, the front camera is the main camera and the rear camera is the sub-camera. The setting value "Rear direction basic (main)" basically reads aloud from the image from the rear camera. Reads aloud from other directions are performed by interrupting only when necessary (for example, when certain conditions are met). For example, in control D of Figure 9C, the rear camera is the main camera and the front camera is the sub-camera. A specific example is shown in Figure 23, which will be described later.

[0110] The setting value "Movement direction base" controls the execution of voice reading based on the image from the camera associated with the movement direction of user U1 (see, for example, FIG. 12 described above). Voice reading for other directions is performed only when necessary (for example, when a predetermined condition is met) by interrupting, etc. Explaining the example of FIG. 12, in state A, the movement direction is forward (+Y), and voice reading based on the image from the front camera in the same direction is performed. In state B, the movement direction is left (-X), and if there is a left camera that matches the movement direction, voice reading based on the image from the left camera is performed. If there is no camera that matches the movement direction, processing can be performed using the image from another camera that is as close as possible to the movement direction. In state C, the movement direction is forward (+Y), and there is no single camera that matches the movement direction, but processing can be performed using the image from the front camera or the right camera that is as close as possible to the movement direction.

[0111] The #3 "Camera Image Capture Interval" item allows the user to set a first time interval (cycle, frequency, etc.) for capturing images to be used for voice reading from a time-series group of camera images. It may also be possible to set a time interval for the camera used primarily (i.e., the main camera) and a time interval for other cameras used (i.e., the sub-camera). This capture interval may be set numerically, or may be set by selecting from several levels, for example.

[0112] The #4 "Reading Interval" item allows the setting of a second time interval (cycle, frequency, etc.) for performing a voice readout based on the captured image. This reading interval may be set numerically or may be set by selecting from several levels. For example, it may be set to high (every second), medium (every minute), or low (every few minutes). In addition, it may be possible to set a time interval for the camera used primarily (i.e., the main camera) and a time interval for other cameras used (i.e., the sub-camera). The processing load of the system can be adjusted according to the setting values ​​of #3 and #34.

[0113] The item #5 "Determination conditions for reading from the main camera" allows the determination conditions for reading from the main camera according to the settings of the reading camera in #2 to be set. Setting value A. "Objects in the direction of movement" is a control that performs reading from the voice of objects detected in the direction of movement / travel of user U1 as much as possible. Note that as another control, a control that performs reading from the voice of all recognized and detected objects as much as possible may also be applied.

[0114] Setting value B. "Approaching object" is a control that prioritizes an object approaching the user U1 as the target for voice reading (described later). It is also possible to set the relative distance and speed thresholds (FIG. 11) used to determine an approaching object in this control.

[0115] Setting value C. "Notable change" is a control that detects characteristic parts that have a significant change between images and prioritizes objects corresponding to the significant change parts as targets for voice reading (described below). A significant change may be, for example, a change or difference between images regarding characteristics or feature points within the images that is greater than or equal to a threshold value.

[0116] Setting value D. "Specific object" is a control that detects a specific object and sets it as the target for voice reading. The specific object can be set by the user (e.g., selected from candidates) on another transition screen. Examples of specific objects include crosswalks, traffic lights, and traffic signs, which will be described later.

[0117] Setting value E. "Danger detection" is a control that detects dangerous conditions / objects in the surroundings and makes them the subject of voice reading (described later).

[0118] The item #6 "Sub-camera reading interrupt timing" allows the user to set the timing (second time interval) of the voice reading interrupt for the sub-camera according to the setting of the reading camera in #2. Setting value A. "Every specified time" allows the user to set the interrupt cycle for the sub-camera. This setting may also be a setting of the ratio to the main camera. For example, in the example of control B in FIG. 9C, the main camera (front camera) reads three times, while the sub-camera (rear camera) reads once, resulting in a setting of 3:1, or 1 / 4.

[0119] Setting value B. "When an approaching object is detected", a voice readout is interrupted when an object approaching user U1 is detected. Setting value C. "When a significant change is detected", a voice readout is interrupted when an object with a significant change between images is detected. Setting value D. "When a specific object is detected", a voice readout is interrupted when a specific object is detected. Setting value E. "When a danger is detected", a voice readout is interrupted when a dangerous condition / object is detected in the surrounding situation.

[0120] The #7 "Reading Object Priority" item allows the user to set priorities for objects and conditions for voice reading. This priority is particularly the priority for voice reading for various objects that can be recognized and detected from an image. For example, there are objects and conditions similar to the setting values ​​A to E in #5 described above, and these may be prioritized or rearranged. Alternatively, the priority for each object or condition may be set by selecting from several levels, such as high, medium, or low. Furthermore, a default setting for the system may be provided, and for example, "danger" may be given the highest priority, followed by "approaching objects."

[0121] #8 "No change reading function" allows you to set whether or not to read out aloud the state of no change when it is determined that there is not much change in the object between images (described later).

[0122] #9 "Reading function using different expressions" allows you to set whether or not to read aloud the text representing the object using different expressions when repeatedly reading aloud the same detected object (described below).

[0123] #10 "Safe Space Guide Function" allows you to set whether to perform a voice readout to guide the user to move to a safe space so that the user does not come into contact with the object, for example, when the object approaches the user, based on the positional relationship between the user and the object (described below).

[0124] #11 "History retention period" allows the user to set the period for retaining the voice reading history data. The history retention period may be set separately for the device 1 and the server device 2. The options for the setting value may be, for example, forever, one year, three months, no retention, etc.

[0125] #12 "Other detailed settings" are detailed settings other than #1 to #11 above. For example, image resolution can be set. For example, image resolution can be set to high (600 dpi), medium (350 dpi), low (72 dpi), etc.

[0126] [Recorder Function] One of the functions of the system is a recorder function related to image reading. Depending on the user settings (#11 in FIG. 14), historical data of the image reading function's processing (including image capture 802, recognition 803, and voice reading 804 in FIG. 8) can be recorded and saved in a storage resource (device 1 or server device 2). This recorder function (in other words, a walking recorder) functions like a drive recorder in an automobile. That is, this function records the surrounding circumstances of the user while walking as a history and allows the history to be reviewed later. The saved historical data can be output later (screen display or audio output). This function allows the surrounding circumstances at the time to be reviewed as evidence. The historical data may include date and time, captured images, text of the recognition results, etc.

[0127] [Recognizing Objects in an Image and Reading Them Aloud] FIG. 15A is a schematic explanatory diagram showing an example of object recognition and reading aloud in an image captured by the camera 403 of the device 1. FIG. 15A schematically shows an example of recognition by the AI ​​20 from a captured image and an example of changes in the object in the image. In particular, this example shows an example of an object approaching the user U1 (i.e., the device 1). The upper part shows an image 1501 at a first time point, and the lower part shows an image 1502 at a second time point. Assume that the object OB1 in the image is a pedestrian, a stranger, walking toward the user. The AI ​​20 recognizes such objects / situations / changes as features in the image and converts them into text representing the object / situation / change. The AI ​​20 expresses the result of its inference of the identity of the object OB1 in text.

[0128] First, when a pedestrian (pedestrian 1001 in FIG. 10 ) is recognized and detected as object OB1 from image 1501, the text representing object OB1 may be converted into, for example, "pedestrian" (or "person"), and the "pedestrian" may be read aloud. Within image 1501, region r1 is a rectangular image region that includes object OB1 (OB1a). Point p1 is an example of position coordinates representing object OB1 / region r1. The AI ​​20 of the server device 2 outputs, as a response, information such as the position coordinates, size, and shape of the region of object OB1, and text representing object OB1.

[0129] Furthermore, if it is determined between images that the object OB1 is approaching the user, text describing the situation in which the object OB1 is approaching may be obtained and read aloud, for example, "A pedestrian is approaching."

[0130] In image 1502, region r2 is a rectangular image region that includes object OB1 (OB1b). Point p2 is an example of position coordinates that represent object OB1 / region r1. For example, in image 1502, the position (point p2) of region r2, which is a pedestrian that is object OB1, has moved and changed in size more significantly compared to the state of the position (point p1) of region r1 in image 1501. In the example of FIG. 15A , the amount of change in size is large.

[0131] The device 1 or the server device 2 may determine whether the object OB1 is approaching the user based on the amount of change in size of the region (r1, r2) of the object OB1 between images. Alternatively, the device 1 or the server device 2 may determine the relative distance and speed of the object from the user based on the images, and determine that the object is approaching when the relative distance falls within a predetermined range (threshold), as shown in FIG.

[0132] Furthermore, the determination of approach is not limited to a binary determination of whether or not there is approach, but may be a multi-value determination of the degree of approach.

[0133] In the example of Figure 15A, the word "pedestrian" or the situation of a pedestrian approaching may be read out loud at the time of image 1501, or the word "pedestrian" or the situation of a pedestrian approaching may be read out loud at the time of image 1502.

[0134] The present system performs control so that reading out is performed with priority given to objects that are closer to the user. In other words, the present system performs control so that reading out is performed with priority given to objects that are closer to the user.

[0135] FIG. 15B shows another example. In image 1511, the detected object OB2 (OB2a) is a pedestrian approaching from the left side of the intersection and moving to the right. In image 1512, the detected object OB2 (OB2b) is a pedestrian approaching the center of the intersection. Compared to image 1511, in image 1512, the size of region r3 of object OB2 remains almost the same as region r4, but the position coordinates have changed significantly from p3 to p4. The device 1 or the server device 2 may determine the movement status of object OB2 based on the amount of change in the position coordinates of region (r3, r4) of object OB2 between images.

[0136] From the change in the size of the feature increasing as in Fig. 15A, it can be understood that the object is approaching the user in the forward or backward direction, and from the change in the position coordinates of the feature as in Fig. 15B, it can be understood that the object is moving, for example, left or right relative to the user and crossing in front of the user.

[0137] Using the amount of change in the size, position, etc. of features corresponding to the object between the images as described above, the device 1 or the server device 2 can grasp the positional relationship between the user and the object. Based on this understanding, predetermined control is possible. Examples of predetermined control include determining an object as a target for voice reading if the amount of change in the object is large, or prioritizing an object as a target for voice reading if the object is close to the user.

[0138] When controlling the device 1 to give priority to reading out an object approaching the user's position, the device 1 may determine the priority of reading out depending on the relative distance from the user, in other words, depending on the distance range as shown in Fig. 11. For example, a plurality of distance thresholds may be set, such as 1 m, 3 m, 5 m, and 10 m. For example, if there are both an object that is 10 m and an object that is 5 m around the user at the same time, the closer object is given a higher priority for reading out and is read out first.

[0139] Alternatively, the system may control the volume of the voice reading to be louder the closer the object is to the user (or the closer the relative distance is to the object), or the louder the volume may be the faster the relative speed (FIG. 11) of the approach.

[0140] Furthermore, the present system may apply not only control of changing the volume but also control of changing the tone according to the degree of approach, for example, when the degree of approach exceeds a threshold, the tone of the voice reading may be changed from a first tone to a second tone.

[0141] The system may also be controlled to add and output an alert sound or vibration depending on the degree of approach. The cycle of the alert sound output may be shortened depending on the degree of approach. For example, when the degree of approach exceeds a threshold, an alert sound or vibration may be added to the reading voice, or the reading voice may be changed to an alert sound or vibration.

[0142] In addition, if a non-visually impaired user is using smart glasses 1C or the like, a display of some kind that accompanies the voice reading, such as a display that indicates that an object is approaching, may be added to the display surface of the device 1.

[0143] The degree of approach may also be regarded as a degree of danger, which will be described later. That is, the greater the degree of approach, the greater the degree of danger, and voice reading or alert output may be applied according to the degree of danger.

[0144] The movement or change of the object (corresponding feature) is not limited to the above example, but may also be read aloud when, for example, a new object appears and is detected within the user's field of view (in the corresponding image), or when the direction of movement of the object changes (see below). Furthermore, a change in the recognition of the object by the AI ​​20 may also be read aloud.

[0145] [Significant Changes in Objects Between Images] Based on a group of captured images, the device 1 or server device 2 identifies and detects portions of the image in which there are significant changes in the surrounding environment / object in a time series, determines those portions as targets for voice reading, and controls the device 1 or server device 2 to prioritize voice reading. Based on the recognition results between images, the device 1 or server device 2 identifies and detects portions of the image in which there are significant changes based on the amount of change in features between the surrounding environment at a first time point (an earlier time point) and the surrounding environment at a second time point (a later time point). For example, portions that are moving differently are detected as objects with significant changes. The device 1 or server device 2 then controls the device 1 or server device 2 to prioritize the portions of the image in the target for voice reading over other portions (i.e., portions with little change).

[0146] As a simple example, the voice reading of an object that has undergone a significant change may be a voice output of only words such as the name of the object. For example, words such as "person," "pedestrian," "bicycle," and "car" may be used. When using short text such as words, it is possible to read out a large number of objects per unit time. More specifically, this voice reading may be a voice output of a written description that describes the situation of the object. For example, written descriptions such as "A person is approaching," "A bicycle is approaching from behind," and "A car is crossing." When written descriptions are used, it is easier for the user to understand the surrounding situation.

[0147] Furthermore, when reading out an object that has undergone a significant change, the system may output a voice based on text describing how the object changed before and after the change. For example, the voice output may be something like, "The stopped car has started moving toward you," or "The traffic light has changed from red to green."

[0148] If the system is set to read out loud only objects that have undergone significant changes and not those that have not undergone significant changes, the same object that has not undergone significant changes will not be read out repeatedly. This reduces the amount of audio information and annoys the user, and may make it easier for them to understand the surrounding situation.

[0149] FIG. 16A illustrates an example of a significant change in an object between images captured by the camera 304. First, assume that no particularly noticeable object (especially a person) is detected on the sidewalk ahead of the user in image 1601 captured at a first time point. Next, in image 1602 captured at a second time point, a pedestrian appears as object OB3 on the sidewalk ahead of the user. This object OB3 is, for example, a pedestrian who appears in the image as if coming out of the road on the left side onto the sidewalk where the user is located. The device 1 or the server device 2 detects such changes and differences between images and, based on these changes and differences, detects the object as having a significant change. The device 1 or the server device 2 then determines the detected object with a significant change as a priority target for voice reading. In the example of FIG. 16A , the text and voice reading of the recognition result for object OB3 may be, for example, "pedestrian" or "pedestrian coming from the left."

[0150] 16B shows another example. In this example, the direction of movement of the object changes significantly between images. In image 1611 at a first time point, the direction of movement of a pedestrian, represented by object OB4 (OB4a), is to the right (+X). In image 1612 at a second time point, the direction of movement of a pedestrian, represented by object OB4 (OB4b), has changed to a forward direction (-Y) toward the user.

[0151] The system also determines the amount of change in the direction of movement of the object corresponding to the features of the recognition result. The device 1 or the server device 2 detects such changes and differences between the images and, based on the changes and differences, detects the object OB4 as having a significant change. The device 1 or the server device 2 then determines the detected object OB4 having a significant change as the object to be prioritized for voice reading. In the example of FIG. 16B , the text and voice reading of the recognition result for the object OB4 may be, for example, "pedestrian" or "a pedestrian coming from the left is turning toward you."

[0152] In addition, with regard to the above control example, if the server device 2, rather than the device 1, makes the judgment through analysis processing, the response information from the server device 2 to the device 1 may include the judgment result (e.g., distance information from the target object) in addition to the text of the recognition result.

[0153] [Processing Sequence] Figure 17 shows an example of a processing sequence of the device 1 and the server device 2 in the system of Figure 1. In step S101, the device 1 turns on the text-to-speech function. For example, the text-to-speech function may be turned on automatically according to settings when the device 1 is started, or the user may perform an operation input to turn on the text-to-speech function. Alternatively, the device 1 may determine the state of the device 1, such as its posture, and automatically turn on the text-to-speech function.

[0154] In step S102, the device 1 reads the setting information ( FIG. 14 ) of the user U1 from the non-volatile memory 303 or the like. Furthermore, in step S102, the device 1 may determine a rule (also referred to as an analysis rule) for analysis (the processing of steps S3 and S4 in FIG. 4 ) or a control mode corresponding to the analysis rule based on the read user setting information. The analysis rule is described below, for example, in the case of the setting of #5 in FIG. 14 ("Determination Condition for Reading from the Main Camera"). For example, the determination conditions may be "B. Approaching Object" as the first priority, "D. Specific Object" as the second priority, and "A. Object in the Direction of Movement" as the third priority, and the analysis rule may be such that analysis (the processing of steps S3 and S4 in FIG. 4 ) is performed based on these priorities / priorities. Furthermore, the control mode may be selected from the control modes shown in FIG. 13 .

[0155] In step S103, in response to the start of use of the image reading function, the device 1 turns on the image input of the camera 403 and inputs images (time-series image frames) from the camera 403.

[0156] In step S104, the device 1 starts communication for cooperation with the server device 2. The device 1 establishes a communication connection with the server device 2 using the communication unit 330 ( FIG. 7A ). The device 1 transmits the user setting information (or information obtained by processing the same) of step S102 to the server device 2 over the established communication. At this time, the device 1 may also transmit user information of user U1, device information of the device 1, and other information to the server device 2. Furthermore, if the device 1 determines an analysis rule or a control mode in step S102, the device 1 may transmit information on the analysis rule or control mode to the server device 2 in step S104.

[0157] In step S105, the server device 2 uses the communication unit 430 (FIG. 7B) to communicate with the device 1 for cooperation and receives information such as user setting information, analysis rules, and control modes from the device 1. Based on this information, the server 2 determines the content of the analysis process to be performed by the server device 2 (the processes of steps S3 and S4 in FIG. 4). In step S105, if the information received from the device 1 does not include an analysis rule or a control mode, the server device 2, rather than the device 1, may determine the analysis rule and the control mode based on the user setting information from the device 1.

[0158] In step S106, the device 1 monitors whether or not there is an operation input by the user U1, and determines whether or not there is a stop instruction, etc. A stop instruction is an instruction to end or pause the image reading function (particularly the series of processes from image capture to voice reading). If there is a stop instruction, etc., the device 1 transmits the stop instruction, etc. to the server device 2, and stops the processing in the device 1. In step S107, the server device 2 stops the corresponding processing in the server device 2 in response to the stop instruction, etc. from the device 1. If there is no stop instruction, etc., the processing from step S108 onwards is similarly repeated as a loop. Note that if the processing is paused in accordance with the stop instruction, and then a start instruction is input based on the operation input by the user U1, the device 1 and the server device 2 resume the processing.

[0159] In step S108, the device 1 waits for a predetermined time period according to the user settings. In this embodiment, the device 1 captures images at a predetermined cycle based on the setting of #3 (first time interval) in FIG. 14, and performs voice reading at a predetermined cycle based on the setting of #4 (second time interval). Step S108 is a standby time for the image capture. Capture is possible at each cycle of the main camera and the sub camera.

[0160] In step S109, device 1 captures and acquires images from the input video of camera 403 and stores them in memory. When multiple cameras are used, the timing of image capture for each camera may be controlled in a predetermined manner, for example, sequential capture as shown in FIG. 9A. The image processing unit 305 in FIG. 7A controls such image capture and stores the acquired images in the image storage unit 306. Note that if the device is equipped with predetermined hardware and does not pose a problem with the processing load, simultaneous capture may be performed by multiple cameras. Furthermore, when capturing images, image information such as the capture date and time, image identification information, and camera identification information is also added to each image data.

[0161] In step S110, the device 1 checks and determines the communication status with the server device 2. The device 1 determines whether communication with the server device 2 is possible or impossible, or whether communication is good or bad. If the communication status with the server device 2 is impossible / poor, the device 1 temporarily disconnects communication with the server device 2, or temporarily suspends transmission of requests and the like while maintaining the communication connection. Examples of an impossible / poor communication status include congestion on the communication network 9 and a large processing load on the AI ​​20 of the server device 2. When disconnecting / stopping communication with the server device 2, the device 1 stores in memory the history of processing up to that point.

[0162] In step S111, the server device 2 temporarily disconnects communication with the device 1 as necessary, based on the communication status between the device 1 and the server device 2 in step S110, or temporarily suspends transmission of responses and the like while maintaining the communication connection. When the server device 2 disconnects / stops communication with the device 1, it saves a history of processing up to that point in memory or a database. This history is retained for a predetermined period or more. The system sets a control flag depending on the communication possible / unable state in steps S110 and S111, and performs processing from step S112 onwards depending on the flag.

[0163] If communication is impossible / poor in step S110, the system controls the server device 2 not to perform analysis processing, including the AI ​​20's recognition processing, and controls the device 1 to perform analysis processing in step S112. If the device 1 is equipped with a function for performing analysis processing on its own, the device 1 can perform analysis processing on its own in step S112. If the device 1 is not equipped with a function for performing analysis processing on its own, the device 1 cannot perform analysis processing on its own in step S112, and both the device 1 and the server device 2 stop analysis processing. In this case, for example, the device 1 notifies the user U1 of the current status and waits for communication with the server device 2 to be restored.

[0164] In step S112, if a function for performing simple analysis processing is implemented on the device 1 side, the simple analysis processing may be performed on the device 1 side. Normally, analysis processing is performed on the server device 2 side, which has abundant computational resources, but if the analysis processing on the server device 2 side is unavailable due to a communication failure, simple analysis processing is performed on the device 1 side (described below). Generally, computational resources on the device 1 side are more limited than on the server device 2 side in the cloud computing system 3. Therefore, the simple analysis processing on the device 1 side refers to processing that is simpler than the analysis processing on the server device 2 side. If the device 1 side has abundant computational resources, it is also possible to perform analysis processing only on the device 1 side, as shown in FIG. 3.

[0165] If communication is possible / good in step S110, the system controls in step S113 to perform analysis processing, including recognition processing by the AI ​​20, on the server device 2 side, rather than on the device 1 side. In step S113, the device 1 transmits an analysis request and an image (the captured image and additional image information in step S109) to the server device 2 for analysis processing on the server device 2 side. Note that when transmitting the request, supplemental information may be attached in addition to the image. The supplemental information is information that can be used in the analysis processing, and may include information that can be detected by sensors of the device 1 (e.g., the position detection unit 311 and the orientation detection unit 312 in FIG. 7A ), such as position, orientation, direction, and acceleration.

[0166] In step S114, the server device 2 receives the analysis request and image from the device 1 and stores them in memory. In response to the request, the server device 2 starts an analysis process including a recognition process by the AI ​​20. First, the server device 2 recognizes the image using the AI ​​20 to extract features related to the object from the image. In other words, the object is detected from the image. Furthermore, the server device 2 converts the object associated with the features through the recognition into text representing the object. As an output from the AI ​​20, feature information including text information associated with the features, position coordinates, size, etc. is obtained. The above-described process for each captured image is repeated in a similar manner at a predetermined cycle.

[0167] Next, in step S115, the server device 2 detects changes (i.e., change points) between the current image and the previous image in the time-series captured image group. The change points correspond to, for example, changes in the position coordinates, size, shape, direction, etc. of the object (i.e., image area) associated with the feature.

[0168] Next, in step S116, the server device 2 determines the distance, speed, and direction of change (in other words, the direction of movement) of the object in the image and between the images, depending on the applied analysis rule, control mode, etc. The server device 2 may determine the distance, speed, and direction of change by analyzing the image. The server device 2 may also determine significant changes, the degree of approach, the degree of danger, and the safe space, etc., related to the object. Note that, in this embodiment, the server device 2 makes this determination in step S116, but such a determination may also be made on the device 1 side after obtaining information about the object from the server device 2 side.

[0169] In step S117, the server device 2 determines the image, object / surrounding circumstances, text, etc. to be read aloud based on the results of the analysis process and determination process up to step S116. The server device 2 selects and determines the object, etc. to be read aloud based on the characteristics, change points, distance, speed, change direction, etc. from the image in accordance with the conditions of the analysis rule, in accordance with the analysis rule, control mode, etc., in accordance with the priority of #7 in Fig. 14, for example. The text associated with the determined object is also determined.

[0170] In step S118, the server device 2 transmits a response to the device 1 as the analysis result, including information such as text about the object to be read aloud determined in step S117. When transmitting the response from the server device 2 to the device 1, the server device 2 may also add processing result information (e.g., priority / priority order for reading aloud multiple objects) based on the analysis rule / control mode, etc., of the analysis process or determination process performed by the server device 2. In step S119, the device 1 receives the response from the server device 2 and, based on the text and additional information about the object to be read aloud, converts the text into speech and outputs the speech from an audio output device such as a speaker, thereby performing a speech readout. If the target text is assigned a priority / priority order, the device 1 controls the speech readout accordingly. For example, the device 1 reads aloud multiple texts in the order of priority. Furthermore, for example, the device 1 changes the volume of the speech readout depending on the degree of proximity. Furthermore, for example, if an alert output instruction is assigned, the device 1 outputs an alert sound from the audio output device.

[0171] The voice reading in step S119 is set to be performed at a predetermined cycle (second time interval). Therefore, the system controls the processes from step S112 to step S119 to be performed at the timing according to the cycle. Note that if there is no object to be read out during the periodic voice reading, for example, if there is no significant change, the response in step S118 will be a response indicating that there is no object to be read out (no change) and that there is no target text.

[0172] 17 illustrates a case where a series of processes from image capture to voice reading are automatically and periodically executed, but when a read-aloud instruction (a read-aloud re-instruction, described later) is input by the user U1 at any timing, the following control may be performed. That is, when the read-aloud instruction is input, if the communication state with the server device 2 is available / good, the device 1 captures an image and transmits the image and a request corresponding to the read-aloud instruction to the server device 2. In response to the request, the server device 2 similarly performs an analysis process and transmits a response including the text of the analysis result to the device 1. The device 1 performs voice reading based on the text of the response.

[0173] In this embodiment, as in the above example, a series of processes from image capture to voice reading is automatically executed at a predetermined cycle from multiple cameras of the device 1. There is basically no need for a user to perform instruction operations such as image capture, which reduces the effort and is highly convenient.

[0174] This system is also capable of reading out all objects recognized in an image, as long as time and computational resources allow. However, reading out a large number of objects at once may be difficult for the user to understand and may confuse them. Therefore, in this embodiment, the object to be read out is determined according to priority, proximity, etc., during the automatic image reading cycle, and the most important objects are read out first. This allows for easy and convenient support for the user. The user can use different user setting information, control modes, etc. depending on the usage scenario, app, etc.

[0175] Furthermore, if no object is detected in the time series or if there is little change in the object, the number of voice readouts may decrease or the same voice readout may be repeated. In such cases, the user may feel uneasy. Therefore, this embodiment also supports voice readout in response to a user's input of a readout instruction at any time, and, as will be described later, has a function to read out that there is no change in the surrounding situation when there is no change. This allows for easy-to-understand and suitable support for the user.

[0176] 17, in step S118, the server device 2 may convert the text representing the object into voice data for reading aloud, and transmit a response including the voice data to the device 1. In this case, the device 1 does not need to convert the text into voice.

[0177] [Basic Control] Fig. 18 is an explanatory diagram regarding the basic control of automatic image capture and voice reading in time series. Voice reading cycle timing 1801 shows, for example, time points 1 to 7. Images 1802 at each time point are example images at the timing of image capture corresponding to the time point in the voice reading cycle. Voice reading 1803 shows an example of the audio content of the voice reading at that time point. Repetition 1804 indicates the number of times the voice reading is repeated in time series, etc.

[0178] At time point 1, it is assumed that no particular object is detected in the image. At time point 1, no voice reading is performed (illustrated as "none"). At time point 2, a pedestrian 1811 is detected as an object 1811, as an example of a significant change, similar to FIG. 16A . At time point 3, the pedestrian 1811 is present on the sidewalk ahead of the user U1. At time points 4 and 5, the state is the same as at time point 3, with almost no change in the pedestrian 1811. At time points 6 and 7, the pedestrian 1811 is moving closer to the user U1. At each time point after time point 2, a "pedestrian" is detected as the object 1811 in the image. Therefore, in the control example of FIG. 18 , at each time point after time point 2, a voice reading such as "pedestrian" (or, specifically, "pedestrian approaching") is output, and the same voice reading "pedestrian" is output multiple times. Time point 2 is the first repetition, and time point 7 is the sixth repetition.

[0179] As in the above example, the user U1 can recognize that there is a "pedestrian" ahead from the voice reading at each point in time, without having to perform any particular operation input.

[0180] The device 1 stores historical data of processes, including image capture and voice reading, related to the above-described control in a storage resource. The device 1 stores the historical data in the memory of its own device or the server device 2 up to an upper limit, such as a predetermined time, a predetermined number of times, or a predetermined data amount. If the historical data exceeds the storage limit, the device 1 or the server device 2 overwrites and erases the data, starting with the oldest.

[0181] [Control for not performing voice reading if there is no change] Next, Fig. 19 is an explanatory diagram regarding control for not performing voice reading if there is no change in the object, based on the basic control of Fig. 18. The example of "pedestrian" as the object 1811 is the same as Fig. 18. In the transition from time point 3 to time point 4, in other words, between the images, there is almost no change in "pedestrian" as the object 1811. Therefore, in this control example, at time point 4, device 1 does not perform voice reading (illustrated as "none"). Similarly, in the transition from time point 4 to time point 5, there is no change, so voice reading is not performed.

[0182] The image reading system may determine that there is no change in the object between images if the amount of change in the object (for example, the amount of change in the position coordinates or size of features) is less than a threshold value based on the recognition results of AI20.

[0183] In the transition from time point 5 to time point 6, the "pedestrian" as the object 1811 approaches the user U1. The system determines that there is a change in the "pedestrian" as the object 1811 from the amount of change between the images. Therefore, at time point 6, for the object 1811 with the change, a voice readout such as "pedestrian" (or specifically "pedestrian approaching") is output. Considering the number of repetitions from time point 2, time point 6 is the third time. Similarly, at time point 7, it is determined that there is a change, so a fourth voice readout is performed.

[0184] As in the above example, if there is no change in the object, there is no voice reading, so the volume and frequency of voices heard by the user U1 are reduced.

[0185] [Control for Reading Out a Message Indicating No Change When No Change Occurs] In FIGS. 18 and 19 , for example, from time point 3 to time point 5, there is almost no change in the objects surrounding user U1. When such a state of no change in the objects continues, a message may be read out periodically to inform the user that there is no change, as in the control example of FIG. 20 . For example, the system determines whether the degree of change (e.g., amount of change) in the objects remains below a threshold for a predetermined period of time or longer. The time count may be performed in units of a predetermined cycle or image. The predetermined period of time for this determination is configurable. If this state continues for a predetermined period of time or longer, the system determines to read out a message indicating that there is no change. The device 1 reads out a message indicating that there is no change, such as "No change."

[0186] In the example of FIG. 20 , at time points 2 and 3, the object 1811, "pedestrian," in front of the user U1 is detected as having changed, and therefore "pedestrian" is read out as described above. At time point 4, it is determined that there is no change in the "pedestrian" as the object 1811. If this state of no change continues for, for example, one cycle of the read-out cycle, it is determined that there is no change to be read out. Therefore, at time point 4, "no change" is read out. Similarly, at the next time point 5, it is determined that there is no change in the "pedestrian," and this state of no change continues for one cycle, so "no change" is read out.

[0187] In another example, if the specified time during which no change continues is set to two cycles of the voice reading period, at time point 4, no voice reading will be performed, and at time point 5, the state of no change has continued for two cycles, so the voice reading will say "No change."

[0188] As in the above example, if there is no change in the object, a voice readout is made to the effect that there is no change, so the anxiety of the user U1 when there is no voice readout is reduced.

[0189] [Function of Instruction to Read Again] Fig. 21 is an explanatory diagram showing a specific example of the function of instruction to read again. User U1 can input an instruction to read again (in other words, an instruction to forcibly read aloud) to device 1 at any timing. The instruction to read again by user U1 can be, for example, a method of voice input using a microphone or the like (microphone / voice input unit 316 in Fig. 7A), or can be input by pressing a predetermined button on remote control 1c in Fig. 5A. For example, pressing a predetermined button on remote control 1c represents an instruction to read again.

[0190] When this read-again instruction is input, device 1 always executes a voice readout of the object at that time and outputs some kind of audio. In response to the operation input of the read-again instruction, device 1 captures an image from the input video, causes AI 20 to recognize the image, and performs a voice readout of the object / surrounding situation. Even if there is no change in the object at the time of input of the read-again instruction, device 1 performs a voice readout that conveys the object or surrounding situation.

[0191] The images at each time point in FIG. 21 are generally similar to those in FIG. 18 , etc. At time points 2 and 3, a "pedestrian" is detected as the object 1811 as a significant change, and the device 1 automatically outputs, for example, "pedestrian" as a read-out voice. Assume that there is almost no change in the "pedestrian" as the object 1811 from time point 4 onward. Since it is determined that there is no change in the "pedestrian" as the object 1811 at time point 4, the device 1 does not perform read-out voice unless there is a particular action from the user U1. Similarly, since it is determined that there is no change at time point 5, the device 1 does not perform read-out voice unless there is a particular action from the user U1.

[0192] On the other hand, for example, assume that a read-aloud instruction is input by user U1 immediately before time point 5. In this case, device 1 follows the read-aloud instruction and performs a read-aloud for object 1811 at time point 5. Device 1 outputs, for example, "pedestrian" as the read-aloud. Alternatively, as a modified example, similar to FIG. 20 , since there is no change in the object, device 1 may output "No change" or the like.

[0193] At time point 6, similarly, it is determined that there is no change in the object 1811, and therefore the device 1 does not perform voice reading unless there is a particular action from the user U1. Assume that the user U1 again inputs a read-aloud instruction at time point 7. In this case, the device 1 similarly follows the read-aloud instruction and performs voice reading at time point 7, outputting, for example, "pedestrian" (or "no change").

[0194] In the above example, the timing at which the read-aloud instruction is input is almost the same as the periodic timing of the read-aloud, such as time point 5 or time point 7, but these timings may be different. In the above example, at time point 5 or time point 7, the device 1 basically transmits a request to the server device 2 in response to the read-aloud instruction, receives a response from the server device 2, and performs read-aloud using the response. Even if there is no change in the object compared to the results of the previous recognition and read-aloud, the device 1 performs read-aloud again based on the current recognition.

[0195] As a variant example, when a reading instruction is input again (especially at a time other than the periodic timing of the voice reading), the device 1 may omit communication with the server device 2 and perform the voice reading using the results of the previous voice reading (e.g., the same text) that have been stored as history up to that point.

[0196] As in the above example, in addition to the automatic voice reading by this system, user U1 can check the surrounding situation by having voice reading performed immediately at the timing of his / her choice. Furthermore, the function of instructing to read aloud again can be used effectively in combination with the function of voice reading with different expressions shown in Figure 22 below.

[0197] [Function of Reading Aloud Using Different Expressions (1)] As in the above control example, reading aloud is repeated multiple times in a time series by automatic and periodic reading aloud or by reading aloud in response to a user U1's instruction to read aloud again. When the surrounding circumstances / object are almost the same at each time point and there are no significant changes, reading aloud modes such as those in each control example are possible, and the same audio output may be repeated. However, as another example, the following may also be used. In another example, when reading aloud multiple times for a similar surrounding circumstances / object, control may be exercised so that different text is read aloud each time, explaining it from a different perspective or using different expressions.

[0198] FIG. 22 is an explanatory diagram showing a specific example of the function of text-to-speech with different expressions. The situation of the object in the image in FIG. 22 is assumed to be the same as that in FIG. 16B . In FIG. 22 , a "pedestrian" coming out from the left side of user U1 is detected as object 2201 in image 2200. Assume that there is almost no change in the situation of object 2201 in the time-series image group. Assume that the function of text-to-speech with different expressions is set to ON in device 1. For example, at time point 1, device 1 detects a significant change in the response from server device 2 and outputs, for example, "A pedestrian is coming from the left" as the first automatic text-to-speech.

[0199] Next, time point 2 is assumed to be the timing of a predetermined cycle of voice reading, or the timing when user U1 inputs an operation to instruct reading again as in FIG. 21 . At this time point 2, device 1 executes a second voice reading for the object 2201, "pedestrian." When repeatedly reading aloud the same object 2201 that has not changed, device 1 controls the reading to be performed using a different expression from the previous (first) reading. Based on the response from server device 2, device 1 outputs, for example, "There is a person diagonally ahead of you to the left," in the second voice reading.

[0200] Next, time point 3 is assumed to be the timing of a predetermined cycle of voice reading or the timing when user U1 performs an operation input for a read-out instruction again. At this time point 3, device 1 executes a third voice reading for object 2201. Device 1 controls to perform voice reading using an expression different from the first and second times. Based on the response from server device 2, device 1 outputs, for example, "A person is coming out of the road on the left" in the third voice reading.

[0201] To achieve such a function, the device 1 may cause the AI ​​20 to perform recognition again at each timing of voice reading, and acquire text in a different expression from the AI ​​20. Alternatively, the device 1 may acquire and store text in multiple expressions collectively from the AI ​​20 when the AI ​​20 recognizes the text in one request-response, and select one expression from the multiple expressions acquired and output it aloud at each voice reading.

[0202] Another specific example is as follows. Suppose that image recognition detects a situation in which a man walking in front of the user pushing a bicycle approaches the user as an object. At a first point in time, the first voice readout is, for example, "Man and bicycle." At a second point in time, the second voice readout is expressed differently, for example, "A man is walking pushing a bicycle." At a third point in time, the third voice readout is expressed even differently, for example, "A man pushing a bicycle is approaching."

[0203] As in the example above, there can be multiple ways to describe a similar situation / object. This feature reads out different sentences at different times, making it easier for the user to understand the situation.

[0204] [Function of Reading Aloud Using Different Expressions (2)] Furthermore, the following control is also possible as a modified example of the function of reading aloud using different expressions. The system may perform reading aloud by adding or synthesizing the current recognition result to the content read aloud of the previous recognition result at each timing of reading aloud, such as automatically, periodically, or in response to an instruction to read aloud again. In other words, if the system determines at each timing that there has been a change in the object, it may perform reading aloud of both the audio representing the previous object and the audio representing the part that has changed this time.

[0205] As an example, at a first time point, in response to a first re-reading instruction, the AI20 recognizes a "pedestrian" as an object in front of the user U1, and reads out "pedestrian." Next, at a second time point, in response to a second re-reading instruction, the AI20 recognizes a "bicycle" (or a "person riding a bicycle") as an object separate from the previous "pedestrian." In this case, the device 1 outputs a voice reading that combines both the previous "pedestrian" and the current "bicycle," such as "pedestrian and bicycle." Next, at a third time point, in response to a third re-reading instruction, the AI20 recognizes a "dog" as another object separate from the previous "pedestrian" and "bicycle." The device 1 outputs a voice reading that combines these, such as "person, bicycle, and dog." In this example, the nouns of multiple objects are connected in parallel and read aloud, but this is not limited to this. It is also possible to make a sentence that expresses the situation of one or more objects in more detail, for example, "A person riding a bicycle is walking a dog."

[0206] [Function of Reading Out Loud with Different Expressions (3)] The following modification is also possible. The present system may read out aloud an object using a different expression in response to updating / correcting the recognition by the AI20 at each timing of reading out loud, such as automatically or in response to a read-out instruction. For example, at a first time point, in response to a first read-out instruction, a "pedestrian" is detected in front of the user U1, and, for example, "pedestrian" is output as a voice. At a second time point, in response to a second read-out instruction, as a result of the AI20 performing the second recognition, the recognition of the previous object, "pedestrian," is updated / corrected, and the object is more accurately detected as a "bicyclist," and, for example, "bicyclist." Furthermore, at a third time point, in response to a third read-out instruction, as a result of the AI20 performing the second recognition, the recognition of the previous "bicyclist" is updated / corrected, and the object is more accurately detected as a "bicyclist with a dog," and, for example, "bicyclist with a dog" is output as a voice.

[0207] Furthermore, when device 1 performs a voice readout based on the result of correcting the previous content to the current content as described above, the current voice readout may be a voice readout that notifies the user of the correction. For example, the second voice readout may output a voice such as "Correction: This is not a pedestrian, but a person riding a bicycle."

[0208] [Control Example of Front and Rear Cameras] Figure 23 shows an example of control of front and rear cameras, assuming the use of multiple cameras as in Figure 9A, etc. Figure 23 particularly shows an example of control that switches between control in which the front camera is the main camera and control in which the rear camera is the main camera. This control enables functions such as interruption. For example, the front and rear cameras are used in accordance with the control mode (M1) in Figure 13 described above. In Figure 23, as an example of a situation, assume that a user U1 is walking on a sidewalk 901, and there is a pedestrian 2301 in front of him and a bicycle 2302 behind him. Assume that the device 1 of user U1, for example, a smartphone 1A, has a front camera C1 and a rear camera C2, similar to Figure 9A.

[0209] Under normal circumstances, the present system performs control 1 shown in the figure, for example, by frequently and preferentially capturing, recognizing, and reading out images in the forward direction (+Y) using the front camera C1. Control 1 corresponds to control B in FIG. 9C, and the processing ratio between the front and rear cameras is set to 3:1. In control 1, changes in the pedestrian 2301 can be detected with high accuracy from the image captured by the front camera C1, and a voice reading can be performed, for example, saying "A pedestrian is approaching."

[0210] The present system automatically switches from Control 1 to Control 2 as a result of a predetermined judgment or the like. For example, during Control 1, at a periodic timing such as time point 4, an object located behind the user U1 is detected from the image captured by the rear camera C2. For example, a bicycle 2302 moving toward the user U1 can be detected as a significant change. For example, when an object with a significant change in the rear direction is detected from the image captured by the rear camera C2, the device 1 switches to Control 2.

[0211] Control 2 shown in the figure prioritizes the frequent capture, recognition, and reading of images in the rear direction (-Y) using the rear camera C2. Control 2 corresponds to control D in FIG. 9C, and the processing ratio between the front and rear cameras is set to 1:3. In control 2, changes in the bicycle 2302 can be detected with high accuracy from the image captured by the rear camera C2, and a voice reading can be performed, for example, saying "A bicycle is approaching."

[0212] The above control example of switching from Control 1 to Control 2 can achieve a function and effect of interrupting the recognition of the rearward direction for user U1 with the recognition of the forward direction. Automatic switching from Control 2 to Control 1 can also be achieved in a similar manner. When an object with a significant change is detected from the image of the front camera during Control 2, the system switches to Control 1. In other words, this function allows the recognition and voice reading of the sub-camera to be realized as an interruption to the recognition and voice reading of the main camera. Similar control is possible when using front, rear, left, and right cameras.

[0213] In a variant, when objects are detected both forward and backward, the system may compare the degree of change and priority of the objects in front and behind the vehicle to determine which of the two is more important, in other words, which object requires more attention, and select between Control 1 and Control 2 based on the determination result. In the example shown in the figure, when a pedestrian 2301 and a bicycle 2302 are detected, it is determined that the bicycle 2302 has a higher priority, for example, based on the relative distance and speed between them. In that case, the system switches to Control 2 or maintains Control 2 and prioritizes voice reading for the bicycle 2302.

[0214] As a modification, a control that treats the front and rear directions equally, such as the state A in Fig. 9A or the control C in Fig. 9C, may be interposed between control 1 and control 2. For example, the present system normally maintains control C, and transitions to control B when an object with a significant change is detected in the image from the front camera, and transitions to control D when an object with a significant change is detected in the image from the rear camera.

[0215] In a modified example, control 1 in FIG. 23 may be replaced with control A in FIG. 9C described above, and control 2 may be replaced with control E described above, and switching may be performed between control A and control E. However, in this case, for example, when control A is in use, there is no detection in the image from the rear camera, so the predetermined judgment for switching must be based on a different condition. An example of such a condition is when, when control A is in use, a state in which no object with a significant change is detected in the image from the front camera continues for more than a predetermined time, and then switching to control E is performed. Alternatively, the front and rear controls may be switched simply in response to a button operation input or voice input by user U1.

[0216] [Control Using Different Front and Rear Thresholds] FIG. 24 illustrates a control example for determining significant changes in an object (when the object moves differently from previous images) using front and rear cameras capturing images in the front and rear directions relative to user U1. In particular, the control example of FIG. 24 illustrates a case in which different thresholds are used for the images and recognition from the front and rear cameras of device 1. This control example is based on the control related to the relative distance and speed between user U1 and the object, as described above ( FIG. 11 ). In FIG. 24 , for images captured by front camera C1 capturing images in front of user U1, a distance threshold TD1 and a speed threshold TV1 are used as the thresholds for determination. On the other hand, for images captured by rear camera C2 capturing images behind user U1, a distance threshold TD2 and a speed threshold TV2 are used as the thresholds for determination. Situation A in FIG. 24 is an explanatory diagram illustrating the determination of relative distance, and situation B is an explanatory diagram illustrating the determination of relative speed.

[0217] For example, as in control 1 in Figure 23, assume that the front camera C1 is the main camera and the rear camera C2 is the sub-camera, and images, detection, and voice reading are performed with a higher frequency, with the front being given priority over the rear. In this case, from the viewpoint of frequency, rear objects are more difficult to detect than front objects. Therefore, the thresholds (TD2, TV2) for determining whether to read out an object from the image captured by the rear camera C2 may be set to different thresholds than the thresholds (TD1, TV1) for determining whether to read out an object from the image captured by the front camera C1, so that the detection is easier.

[0218] In the example of situation A, the rear distance threshold TD2 is set to a larger value than the front distance threshold TD1 (TD1<TD2) to prioritize the rear. In this example, TD2 is about twice TD1. In situation A, a pedestrian 2401 is proceeding in the -Y direction from in front of the user U1, and another pedestrian 2402 is proceeding in the +Y direction from behind the user U1. On the front side, based on analysis of the image from the front camera C1, if the distance D2401 to the pedestrian 2401 is within the distance threshold TD1, the pedestrian 2401 is determined to be a target for voice reading. On the rear side, based on analysis of the image from the rear camera C2, if the distance D2402 to the pedestrian 2402 is within the distance threshold TD2, the pedestrian 2402 is determined to be a target for voice reading. For example, even if the distances D2401 and D2402 are approximately the same, the distance thresholds are different, so the pedestrian 2402 at the rear is detected first and given priority, and a voice reading (e.g., "A pedestrian is approaching from behind") is performed.

[0219] In situation B, assume that user U1 is stationary, speed v1 = 0, a pedestrian 2401 in front is moving toward user U1 in the -Y direction at speed v2A, and a pedestrian 2402 behind is moving toward user U1 in the +Y direction at speed v2B. In the example of situation B, rear priority is given to the pedestrian, and a rear speed threshold TV2 is set to a smaller value than a front speed threshold TV1 (TV1 > TV2). On the front side, based on analysis of images from the front camera C1, if the relative speed of pedestrian 2401 with respect to speed v1A is equal to or greater than speed threshold TV1, the pedestrian 2401 is determined to be a target for voice reading. On the rear side, based on analysis of images from the rear camera C2, if the relative speed of pedestrian 2402 with respect to speed v2B is equal to or greater than speed threshold TV2, the pedestrian 2402 is determined to be a target for voice reading. For example, even if the speeds v1A and v1B are approximately the same, the speed thresholds are different, so the pedestrian 2402 behind is detected first and then read out. Similar control is possible using the relative speeds even when the user U1 is moving.

[0220] 24 can also be applied when recognition and voice reading are performed equally by the front and rear cameras, as in state A in Fig. 9A . For example, if user U1 is not visually impaired, it is possible to prioritize the rear and set the threshold value for the rear side to be more sensitive so that objects at the rear can be detected more sensitively than objects at the front, taking into consideration that the front is easy to see and the rear is difficult to see.

[0221] Furthermore, the above-described determination of relative distance and determination of relative speed may be used alone or in combination. In the case of combined use, for example, an object may be determined as a target for voice reading if an AND (logical product) condition using both thresholds is met. Similar threshold control can be applied to each direction when front, rear, left, and right cameras are used. Various settings for the above-described control thresholds are possible depending on the user, app, usage scenario, etc. Various thresholds can be variably set using the user setting function, and a mode corresponding to each setting can be selected and used.

[0222] [Support for the Visually Impaired] FIG. 25A is an explanatory diagram of a usage scenario in which support for the visually impaired is provided. For example, this is a Y-Z plan view of a situation in which a visually impaired user U1 is standing still or walking on a sidewalk 2500. For example, similar to FIG. 6B and other figures, a device 1 such as a hat 1D has cameras (in this example, a front camera C3 and a rear camera C4) that capture images in the front and rear directions. It is preferable that the front camera C3 and other cameras have a capture range that can capture images up to the feet of the user U1. The user U1 walks while tapping the road surface, such as the sidewalk 2500, with a white cane, for example. In Situation 1, there is an obstacle 2501 on the road surface ahead of the user U1. Even in this case, the obstacle 2501 is captured, for example, based on an image from the front camera C3, recognized as an object by the AI ​​20, and converted into text (e.g., "obstacle" / "situation with an obstacle ahead"). The device 1 then reads out the object aloud (for example, "There is an obstacle ahead"), thereby supporting the walking of visually impaired people.

[0223] In situation 2, a person 2502 is approaching the user U1 from behind. Even in this case, the person 2402 is captured, for example, based on an image from the rear camera C4, recognized as an object by the AI ​​20, and converted into text (e.g., "person" / "A situation in which a person is approaching from behind"). The device 1 then reads out the object aloud (e.g., "A person is approaching from behind"). In particular, the system may be controlled to increase the degree of approach or danger and output an additional alert depending on the distance and speed of the approaching object. For example, a voice or alert sound such as "Please be careful" is output. This allows for more effective support for the walking of visually impaired people.

[0224] For example, in the case of a visually impaired user who cannot see clearly both forward and backward, it is desirable to control multiple cameras so that both forward and backward voice readings are performed. In this case, for example, control with an equal ratio between the front and rear may be applied as shown in Fig. 9A above, or control that switches between the front and rear controls may be applied as shown in Fig. 23.

[0225] [Nighttime Road Assistance] FIG. 25B is an explanatory diagram of nighttime road assistance as one example of a usage scenario. For example, on a nighttime road 2510, images from the front and rear cameras of a smartphone 1A or the like are used as the device 1 of the user U1. The system also uses a rear camera C2 to detect whether a suspicious person (person) 2511 is following the user U1 from behind, and performs a voice readout to warn the user (e.g., "Someone is approaching from behind," "Be careful"). In addition to the voice readout, an alert sound may be output, the device 1 may vibrate, or the volume of the voice readout may be increased above normal. Furthermore, the brightness of the light emitted from the screen of the smartphone 1A or the like may be temporarily increased. For example, increasing the volume of the voice readout may have the effect of intimidating the suspicious person 2511.

[0226] The system may determine the degree of suspiciousness of a suspicious person (person) 2511, which is an object detected behind the user U1, and change the manner of voice reading (for example, volume, tone, alert, etc.) depending on the degree of suspiciousness. Examples of methods for determining the degree of suspiciousness include determining whether a person detected behind the user U1 is following the user U1 while maintaining a certain distance between them, or whether the person appears to be hiding.

[0227] Furthermore, when the device 1 is applied to nighttime road assistance, the camera 403 of the device 1 may be an infrared camera that can capture images even in dark environments.

[0228] [Function for Handling "Walking While Using a Smartphone"] FIG. 25C is an explanatory diagram illustrating a function for handling "walking while using a smartphone" as one usage scenario. In a generally safe situation, such as a sidewalk, user U1 holds a smartphone 1A as device 1 and walks while searching for a destination while viewing a map using a map app, for example. Using images captured by the front and rear cameras of smartphone 1A, objects are read aloud to user U1 from an image of the surroundings. For example, various objects listed on the map, particularly facilities and intersections along the route to the destination, may be detected as objects and read aloud. Furthermore, to ensure safety while "walking while using a smartphone," objects approaching the user (such as people, bicycles, and cars) are detected and read aloud in the same manner as in the above-described control. In this case, approaching objects are given priority over facilities and other objects. In the illustrated example, "intersection A" and "convenience store A" are detected in the user U1's direction of travel and read aloud. Furthermore, when a bicycle 2521 approaches the user U1, it is detected and a voice readout is performed, such as "A bicycle is coming from the right. Please be careful." In this way, safety while "walking while using a smartphone" can be improved, and it is also possible to link with map apps, etc.

[0229] [Vehicle Driving Assistance] FIG. 25D is an explanatory diagram relating to vehicle driving assistance as one usage scenario. For example, a user U1 is riding an electric bicycle 2531. While driving the electric bicycle 2531, the camera 403 of a device 1, such as smart glasses 1C or a smartphone 1A, worn by the user U1 captures images of the surroundings, detects objects, and performs voice reading. For example, if there are objects (people, bicycles, cars, etc.) approaching the user U1 riding the electric bicycle 2531 in directions such as the front, back, left, and right, voice reading of those objects is performed preferentially. This can assist driving and contribute to traffic safety.

[0230] [Example of Control Depending on Whether Communication with the Server is Possible] FIG. 26 illustrates an example of control depending on whether communication with the server device 2 is possible, and illustrates an example of control relating to step 110 of FIG. 17 . State A is an example of control in a state where communication with the server device 2 is possible, in which analysis processing is performed on the server device 2 side. The device 1 (e.g., smart glasses 1C) applies, for example, control A1 to the four cameras (C10, C20, C30, C40) in the front, rear, left, and right directions. Control A1 sets equal ratios for the four directions and sequentially transmits requests to the server device 2 to perform analysis processing including recognition by AI20. For example, at a first point in time, request 1 is transmitted to analyze an image from the front camera C10. At a second point in time, request 2 is transmitted to analyze an image from the rear camera C20. At a third point in time, request 3 is transmitted to analyze an image from the right camera C30. At a fourth point in time, request 4 is transmitted to analyze an image from the left camera C40. From the fifth point onward, the process is repeated in the same manner, starting with the front camera C10. The server device 2 performs analysis processing in response to each request and transmits each response. In response to each response, the device 1 performs voice reading if it is time for voice reading and there is a target. As mentioned above, it is also possible to set a ratio and prioritize a camera in a specific direction as the main camera.

[0231] State B is a control example in a state where communication with the server device 2 is not possible, and is a case where analysis processing is performed on the device 1 side. Here, it is assumed that the device 1 has fewer computational resources than the server device 2, in other words, lower computational power, and the analysis processing takes time. Therefore, here, the analysis processing on the device 1 side is simpler than that of the server device 2. Here, the simple analysis processing is analysis processing that keeps the processing load on the device 1 low. In order to reduce the processing load, control B1 sets the processing ratio so that only some of the cameras in the four directions are used (on state). For example, only the front and rear cameras are used, the left and right cameras are turned off, and the front and rear cameras are set to an equal ratio.

[0232] For example, at a first point in time, the device 1 analyzes the image captured by the front camera C10. At a second point in time, the device 1 pauses the analysis process or continues the analysis process of the image captured by the front camera C10 at the first point in time. At a third point in time, the device 1 analyzes the image captured by the rear camera C20. At a fourth point in time, the device 1 pauses the analysis process or continues the analysis process of the image captured by the rear camera C20 at the third point in time. This process is repeated from the fifth point in time onwards. At each point in time at which the simple analysis process is performed, the device 1 performs voice reading if it is time for voice reading and there is a target. In the case of control B1, the cycle for image capture and voice reading is longer than, for example, control 1 of FIG. 9A, resulting in a lower processing load.

[0233] In this way, even when communication with the server device 2 is not possible, the user U1 can be assisted by simple analysis on the device 1 side. Similarly, control can be performed to limit analysis processing to cameras in a specific direction. Control B2 is an example of limiting analysis processing to only the front camera C10. At the first and third time points, the device 1 analyzes images from the front camera C10. Furthermore, if the analysis is set to be performed only at the first time point, the processing load can be further reduced.

[0234] As another example of control, control may be performed according to a multi-value state, not limited to the binary state of whether or not communication with the server device 2 is possible. For example, the present system may grasp the communication speed or the like as the performance of communication between the device 1 and the server device 2, and apply different control depending on the communication speed or the like. For example, normal control may be applied when the communication speed is equal to or greater than a first threshold, low-speed control may be applied when the communication speed is less than the first threshold, and even lower-speed control may be applied when the communication speed is less than a second threshold.

[0235] [Priority of Multiple Objects] Figure 27 shows an example of control for performing voice reading according to the priority of objects when multiple objects are detected in the surroundings of user U1. Assume that user U1 is walking forward (+Y) on the left side of a road 2700. A car 2701 detected ahead of user U1 is stopped. Pedestrians 2702 and 2703 are also detected ahead, walking in the -Y direction toward user U1. A bicycle 2704 is also detected behind, traveling in the +Y direction toward user U1. Assume that device 1 detects these four objects at approximately the same time.

[0236] The device 1 determines the priority of these four objects and performs voice reading in order of priority. Examples of priority determination are as follows. First, for example, since the car 2701 is stopped, it is assigned a low priority based on its relative distance and speed from the user U1. Since the pedestrian 2702 is approaching the user U1, it is assigned a high priority based on the above-mentioned relative distance, speed, and degree of approach. Furthermore, for the moving pedestrians 2702, 2703, and bicycle 2604, their movement paths and the possibility of contact with the user U1 are determined. The pedestrian 2703 is approaching the user U1 in the Y direction, but its predicted path k3 is different from the path k1 of the user U1's movement. In other words, the pedestrian 2703 is assigned a low priority because it is unlikely to come into contact with the user U1. The pedestrian 2702 is approaching the user U1 in the Y direction, and the predicted path k2 is close to the path k1 of the user U1's movement (e.g., the distance in the X direction is small). In other words, the pedestrian 2702 is likely to come into contact with the user U1, and therefore is given a high priority.

[0237] The bicycle 2704 is approaching the user U1 in the Y direction and is assigned a high priority based on its relative distance and speed to the user U1. Furthermore, the bicycle 2704's predicted route k4 is close to the route k1 of the user U1's movement. In other words, the bicycle 2704 is assigned a high priority because it may come into contact with the user U1.

[0238] Furthermore, the device 1 or the server device 2 compares the pedestrian 2702, which is assigned a high priority, with the bicycle 2704, and assigns priorities by, for example, comparing their relative speeds. For example, if the bicycle 2704 has a higher relative speed than the pedestrian 2702, the bicycle 2704 is assigned the first priority and the pedestrian 2702 is assigned the second priority. Based on these priorities, the device 1 reads out the text aloud in the order of bicycle 2704, pedestrian 2702, and car 2701, for example. For example, the voice output may be, first, "A bicycle is coming from behind," second, "A pedestrian is coming from the front," and third, "A car is stopped in front."

[0239] In this way, voice reading is performed in order of priority taking into consideration the degree of proximity, route, etc., thereby increasing the safety of the user U1 and those around him.

[0240] [Control According to the Risk Level of Risk Factors and Guidance of Safe Spaces] Figure 28 shows an example of control according to the risk level of risk factors and an example of control for providing guidance on safe spaces. The device 1 and the server device 2 may set a specific object as a risk factor, or may determine the object as a risk factor based on the relative distance and speed to the object. Here, a risk factor is an object that may come into contact with the user U1 and pose a risk to the user U1.

[0241] As an example of setting a predetermined object as a risk factor in advance, the roadway 903 and the car 1003 running on the roadway 903 in Fig. 10 may be set as risk factors. In another example, a tool capable of killing or injuring may be set as a risk factor.

[0242] In the example of FIG. 25 , it is assumed that user U1 is walking forward (in the +Y direction) on the left side of road 2800. Device 1 or server device 2 assumes and calculates a route k10 in the direction of movement and travel of user U1 based on images from camera 403 and sensors. It is assumed that a bicycle 2801 and a pedestrian 2802 are detected as objects ahead of user U1. The bicycle 2801 is traveling in the -Y direction toward user U1. The pedestrian 2802 is walking in the +Y direction, in the same direction as user U1. Device 1 or server device 2 assumes and calculates a route k11 for bicycle 2801. It is determined that route k11 for bicycle 2801 is close to route k10 for user U1, and that there is a possibility of contact if things continue as they are.

[0243] The device 1 or the server device 2 determines and sets the bicycle 2801 as a risk factor based on the determination. The device 1 or the server device 2 may also determine and set a risk level for the risk factor. For example, the risk level may be calculated based on the relative distance or speed between the user U1 and the bicycle 2801, or based on the proximity between the route k10 of the user U1 and the route k11 of the bicycle 2801. Even if the route of the user U1 and the route of the object are in different directions, the object may be set as a risk factor if the routes are expected to intersect.

[0244] The risk level of a risk factor may be a single value, but may also be determined using multiple levels / multiple values, such as high / medium / low or first / second / third. In this example, the risk level of the bicycle 2801 is set to medium based on the relative distance D1 or relative speed V1 between the user U1 and the bicycle 2801. The device 1 controls the reading of the bicycle 2801, which is a risk factor, with high priority based on the risk level.

[0245] For example, at time point 1, bicycle 2801 is detected and a voice readout equivalent to a basic notification such as "A bicycle is approaching from the front" is performed. Next, at time point 2, the system determines that bicycle 2801 is a risk factor and sets the risk level to medium based on the fact that the relative speed V1 between user U1 and bicycle 2801 is above a threshold and that the route is close. The system determines that bicycle 2801, which is a risk factor, is a high priority target for voice readout depending on the risk level being medium. Furthermore, the system adds an alert to the voice readout for bicycle 2801 depending on the risk level being medium. The alert may be a voice output such as "Be careful," or a predetermined alert sound. The alert sound may be a sound that repeats at a predetermined interval (e.g., beep beep...) or may generate a vibration.

[0246] Furthermore, when risk factors are read out aloud, the manner of the read out aloud / alert may be changed depending on the level of risk. For example, the higher the level of risk, the louder the volume may be. The tone of the voice / alert sound may be changed depending on the level of risk. The higher the level of risk, the shorter the period of the alert sound or vibration may be.

[0247] Furthermore, this system has a function of providing voice guidance for user U1 to move to a safe space to prevent or avoid danger factors related to risk factors. This will be explained using the same FIG. 28 . The device 1 and server device 2 predict the route k11 of the bicycle 2801 relative to the route k10 of the user U1, and perform normal control until the bicycle 2801 approaches within a certain distance D11 from the user U1 on the route k10. For example, at time point 1, the voice reads, "A bicycle is coming from the front." In normal control, the system basically maintains the position and route k10 of the user U1, and waits in the hope that the bicycle 2801 will change the route k11 taking the user U1 into consideration.

[0248] If the bicycle 2801 enters within a predetermined distance D11 on the route k10 without changing the route k11, the system determines that there is a possibility of contact and performs control related to the guide of the safe space. Based on an analysis of the image from the camera 403 (e.g., the front camera), the system calculates the safe space and the direction of movement to the safe space (in other words, the direction of escape) to avoid contact with the bicycle 2801 coming toward the user U1 on the route k10 of the user U1.

[0249] In the illustrated example, the system first determines whether it is possible to move to the left (-X) or right (+X) from the current position on road 2800. In this example, there is no space to the left due to a wall or the like, and there is an empty space 2810 (in other words, a space with no objects) to the right. The system determines that movement to the left is not possible, but that movement to the right is possible. In addition, at this time, the system also determines the distance from surrounding objects to the empty space 2810 to the right, for example. In particular, the system compares the empty space 2810 with the path k11 of the bicycle 2801 and the paths of other objects, and determines whether those paths intersect with the empty space 2810. As a result of this determination, the system determines that the empty space 2810 is safe, or in other words, sets the empty space 2810 as a safe space.

[0250] The device 1 guides the user U1 to move to the right, using the empty space 2810 to the right (+X) of the user U1 as a safe space. For example, at time point 3, the voice output corresponding to the guidance is "Please move to the right" or "Please move to the right."

[0251] As described above, the safety space is selected as a space that avoids the object, taking into consideration the path of the object approaching the user U1, thereby preventing or avoiding contact between the user U1 and the object, such as the bicycle 2801.

[0252] [Control for Prioritizing Voice Reading of Specific Objects] Figure 29 shows an example of control for prioritizing voice reading of specific objects. In this system, the objects to be read aloud may be set in advance to be limited to specific objects. This setting may be a system or app setting, or may be a user setting. For example, specific objects related to assistance for the visually impaired, traffic safety, accident prevention, etc. may be set as specific objects.

[0253] In the illustrated example, a user U1 is walking forward (+Y) on a sidewalk 2900 with a roadway on the right side. Objects detected ahead of the user U1 include a traffic light 2901, a crosswalk 2902, a pedestrian 2903, a crosswalk 2904, and a road sign 2905. The system assumes that traffic lights, crosswalks, and road signs are set in advance as specific objects. The system prioritizes detecting the specific objects, such as the traffic light 2901 and the crosswalk 2902, from within the camera image, and prioritizes reading out the detected traffic light 2901 and the crosswalk 2902.

[0254] For example, when the device 1 first detects the crosswalk 2902, it reads out "There is a crosswalk on your right." Next, when the device 1 first detects the traffic light 2901 (for example, when the traffic light is green), it reads out "The traffic light is green." The system also detects changes in a specific object and reads out the changes. For example, a change from a green light to a red light is detected as a change in the status of the traffic light 2901. When the traffic light changes, the device 1 reads out the change in the traffic light. For example, it outputs "The traffic light has turned red."

[0255] The system also places importance on traffic safety and calculates the safety and risk when user U1 crosses the crosswalk 2902. For example, when the traffic light 2901 is red, the system may determine that the risk is high and output an alert according to the risk level. For example, the system may output a message saying, "The traffic light has turned red. Please be careful."

[0256] Similarly, other specific objects, such as road signs 2904, can be read out with high priority. In the illustrated example, when the road sign 2904 indicating a road closure for pedestrians is detected, a message such as "There is a road closure sign for pedestrians" is read out with priority. Other examples of specific objects include intersections and tactile paving blocks.

[0257] As described above, by giving priority to specific objects when reading them out loud, efficient support can be achieved according to the usage scenario, such as support for the visually impaired.

[0258] [Control Example Using a 360-Degree Camera] FIG. 30 shows a control example when a 360-degree camera (in other words, a celestial camera) is used as the camera 403. The illustrated image 3001 is a schematic diagram of an image captured by the 360-degree camera. In image 3001, the center point of the circular / ring-shaped area corresponds to the position of user U1, i.e., the position of the 360-degree camera. The circumferential direction of the circular / ring-shaped area corresponds to the front-to-back, left-to-right directions (in other words, azimuth angles). The radial direction of the circular / ring-shaped area corresponds to the up-to-down direction (in other words, elevation and depression angles). Objects can be detected within a 360-degree celestial sphere image corresponding to such a circular / ring-shaped area. For example, if a pedestrian is diagonally ahead and to the left of user U1, the pedestrian is captured as pedestrian image g1 in image 3001.

[0259] For example, user U1 in FIG. 6B may be provided with a 360-degree camera with a vertically upward optical axis on a hat 1D or the like. Alternatively, a 360-degree camera with a forward-facing optical axis on the front of the hat 1D and a 360-degree camera with a backward-facing optical axis on the rear may be provided. Such an image 3001 may be obtained using a single 360-degree camera, or may be obtained by synthesizing such images 3001 using multiple cameras. Alternatively, a panoramic image 3002 such as that shown in the lower part of FIG. 30 may be used.

[0260] [Control Example for Changing the Audio Output Mode Depending on the Camera Direction] Figure 31 shows a control example for changing the audio output mode (e.g., tone) of a text-to-speech voice using the front and rear cameras in the front-to-back direction relative to user U1. For example, suppose there is an object 3101 detected in the space in front of user U1 based on an image captured by the front camera C1 of device 1 (smartphone 1A), and an object 3102 detected in the space behind user U1 based on an image captured by the rear camera C2. Device 1 changes the audio output mode to make it easier to distinguish between the text-to-speech voice of object 3101 based on the image captured by the front camera C1 and the text-to-speech voice of object 3102 based on the image captured by the rear camera C2. For example, the tone and volume of the voice may be changed. In the illustrated example, the text-to-speech voice of the front object 3101 is controlled to be output in a male voice (first tone, relatively low voice), and the text-to-speech voice of the rear object 3102 is controlled to be output in a female voice (second tone, relatively high voice).

[0261] This makes it easier for the user U1 to recognize the front and rear positions. Furthermore, with this control, even if multiple voice readings for multiple objects are performed almost simultaneously, it becomes easier to distinguish between them based on the difference in tone.

[0262] [Control Example Using 3D Audio Output] Figure 32 shows a control example using 3D audio output. In the illustrated example, user U1 is wearing smart glasses 1C and is near an intersection. Assume that there are an object 3201 detected from the front direction (+Y) and an object 3202 detected from the right direction (+X) of user U1. In this control example, device 1 is equipped with an audio output device having a 3D audio output function. This 3D audio output function is a function that sets the source of audio in a 3D space and outputs audio from a speaker or the like so that the audio sounds as if it is coming from that source.

[0263] In the illustrated example, the device 1 determines the position L1 of the object 3201 detected in front of the user U1 and the position L2 of the object 3202 detected to the right of the user U1 based on the distance determination described above. These positions may be approximate. The device 1 then controls the 3D audio output for the voice reading of the object 3201 so that the voice sounds as if it is coming from the front position L1, and controls the 3D audio output for the voice reading of the object 3202 so that the voice sounds as if it is coming from the right position L2. This makes it easier for the user U1 to recognize the direction and position of the object. Furthermore, with this control, even if multiple voice readings of multiple objects are performed almost simultaneously, the user U1 can distinguish between them using the 3D audio output.

[0264] The following reference information is an example of technology related to 3D audio output functions: [Reference information] SRS-RA3000 Active Speaker / Neck Speaker Sony (sony.jp)

[0265] [Effects, etc.] As described above, according to the first embodiment, the image reading function can provide more suitable support to visually impaired and non-visually impaired users. For example, it is possible to improve the safety of visually impaired users when walking and to assist them in using transportation systems and facilities. Even for non-visually impaired users, it is possible to achieve, for example, safety measures and crime prevention on roads at night, and improved safety when driving while "walking while using a smartphone."

[0266] <Embodiment 2> A response output system (in other words, an AI conversation system) according to embodiment 2 will be described with reference to Figure 33 and subsequent figures. The basic configuration of embodiment 2 is the same as or common to embodiment 1, and the following describes the components of embodiment 2 that differ from embodiment 1.

[0267] The following issues can be cited: Conventionally, there is AI (i.e., conversation AI / chat AI, etc.) that generates conversations (i.e., chats, etc.) based on content input by a user to a response input / output interface, but when using conversation AI in situations such as when the user is walking around town, the surrounding circumstances are not taken into consideration. Even when the user captures the surrounding circumstances with a camera and inputs them into the conversation AI, the capturing operation is necessary, which is cumbersome.

[0268] Therefore, in the second embodiment, when a user uses a conversational AI in a situation such as walking around town, multiple images capturing the surrounding environment are automatically captured at various times using multiple cameras on the device and input to the conversational AI. The conversational AI receives the multiple images along with the conversational text input by the user, recognizes the surrounding environment from the multiple images, and generates and outputs a response conversational text that reflects the recognized surrounding environment in the topic of conversation. As a result, in the second embodiment, the surrounding environment is reflected during a conversation with the AI, realizing a favorable conversation without requiring complicated operations.

[0269] In the first embodiment described above, a system that assists a user in walking, etc., based on input of multiple camera images from the device 1 was described. Since the second embodiment also uses input of multiple camera images, many of the components overlap with the first embodiment. The following description will focus on the components added in the second embodiment. The functions of the second embodiment share a common configuration with the first embodiment in that the AI ​​uses time-series captured images from the multiple cameras of the device 1 as input, analyzes them, recognizes the user's surroundings, and generates and outputs response text, etc., according to the surroundings. The functions of the second embodiment differ from the functions of the first embodiment in that the user and the AI ​​converse, and the AI ​​generates conversational sentences using text and images. The AI ​​conversation function described in the second embodiment can be implemented in addition to the text-to-speech function described in the first embodiment, and can also be used in combination. That is, the function of the first embodiment can assist the user in walking by text-to-speech reading of the surroundings, while the function of the second embodiment can enable a conversation between the user and the AI ​​while reflecting the surroundings.

[0270] In the second embodiment, the device 1 has a function for conducting a conversation between the user and the AI ​​(sometimes referred to as an AI conversation function). This conversation may be implemented as a user interface in a specific application. In the second embodiment, the user's device 1 has an interface for conducting a conversation with the AI ​​(sometimes referred to as a conversational AI) using text, audio, and images. The interface includes an interface for inputting and sending text, audio, and images of requests (e.g., instructions, questions, prompts, etc.) to the AI, and an interface for receiving and outputting text, audio, and images of responses (e.g., answers, etc.) from the AI. The conversational AI is implemented, for example, by the AI ​​server 2 (particularly the MLLM), but may also be implemented as a function within the device 1 (particularly the local LLM). Note that the slash ( / ) in "text / audio / image" indicates OR. In other words, this expression indicates one or more of text, audio, and images.

[0271] In the AI ​​conversation function of the second embodiment, the device 1 sends and inputs, as a request, data of multiple images (sometimes referred to as camera images) captured by multiple cameras of the device 1, along with text / audio / images of a conversation based on user input, from the device 1 to the conversation AI of the AI ​​server 2. The conversation AI of the AI ​​server 2 generates text / audio / images of a response to the conversation from the input text / audio / images. At this time, the conversation AI of the AI ​​server 2 recognizes the surrounding circumstances of the user and the device 1 by performing image recognition processing on the multiple input images (camera images). The conversation AI of the AI ​​server 2 then controls the response conversation so that the recognized surrounding circumstances are reflected in the response conversation. In other words, the user's device 1 controls the conversation AI to obtain a response conversation that reflects the surrounding circumstances. For example, the conversation AI detects surrounding objects (in other words, target objects) and reflects the detected surrounding objects as topics in the conversation flow. For example, the conversation AI generates a conversation that uses the surrounding objects as topics and inserts or adds it to the conversation flow. Examples of surrounding objects may include facilities and landmarks defined on a map, or may be objects not defined on a map (e.g., moving objects). This allows the user to have a conversation with the AI ​​that reflects the surrounding situation.

[0272] Furthermore, in the AI ​​conversation function of the second embodiment, the user's face may be included in the camera image. The conversation AI performs image recognition processing on the input camera image to recognize the surrounding situation (surrounding objects other than the user's face, etc.) as well as the user's facial expressions and behavior (e.g., facial orientation, movement, gaze direction, etc.). Facial expressions and emotions have a correspondence relationship. Facial expressions / emotions can be classified using publicly known technology. For example, classifications include joy / anger / sadness / happiness / neutrality.

[0273] The conversation AI then controls the response conversation to reflect the recognized facial expressions and behaviors, along with the recognized surrounding circumstances. In other words, the user's device controls the conversation AI to acquire a response conversation that reflects the facial expressions and behaviors. For example, the conversation AI generates a conversation that focuses on the user's facial expressions and behaviors and inserts or adds it to the conversation flow. Alternatively, the conversation AI may be controlled to change the content of subsequent conversations depending on the user's facial expressions and behaviors. For example, when providing guidance regarding the surrounding circumstances, such as navigating on a map, the conversation AI may control the guide conversation to change or add depending on the user's facial expressions and behaviors. For example, the conversation AI may control the conversation to continue or develop a topic about a surrounding object if the user has a happy expression, or may control the conversation to discontinue the topic or change to another topic if the user has an anxious or sad expression. Furthermore, for example, when the conversation AI is navigating a route to a destination on a map, if the user looks anxious or sad, the conversation AI may be controlled to add guide dialogue to ease the user's anxiety or sadness.

[0274] [System Configuration] Figure 33 shows a system configuration of embodiment 2. The system configuration of embodiment 2 is roughly similar to the system configuration of Figure 1, etc. The system of embodiment 2 (Figure 33) corresponds to an example when applied to an AI chatbot service (in other words, an AI conversation service) on a communication network 9. A business provides the service to each user's device 1 via an AI server 2. Figure 33 differs from Figure 1 in that the device 1 is an AI conversation device compatible with the AI ​​conversation function. Furthermore, the AI ​​server 2 includes, as the AI ​​20, a conversation AI 40 compatible with the AI ​​conversation function. The conversation AI 40 generates a response conversation sentence in response to a conversation sentence input from the device 1 and outputs it to the device 1. This system has an AI conversation function, which is a function of having a conversation with the conversation AI 40 using images captured by multiple cameras equipped on the device 1 of user U1. The configurations of Figures 2 and 3 can also be applied to embodiment 2.

[0275] As in the first embodiment, various devices can be applied as the device 1. The various devices 1 described in Figures 5A, 5B, 6A, and 6B can also be applied to the second embodiment. For example, a smartphone 1A (Figure 6A) can be used, and the multiple cameras can include a front camera C2 (in-camera) and a rear camera C1 (out-camera). User U1 converses with conversation AI 40 through images captured by the front camera C2 (in-camera).

[0276] Conversational sentences created by user U1 inputting text and voice into smartphone 1A are input to conversation AI 40, and multiple images captured by multiple cameras on smartphone 1A are also input to conversation AI 40. The multiple images include an image of the surrounding situation (e.g., the actual view ahead) captured by rear camera C1 (outer camera) and an image of user U1's face captured by front camera C2 (inner camera). During the conversation with the AI, each camera image is also continuously input on the time axis. Note that use of rear camera C1 (outer camera) is not limited to use, and cameras with other shooting directions (e.g., left and right cameras) may also be used.

[0277] The conversation AI 40 generates a response conversation sentence (sometimes referred to as an output conversation sentence) based on the text / audio / images, etc. (sometimes referred to as an input conversation sentence) input from the device 1 of the user U1, and outputs it to the device 1. In doing so, the conversation AI 40 uses input image information from multiple cameras. The conversation AI 40 recognizes the surrounding situation, facial expressions, behavior, etc. from the camera images, and controls the response conversation sentence to reflect this. The device 1 of the user U1 outputs the conversation sentence from the conversation AI 40 in a predetermined format, such as text / audio.

[0278] For example, the conversation AI 40 generates conversational sentences that talk about the recognized surrounding objects and inserts them into the conversation flow. Furthermore, for example, the conversation AI 40 generates conversational sentences that talk about facial expressions and inserts them into the conversation flow. The conversation AI 40 may devise or change the way it responds to conversational sentences by referring to the facial expressions, behavior, etc. of the user U1.

[0279] During a conversation, if the conversation AI 40 recognizes, based on camera images, a surrounding situation or surrounding object, such as a bicycle coming from in front of the user U1, it generates a conversational sentence such as "There's a bicycle coming from ahead" to draw attention to the surrounding object and inserts it into the conversation flow. Furthermore, if the conversation AI 40 recognizes, for example, a famous landmark as a surrounding situation or surrounding object, it generates a conversational sentence that discusses the surrounding object, such as "The building you see in front of you on the right is the famous Roppongi Hills, right?" and inserts it into the conversation flow. In this way, the conversation AI 40 generates conversation topics based on the recognition of camera images.

[0280] Furthermore, the conversation AI 40 may use, in addition to images of surrounding objects captured by the outer camera and images of the face captured by the inner camera, detection information on the posture (vertical / horizontal or tilted state), position, and direction of the device 1, and the line of sight (line of sight, point of gaze, etc.) of the user U1. For example, the conversation AI 40 may generate conversational sentences so as to prioritize the topic of surrounding objects in the direction the device 1 is facing or in the line of sight of the user U1.

[0281] In the second embodiment, not only the multiple camera image input described in the first embodiment, but also voice input by the user or text of the voice recognition result, text of the character input result from key buttons, etc. are used as input information to the conversation AI. The conversation AI can generate text / voice of a response conversation sentence based on the text / voice of the input conversation sentence.

[0282] It should be noted that technology for conducting AI chat in response to input of text information is already known. For example, Japanese Patent No. 6433614 discloses a system that uses log information and chatbot evaluation information in a normal chatbot system to realize chat more efficiently.

[0283] The technical aspects of the AI ​​chat in the second embodiment (especially the conversation AI 40) can be implemented using known AI technologies for chatbots. Alternatively, a dedicated AI for this embodiment may be implemented.

[0284] 34A shows an example of the configuration of device 1 in embodiment 2, which differs from Fig. 7A in that it has AI chat software 347 (in other words, an AI conversation application) in non-volatile memory 303. In addition, AI data 343 includes not only data related to image reading software 342 but also data related to AI chat software 347.

[0285] 34B shows an example of the configuration of server 2 in embodiment 2, which differs from FIG. 7B in that it has AI chat software 447 (in other words, an AI conversation server) in non-volatile memory 403. In addition, AI data 443 includes not only data related to AI software 442 but also data related to AI chat software 447.

[0286] [Processing Flow: AI Conversation Method] The processing flow in the second embodiment is different from that in Fig. 4. The configurations of settings, image acquisition, image analysis, image readout data creation, image readout, etc. shown in Figs. 8 to 32 can be similarly applied to the second embodiment.

[0287] Fig. 35 shows a processing flow corresponding to the method (AI chat method, AI conversation method, etc.) in embodiment 2. The flow in Fig. 35 is executed in the system shown in Fig. 33 etc. That is, the processor 301 of the device 1 and the processor 401 of the server 2 execute processing according to the flow in Fig. 35.

[0288] In step S201, the device 1 of the user U1 turns on (enables) the AI ​​conversation function when an instruction from the user U1 or the state of the device 1 matches a specific condition.

[0289] The actual state of AI20 (conversational AI40) and AI function 20b is that, in the configuration shown in Figures 34A and 34B, processor 301 and processor 401 load AI chat software 347 on device 1 side and AI chat software 447 on server 2 side, either alone or together with image reading software 342 and image reading software 442, and the processing is executed (in other words, as an execution module, instance).

[0290] The AI ​​20 (particularly the conversational AI 40) is composed of a generative AI and a large-scale language model (LLM). The LLM may be a multimodal LLM (MLLM). In this specification, the actual AI may also be referred to as an AI process circuitry.

[0291] In step S202, the device 1 receives a voice input or text information input from the user U1. In the case of a voice input, the device 1 may recognize the voice input from a microphone or the like and convert it into text (characters). The device 1 sends the voice or text of the voice recognition result, or the text of the text information input result to the AI ​​20 or the AI ​​function 20b in response to an instruction (operation input) from the user U1 or the lapse of a predetermined time since the input.

[0292] In step S203, the device 1 captures and acquires multiple images by capturing images from input images of multiple cameras with different shooting ranges / shooting directions at predetermined timings, either at the same timing as step S202 or at similar timings with a predetermined relationship. The AI ​​function 20b imports the captured camera images (multiple images, image information). The image import at this time is performed by controlling the imaging timing of the multiple cameras according to the settings, as shown in the previous embodiment 1 ( FIG. 8 , etc.).

[0293] For example, in a smartphone 1A (FIGS. 6A and 9A), if there is a camera C2 (in-camera) on the front side of the housing that takes pictures of the rearward direction and a camera C1 (out-camera) on the back side of the housing that takes pictures of the forward direction, while a conversation with AI is taking place through an image input including a facial image from the in-camera, an image input including a forward image from the out-camera is also taken at a predetermined timing. As in the first embodiment, the processing ratio between the image input from camera C1 and the image input from camera C2 can be set (e.g., state A in FIG. 9A). The same applies when using front, rear, left, and right cameras.

[0294] The AI ​​20 or the AI ​​function 20b determines and outputs the content of the response conversation based on the input text or voice of the user U1 and the camera image (multiple images). At that time, the AI ​​20 or the AI ​​function 20b uses an image of the surrounding situation captured by the outer camera or an image of the surrounding situation other than the face captured by the inner camera, and controls the response conversation so that the recognized surrounding situation is reflected. The AI ​​20 or the AI ​​function 20b may also analyze the facial expression and behavior of the user U1 from the facial image captured by the inner camera, and control the response conversation so that the analyzed expression and the like are reflected.

[0295] Furthermore, the AI ​​20 or the AI ​​function 20b may detect the orientation (length, width, tilt, etc.), position, direction, facial orientation, movement, line of sight, etc. of the device 1. Based on the detection information, the system may, for example, estimate a surrounding object close to the line of sight of the user U1 and generate a conversational sentence so as to incorporate the surrounding object into the topic of conversation.

[0296] Furthermore, the AI ​​20 or the AI ​​function 20b may understand, based on an analysis of the facial image, the behavior of the user U1, such as whether the user U1 is looking at the display screen on the front side of the device 1, in which direction the user is looking, whether the user's gaze is changing frequently, or whether the user's gaze is steady and calm. When the user U1 holds the device 1 in his / her hand, the movement of the user's hand, the position and posture of the device 1 held in his / her hand, etc., change from time to time. The state of such changes can also be understood based on the sensor (such as the posture detection unit 312) of the device 1 and the contents of the camera image.

[0297] The AI ​​function 20b (AI conversation application, in other words, client software) on the device 1 side may perform pre-processing on the device 1 side, i.e., sub-processing before the AI ​​20 on the server 2 side performs main processing. The AI ​​function 20b may format the text of the conversation, or may perform format conversion processing such as converting voice to text or image to text. The AI ​​function 20b may also extract points or ranges of interest based on the gaze of the user U1 from the camera image, or may perform image processing such as editing the camera image to reflect points of interest.

[0298] In step S204, the device 1 transmits the input conversational text and camera images (multiple images), as well as preprocessing information, to the AI ​​20 on the server 2 side as a communication request (data of a predetermined communication protocol). Note that the device 1 may transmit the captured information (input conversational text and camera images) to the AI ​​20 on the server side without the AI ​​function 20b performing any action. Also, as in embodiment 1 (FIG. 3), the series of processes may be completed only on the device 1 side, in which case the device 1 may perform the processing using the AI ​​function 20b without transmitting information to the server 2 side.

[0299] In step S205, the server 2 of the present system analyzes the input conversation text, etc., obtained from the device 1, the camera image, and, if any, the preprocessing information, using the AI ​​20 (conversation AI 40). In particular, in step S205, the AI ​​20 analyzes / recognizes the surrounding situation from the camera image (e.g., an image from the outer camera), and as a result, obtains text representing the surrounding situation.

[0300] Step S206 may also be added. In step S206, AI20 analyzes / recognizes the user's facial expressions and behavior from a camera image (e.g., an image from the in-camera), and as a result, obtains information representing the expressions and behavior.

[0301] In step S207, AI20 generates one or more candidate response sentences (text / audio / image) from the results of steps S205 and S206, i.e., the analysis results of the input conversation sentence and the camera image (surrounding circumstances, facial expressions, etc.). Note that in step S207, one response sentence may be determined from the beginning without generating candidates. The response sentence (output conversation sentence) is typically text (in other words, a sentence consisting of character strings) generated using text representing the surrounding circumstances and facial expressions, which are the recognition results. As the response conversation sentence, a conversation sentence using audio or images (still images or video) may be generated and output instead of or in addition to text.

[0302] In step S208, the system (e.g., AI20 of server 2) determines the conversation sentence to be output (output conversation sentence) by selecting one from the candidates based on a predetermined judgment (e.g., judgment of the degree of change described in embodiment 1) and timing based on the response sentence candidates that are the result of step S207, and transmits the information to device 1.

[0303] In step S209, the device 1 receives the information on the output conversation sentence from the AI ​​20 determined in step S208, and automatically outputs the output conversation sentence to the user U1 in a predetermined output format. The output may be a voice readout (in other words, a voice output), a text display on a screen, or an image display. The output format is configurable.

[0304] In step S210, device 1 checks whether to turn off the AI ​​conversation function. If it is to be turned off (YES), the flow ends. If it is to remain on (NO), the flow returns to step S201 and repeats the same process.

[0305] As described above, in the second embodiment, in a system having a conversation AI that responds to conversational sentences input by voice or text from user U1, by inputting information from one or more camera images to the conversation AI at an appropriate timing, it is possible to optimize the response by the conversation AI, thereby realizing a suitable conversation that reflects the surrounding circumstances, facial expressions, etc. of user U1.

[0306] [System Configuration Example: Software] FIG. 36A and other figures show a system configuration example, particularly a software configuration example, according to the second embodiment.

[0307] The system of FIG. 36A (in other words, an AI conversation system, etc.) is a system in which a user terminal, which is device 1 of user U1, and a server 2, which is an AI server, are appropriately connected via communication, similar to FIG. 33 . Device 1 is provided with app 351 as a client application for realizing the AI ​​conversation function. Server 2 is provided with AI 352 as server software for realizing the AI ​​conversation function. App 351 also has a conversation client 353 and an image recognition client 354 as components of a program, etc. AI 352 has a conversation AI 355 and an image recognition AI 356 as components of a program, etc.

[0308] 36A, a conversation client 353 creates an input conversation sentence (instructions, prompts) based on user-input text / voice and transmits it as a request a1 (e.g., a packet on a communication network) to the server 2. An image recognition client 354 acquires camera images (multiple images) and transmits data of the camera images (multiple images) a2 to the server 2 at approximately the same timing as the input conversation sentence (request a1). The data of the request a1 and the camera images a2 may be transmitted together at approximately the same timing, or may be transmitted separately at different timings.

[0309] The conversation AI 355 of the server 2 generates a conversation (output conversation) for response from the input conversation of the request a1 based on a conversation generation model. In doing so, the server 2 also uses the image recognition AI 356. The image recognition AI 356 receives the camera image a2 as input, recognizes the surrounding situation or facial expressions and behavior based on image recognition processing, and obtains information a3, such as text describing the surrounding situation or text describing facial expressions and behavior. The conversation AI 355 receives the information a3 along with the input conversation, and generates an output conversation that reflects the surrounding situation or facial expressions as a topic based on this information. The conversation AI 355 transmits the text / audio of the output conversation to the device 1 as a response a4. The conversation client 353 of the device 1 outputs the output conversation to the user U1 in the form of text / audio / images.

[0310] Regarding the series of steps related to the AI ​​conversation function, a typical example is that the user U1 can input a conversation sentence by voice and the AI ​​can output a conversation sentence in response. Another example is that the user U1 can input the text of the conversation sentence on the screen and the text of the conversation sentence in response from the AI ​​can be displayed on the screen.

[0311] 36A, as a modified example, the image recognition AI 356 may be integrated into the conversation AI 355. Also, the conversation client 353 and the image recognition client 354 may be integrated into one.

[0312] The system of FIG. 36B is a variation of the system of FIG. 36A, and its main difference is that it includes an image recognition AI 357 on the device 1 side. The image recognition AI 357 on the device 1 side is implemented using a local LLM. The image recognition AI 357 on the device 1 side uses camera images from the device 1 as input, recognizes the surrounding situation or facial expressions and behavior based on image recognition processing, and obtains information such as text describing the surrounding situation or text describing facial expressions and behavior. The image recognition AI 357 on the device 1 transmits information a3 to the server 2 at approximately the same time as the input conversation sentence (request a1).

[0313] The conversation AI 355 of the server 2 receives the input conversation sentence of the request a1 and information a3 such as the surrounding situation and facial expressions, and generates a conversation sentence (output conversation sentence) that reflects the surrounding situation or facial expressions as a topic based on a conversation generation model. The conversation AI 355 transmits the text / audio of the output conversation sentence as a response a4 to the device 1. The conversation client 353 of the device 1 outputs the output conversation sentence to the user U1 in the form of text / audio / images.

[0314] For the image recognition AIs 356 and 357, in particular, well-known object detection technologies such as CNN (Convolutional Neural Network) can be applied.

[0315] Although there may be a time difference between the processing of the text / audio of the conversation and the processing of the camera image due to computer processing time, these processes may be synchronized as much as possible by setting and controlling the timing of the transmission of various data and information. For example, in Figure 36A, after image recognition processing by image recognition AI 356 is completed, the input conversation sentence of request a1 and information a3 may be input to conversation AI 355 in a timed manner to generate a response. In Figure 36B, after image recognition processing by image recognition AI 357 is completed, the conversation sentence request a1 and information a3 may be transmitted in a timed manner.

[0316] As another modification, an image recognition AI for recognizing the surrounding situation (surrounding objects, etc.) and an image recognition AI for recognizing facial expressions and behavior may be provided separately. As another modification, an image recognition AI may be provided for each camera corresponding to multiple cameras of device 1. For example, when using two cameras, an in-camera and an out-camera of smartphone 1A, two separate image recognition AIs may be provided: one for the in-camera image and one for the out-camera image.

[0317] The system of FIG. 36C is a variation of FIG. 36A , with the main difference being that it includes a conversation AI 358 on the device 1 side and an image recognition AI 356 on the server 2 side. The conversation AI 358 on the device 1 side is implemented using a local LLM. The image recognition client 354 on device 1 sends a camera image a2 to the server 2. The image recognition AI 356 on the server 2 obtains information a3 representing the surrounding circumstances, facial expressions, etc. from the camera image a2 based on image recognition processing. The image recognition AI 356 on the server 2 sends the information a3 to the device 1 as a response. The image recognition client 354 transfers the information a3 to the conversation AI 358. The conversation AI 358 receives input conversational text entered by the user and the information a3, and generates conversational text (output conversational text) that reflects the surrounding circumstances, facial expressions, etc. as topics based on a conversation generation model. The conversation AI 358 outputs the output conversational text to user U1 in the form of text, audio, and images.

[0318] The system of Fig. 36D is a variation of Fig. 36A, and has a configuration integrated into device 1 similar to Fig. 3, with the main difference being that device 1 is provided with conversation AI 358 and image recognition AI 357. The conversation AI and image recognition AI of embodiment 2 are not required on the server 2 side. When used in conjunction with the text-to-speech function of embodiment 1, the server 2 side only needs to implement the text-to-speech function.

[0319] Although not shown in Fig. 36A etc., the server 2 may be equipped with a function related to the voice reading in the first embodiment. For example, the AI ​​20 in the first embodiment (AI that recognizes an object from an input image, converts it into text, and outputs it) and each AI in the second embodiment (conversation AI and image recognition AI) may be integrated and implemented.

[0320] [AI Conversation Function: Configuration Example] Figure 37 shows a more detailed configuration example of the AI ​​conversation function based on Figure 36A. In the system of Figure 37, the server 2 has a conversation AI 355, which includes an image recognition AI 356. The device 1 has, as functional units, a setting unit 371, an image acquisition unit 372, a conversation input unit 373, a conversation output unit 374, multiple cameras 375, an input device 376, an output device 377, and the like. Each functional unit is realized by processing by a processor or by implementing a circuit.

[0321] The setting unit 371 is a part for performing system settings and user settings related to the AI ​​conversation function, etc., and also provides a user interface for this purpose. The multiple cameras 375 are cameras with different imaging ranges / imaging directions, such as the in-camera and out-camera mentioned above. From the multiple cameras 375, multiple images b1 (corresponding to the camera input image 801 in FIG. 8 ) including images from each camera are obtained.

[0322] The image acquisition unit 372 is a part that acquires capture images (corresponding to 801 in Figure 8) at predetermined timings based on multiple images b1 from multiple cameras 375, and transfers them to the conversation input unit 373 as multiple images b2 (image information).

[0323] The conversation input unit 373 acquires text / audio / image b3 (user input information) input by the user through the input device 376 and creates an input conversation sentence (instructions, prompts, etc.) to be input to the conversation AI 355 based on the user input information. The conversation input unit 373 transmits and inputs the input conversation sentence and multiple images (image information) to the conversation AI 355. When the user input format is voice, the conversation input unit 373 acquires voice data through a microphone, performs voice recognition processing on the voice data to convert it into text, and creates an input conversation sentence in text format. At this time, the conversation input unit 373 also creates a request b4 to be sent to the server 2 by attaching multiple images b2 to the input conversation sentence. The request b4 in FIG. 37 corresponds to the request a1 and camera image a2 in FIG. 36A combined into one. The format for attaching multiple images b2 to the input conversation sentence (the format of the request b4) is not limited, but an example is shown in FIG. 38.

[0324] The conversation input unit 373 sends request b4 to the server 2. The server 2 receives request b4, and the conversation AI 355 extracts the text of the input conversation sentence and multiple images from request b4. The conversation AI 355 analyzes the input conversation sentence and multiple images to generate text / audio / images that will serve as a response sentence (output conversation sentence) for the conversation. The conversation AI 355 recognizes the surrounding situation, facial expressions, etc. from the multiple images using the image recognition AI 356, and obtains text that represents the surrounding situation, facial expressions, etc. The conversation AI 355 receives as input the text of the input conversation sentence and the text representing the surrounding situation, facial expressions, etc., and generates text of a response conversation sentence (output conversation sentence) that reflects the surrounding situation, facial expressions, etc. in the topic. Note that in this case, the conversation AI 355 may generate a first output conversation sentence for the input conversation sentence, generate a second conversation sentence for the text representing the surrounding situation, facial expressions, etc., and output them as a set. Server 2 (conversation AI 355) sends the output conversation sentence to Device 1 as response b5.

[0325] The device 1 receives the response b5, and the conversation output unit 374 extracts the text of the output conversation sentence from the response b5 and outputs the output conversation sentence in a predetermined format from the output device 377. If the output format is audio, the conversation output unit 374 creates audio data from the text of the output conversation sentence by speech synthesis processing, and outputs the audio from the speaker (output device 377) based on the audio data (b6).

[0326] When recognizing the surrounding situation from camera images, the present system may use detection information such as the position, orientation, and posture of the device 1, as well as information from a map database, etc. In this way, the accuracy of recognizing the surrounding situation can be improved. When the AI ​​on the server 2 performs image recognition processing, related information such as detection information may be transmitted from the device 1 along with the camera images.

[0327] [Request Data] Figure 38 shows an overview of an example configuration of request b4. Request b4 includes text data 3801 of an input conversation (prompt) and image data 3802 of multiple images (camera images). The image data 3802 of the multiple images includes image data (first camera image data to Nth camera image data) for each camera (corresponding shooting direction). For example, in the case of two front and rear cameras (C1, C2) as shown in Figure 6A, the first camera image data is image data from the outer camera (C1) that captures the front, and the second camera image data is image data from the inner camera (C2) that captures the rear. Each camera image data also includes information such as the shooting direction and the shooting time. Such data of request b4 is transmitted and received as data, such as packets, according to a communication protocol.

[0328] Furthermore, the illustrated example of request b4 illustrates a case in which both input conversational text and camera images are generated at roughly the same time. However, there are also cases in which only input conversational text or only camera images are generated at a certain time. In such cases, the data content of request b4 is either text data 3801 or image data 3802. Furthermore, the illustrated example of request b4 illustrates a case in which multiple images captured by multiple cameras at roughly the same time are transmitted as a set, but this is not limited to this. As in the aforementioned FIG. 9A , if multiple cameras capture images at different times with a predetermined cycle or ratio, a corresponding data set can be created. For example, in the case of control 1 in FIG. 9A , the data set for image data 3802 may be composed of one camera image data per time point, and these may be transmitted sequentially. As another example, camera image data from multiple nearby time points may be combined into a single request b4. For example, image data 3802 may contain four camera image data sets from time points 1 to 4 of control 1.

[0329] [Map and Surrounding Situation] FIG. 39 illustrates an example of the surrounding situation of user U1 and device 1 when using the AI ​​conversation function of embodiment 2. It is a schematic diagram showing a bird's-eye view of the current positions of user U1 and device 1, roads, surrounding objects, etc. on a map. The map has east-west, north-south, and latitude and longitude (not shown), and corresponds to a spatial coordinate system (X, Y, Z). In this example, the Y direction is north. User U1 (similar to FIG. 6A ) holding a smartphone 1A as device 1 is at position L1 (current position) on the map. Position L1 is expressed by position coordinates (X1, Y1, Z1) and corresponding latitude, longitude, altitude, etc. At position L1, user U1 and smartphone 1A are facing north (+Y direction). Direction DC1 is the shooting direction of the smartphone 1A's outer camera (camera C1), and direction DC2 is the shooting direction of the smartphone 1A's inner camera (camera C2). A route 3901 is an example of a travel route from a current position L1 of a user U1 to a destination 3902 (corresponding to a surrounding object SO3). In this example, the user U1 has a destination 3902 and is traveling toward the destination 3902, but in another example, there may be no particular destination.

[0330] Examples of surrounding objects include surrounding objects SO1 to SO8. In this example, the surrounding objects are buildings, facilities, landmarks, etc. that exist on a map. The surrounding objects are not limited to these. The surrounding objects handled in the second embodiment are any objects that can be recognized and detected from camera images. Note that the surrounding objects may be pre-registered in a service / database such as a map or navigation system, or may not be pre-registered. In the case of an object such as a facility that is registered in a service such as a map, information about the facility can also be provided by the service, but the function of the second embodiment differs in that the surrounding objects of the facility, etc. are reflected in the conversation with the AI. Furthermore, in the second embodiment, surrounding objects that are not registered in a map, etc. can also be reflected in the conversation with the AI.

[0331] Furthermore, gaze direction 3903 is an example of the gaze direction from user U1 at position L1, and is a direction diagonally forward and to the right of north (+Y direction) at approximately 30 degrees. In the case of gaze direction 3903, for example, surrounding objects SO1 and SO2 are visible in the user U1's field of view near gaze direction 3903 (in other words, these surrounding objects are captured in the corresponding camera image), but surrounding object SO4 and the like are not visible. Note that with regard to the field of view and gaze direction, it does not matter whether user U1 actually recognizes surrounding objects SO1 and the like.

[0332] [Multiple Images (Camera Images)] Figure 40 shows an example of multiple images (camera images) captured by device 1 using multiple cameras and transmitted to server 2. This example shows two images captured by two front and rear cameras (C1, C2) as shown in Figure 6A. Image 4001 in (A) is an image captured by the outer camera (camera C1) capturing an image in front of user U1, and image 4002 in (B) is an image captured by the inner camera (camera C2) capturing an image behind user U1 at approximately the same time. The example in Figure 40 corresponds to the example situation in Figure 39.

[0333] In the forward image 4001, for example, a surrounding object SO1 and a surrounding object SO2 are shown, and these surrounding objects can be detected based on image recognition. For example, a store, which is the surrounding object SO1, can be detected from the characters on a signboard. For example, a building, which is the surrounding object SO2, can be detected from the shape of a tower. Other examples of surrounding objects that may be recognized include trees. An x mark 4003 is an example of a gaze point 4003 corresponding to the line of sight direction.

[0334] In the rear image 4002, no particularly conspicuous surrounding objects are captured and detected. Since the user's face (corresponding face area 4004) is captured in the rear image 4002, the face, facial expression, behavior, etc. can be determined and detected based on image recognition. As behavior, the direction and movement of the face can be determined. Furthermore, the gaze direction may be estimated from the state of the eyes. Facial expressions are classified into a plurality of types (emotion types), such as joy, anger, sadness, happiness, and neutral, based on known emotion recognition technology, for example.

[0335] In the function of the second embodiment, for example, peripheral objects in a forward image 4001, such as peripheral objects SO1 and SO2, are recognized and detected and reflected as topics in the AI ​​conversation. In particular, in the case of control using gaze direction, among all peripheral objects in the image, peripheral objects closer to the gaze direction (point of gaze) are given priority. For example, in image 4001, the gaze direction (corresponding point of gaze 4003) is to the right, so even if a peripheral object is detected in an area to the left, the peripheral objects SO1 and SO2, which are to the right, are used preferentially. The distance between the point of gaze 4003 and the peripheral objects (representative position coordinates) may also be determined. An area may be taken with a predetermined radius centered on the point of gaze 4003, and it may be determined whether any peripheral objects are included within that area.

[0336] [Specific Example (1)] A specific example using the AI ​​conversation function of embodiment 2 will be described. Figure 41 shows the situation of user U1 on a map corresponding to this specific example. Figures 42 and 43 show a specific example of the flow of a conversation with the conversation AI corresponding to the situation in Figure 41.

[0337] In the specific example of FIG. 41 , user U1 uses a map app on smartphone 1A to receive navigation (in other words, guidance, etc.) of a route to a destination. This map app is used in conjunction with or integrated with a conversation AI. For example, in FIG. 36A , conversation AI 355 can be considered as a conversation AI integrated with a map service, and conversation client 353 can be considered as a conversation client integrated with a web browser, etc., that receives the map service, etc. The conversation AI navigates the route to the destination through a conversation with user U1. During this navigation conversation, the conversation AI reflects surrounding objects, etc., recognized from camera images as topics of conversation.

[0338] In Fig. 41, initially, user U1 is at position p0 as his current position (present location). User U1's destination 4101 is facility "XXX" on the map, and the surrounding object recognized from the image is SOX. Positions p1 and the like are examples of changing current positions. Route 4110 is a travel route from position p0 to destination 4101, and is an example of a route recommended by navigation.

[0339] In FIG. 42 , first, user U1 launches an app on smartphone 1A (conversation client 353 in FIG. 36A ) and vocally inputs, for example, "Hello. I want to go to 'XXX'." (This is referred to as input conversation sentence 4201.) Smartphone 1A transmits this input conversation sentence 4201 along with a camera image to server 2. Server 2 (conversation AI 355 in FIG. 36A , particularly the map application) generates a conversation sentence in response to input conversation sentence 4201, such as, "Hello, user U1. A route from your current location to destination 'XXX' has been set. A map will be displayed. The destination is a 15-minute walk away." This is then sent back to smartphone 1A (output conversation sentence 4202). The route here corresponds to route 4110 in FIG. 41 . Smartphone 1A outputs output conversation sentence 4202 vocally and displays a map on the screen.

[0340] User U1 follows the map navigation to head to destination 4101. User U1 proceeds along the presented route (path 4110). The conversation AI outputs a navigation message such as "Please go straight ahead (north) for a while" (output response sentence 4203). User U1 proceeds along the navigation message.

[0341] For example, user U1 reaches position p1 just before traffic light A (intersection A). When user U1 is unsure where to turn left, he or she utters, for example, "I wonder if I should turn left at the next traffic light?" (input conversation sentence 4204). At this time, the user's facial expression is normal. In response to input conversation sentence 4204, the conversation AI outputs a navigation message such as, for example, "No. Just follow the road until the next traffic light comes up" (output response sentence 4205). User U1 continues to go straight according to the navigation and reaches, for example, position p2.

[0342] Assume that user U1 continues driving straight ahead but is still unsure where to turn left. He / she is silent and has an anxious expression. A camera image capturing his / her expression at that time is sent to server 2 (silence, camera image 4206). Based on the camera image, the conversation AI recognizes and classifies user U1's expression as sadness, and generates and outputs a navigation response to address the sadness, such as "You're almost there. You can see the traffic light for turning left." (output response statement 4207). User U1 turns left at traffic light B (intersection B) according to the navigation (silence, camera image 4208). At this time, user U1 is silent, for example, and has a neutral expression.

[0343] Furthermore, from the fluctuation of the camera image at this time, the system can also recognize the act of turning left. At this time, for example, a pedestrian 4105 at the left turn destination is recognized and detected from the camera image as a type of surrounding object (moving body). In this case, the presence of the pedestrian 4105 can be reflected in the conversation (navigation). The conversation AI may provide navigation such as, "Proceed toward the pedestrian at that traffic light."

[0344] Next, the conversation AI outputs a navigation such as "Please go straight to 'YYY'" (output conversation sentence 4209). If the conversation AI determines that a facility such as "YYY" is on the route based on map data or recognition of surrounding objects, it can provide such navigation. User U1 goes straight according to the navigation and arrives at position p3, for example. User U1 utters, for example, "I can see 'YYY'" (input conversation sentence 4210). The user's facial expression at this time is normal.

[0345] Based on the camera image, the conversation AI detects, for example, a bicycle approaching from behind as a surrounding object of user U1 (similar to embodiment 1). To alert user U1, the conversation AI generates and outputs a response such as, for example, "A bicycle is approaching from behind. Please be careful" (output conversation sentence 4211).

[0346] User U1 arrives at position p4 near "YYY" just before traffic light C (intersection C) where he should turn right. The conversation AI outputs a navigation message such as "Turn right at the intersection with "YYY"" (output conversation sentence 4212). User U1 turns right at intersection C according to the navigation message (silence, camera image 4213). At this time, the user is silent and has a normal expression.

[0347] Figure 43 is a continuation of Figure 42. User U1 reaches position p5, for example. The conversation AI recognizes stone monument 4103 as a surrounding object based on camera images. The conversation AI outputs a response that reflects the recognized stone monument 4103 as a topic, such as "The stone monument you see on the left is engraved with a famous song by Mr. B" (output conversation sentence 4214). In response to output conversation sentence 4214, user U1 utters, for example, "That's true" (input conversation sentence 4215). The expression at this time is one of joy.

[0348] Furthermore, depending on the facial expression of the user U1 at the time of the reaction (input conversation sentence 4215), the conversation AI may develop or discontinue the topic regarding the stone monument 4103. For example, if the facial expression is one of joy, the conversation AI may output an output conversation sentence that develops the same topic.

[0349] The conversation AI continues navigating the route. User U1 reaches position p6. The conversation AI recognizes signal D (intersection D) where to turn right based on map data or recognition of surrounding objects. The conversation AI recognizes, for example, the church 4104 at signal D where to turn right. The conversation AI outputs navigation such as, "You can see a church with a cross diagonally ahead to your right. Turn right at the traffic light just before it." (output conversation sentence 4216). User U1 utters, for example, "OK" and turns right at signal D (input conversation sentence 4217). The user's facial expression at this time is normal.

[0350] User U1 continues straight along the route after turning right and arrives at position p7. From the camera image, the conversation AI recognizes facility "XXX," which is destination 4101, as a surrounding object SOX. For example, suppose that a tower, which is part of facility "XXX," is detected diagonally forward and to the left from the camera image. The conversation AI reflects the surrounding object and outputs navigation such as, "I can see the tower of destination "XXX" ahead and to the left." (output conversation sentence 4218). In response to output conversation sentence 4218, user U1 utters, for example, "Is that it?" (input conversation sentence 4219). At this time, user U1's line of sight is diagonally forward and to the left, which coincides with the direction in which facility "XXX" is located.

[0351] Based on the camera image and gaze direction, the conversation AI determines that user U1 is looking toward facility "XXX" and outputs navigation such as, "Yes. You will arrive at the entrance in 20 meters." (output conversation sentence 4220). User U1 follows the navigation and arrives at destination 4101 "XXX." User U1 utters, for example, "We've arrived!" (input conversation sentence 4221). The expression at this time is one of joy. Upon arriving at destination 4101, the conversation AI outputs, for example, "We've arrived. Route guidance will end." (output conversation sentence 4222), and ends route navigation.

[0352] According to the above specific example, by reflecting the surrounding circumstances of the user U1 in the AI ​​conversation, more suitable conversations, such as route navigation and guidance, can be realized.

[0353] [Specific Example (2)] Another specific example will be described. FIG. 44 shows an example of the flow of a conversation with an AI in another situation. The situation for user U1 is that user U1 has just arrived at Tokyo Station and has an appointment to meet someone, but there is still time before the appointment, and the user U1 does not have a specific purpose. User U1 inputs, for example, "Hello. I'm at Tokyo Station today" into the app on device 1 (input conversation sentence 4401). Based on the input conversation sentence 4401 and the camera image, the conversation AI recognizes and estimates that user U1 is at Tokyo Station, particularly at the Yaesu Central Exit of Tokyo Station. The conversation AI responds, for example, "Hello. It appears you are at the Yaesu Central Exit of Tokyo Station. Is there anything I can help you with?" (output conversation sentence 4402), reflecting the recognition result (particularly "Yaesu Central Exit").

[0354] In response to output conversation sentence 4402, user U1 utters, for example, "I have plans to meet someone at "AAA," but I still have an hour left." (input conversation sentence 4403). From input conversation sentence 4403 (information "AAA"), the conversation AI can determine that its estimation that user U1 is at Tokyo Station (particularly the Yaesu Central Exit) is correct. Furthermore, from input conversation sentence 4403, the conversation AI can determine that user U1's interests include the destination "AAA" on the map (e.g., a specific facility) and plans for one hour later.

[0355] The conversation AI generates a topic related to what it has understood in response to the input conversation sentence 4403, and outputs a response such as, "To get to "AAA," it's quicker to go to the left and pass through "BBB." There are also souvenir shops on the left." (output conversation sentence 4404). In this example, the conversation AI suggests a route to the destination "AAA" and facilities along the route.

[0356] In response to output conversation 4404, user U1 utters, for example, "Maybe I should buy souvenirs on the way home. Let's have something light to eat." (input conversation 4405). In this example, user U1 reacts negatively to the topic (souvenirs) presented by the conversation AI, and instead responds with an interest in food. From this reaction (input conversation 4405), the conversation AI can determine that user U1 is negatively interested in souvenirs and positively interested in food.

[0357] Assume that user U1 is walking in a certain direction in Tokyo Station after input conversation sentence 4405. The conversation AI recognizes surrounding objects from camera images at that time to recognize the direction and location of user U1. The conversation AI generates a topic based on the recognized direction and location. For example, the conversation AI outputs a response such as, "There's a newly opened "DDD" in the "CCC" area where you're heading now. You like sweets, don't you? Their cakes (¥1,000) are apparently popular." (output conversation sentence 4406). In this example, the conversation AI extracts a restaurant facility "DDD" located in the direction of user U1's travel as a candidate and suggests a meal based on information about the facility (which may be map data or web search information) and user U1's attributes and preferences.

[0358] The conversational AI and app may also generate links, such as URLs, from words in the conversation. Users can obtain detailed information about a word by specifying the link (either by clicking on the displayed information or by voice input) in the conversation. For example, the link for "DDD" or the link for "cake" can provide detailed information about the store or food.

[0359] User U1 agrees with output conversation sentence 4406, for example, and utters, "That sounds good. I'll go and check it out" (input conversation sentence 4407).

[0360] As in the above specific example, conversation with the AI ​​can be realized by providing topics based on the user U1's surroundings and changing the topic depending on the user U1's reaction.

[0361] [Example of Conversation Screen] FIG. 45 shows, as a supplementary illustration, an example of a GUI for a conversation with an AI on the front screen of the smartphone 1A. The example in FIG. 45 is an example in which the text of the conversation (including the text of the voice recognition result) is displayed. In FIG. 45, a touch panel screen 4500 has an in-camera (camera C2), and as an interface for "AI chat," there is a display field 4501 for the AI ​​character's response conversation text and a display field 4502 for the user's input conversation text. In addition, a map 4503 for when using a map application (navigation function) is displayed at the bottom. Note that the map 4503 may be displayed in a separate window.

[0362] [Method of Reflecting Facial Expressions / Emotions] Any of the following methods may be used to reflect facial expressions / emotions in AI conversations. (1) The AI ​​may change the topic of conversation itself depending on the facial expression / emotional state. (2) The AI ​​may change the way words are used in conversation sentences depending on the facial expression / emotional state, without changing the topic of conversation.

[0363] As described above, according to the second embodiment, even if the user U1 does not perform an image capturing operation, it is possible to realize a conversation with the AI ​​that reflects the surrounding circumstances of the user U1 and the device 1. Furthermore, even if the user does not input a conversation sentence, the AI ​​side can automatically generate a conversation sentence in response to the input of a camera image.

[0364] [Modification of Second Embodiment] The following modification of the second embodiment is also possible.

[0365] When both input conversational text and camera images are input at approximately the same time, the response / input / output processing of one conversation AI (e.g., 355 in FIG. 36A ) and the response / input / output processing of the image recognition AI (e.g., 356 in FIG. 36A ) may be performed at a predetermined ratio, rather than at a 1:1 ratio. For example, if the processing load of the image recognition AI 356 in FIG. 36A recognizing the surrounding situation from an image is higher than the processing load of the conversation generation by the conversation AI 355, the ratio of the conversation AI processing to the image recognition AI processing may be set to, for example, 2:1 to reduce the overall load. Conversely, if the processing load of the image recognition AI 356 in FIG. 36A recognizing the surrounding situation from an image is lower than the processing load of the conversation generation by the conversation AI 355, the ratio of the conversation AI processing to the image recognition AI processing may be set to, for example, 1:2 to reduce the overall load.

[0366] There are several patterns of surrounding conditions recognized from camera images. For example, there are cases where fixed surrounding objects are detected in a space such as a facility, and cases where objects that require caution when walking (moving objects such as cars, bicycles, and other people) are detected as described in embodiment 1. Priorities / priorities may be set among these surrounding objects. Objects that require caution are set with a relatively high priority. When the conversation AI detects an object that requires caution as a surrounding object during a conversation with a user, it inserts a conversational sentence indicating that caution should be exercised over normal conversational sentences (e.g., a topic about the facility).

[0367] Additionally, surrounding objects on the map may be categorized into types / categories, and priorities may be set among the types / categories of surrounding objects according to the user's level of interest, and similar control may be exercised over the output of dialogue. Examples of types / categories include history, nature, food, shopping, etc. The user may be able to set their level of interest.

[0368] When the text-to-speech function of the first embodiment (i.e., the function of reading out surrounding objects and situations) and the AI ​​conversation function of the second embodiment (i.e., the function of reflecting the surrounding situation in the conversation as a topic) are used in combination, a priority or processing ratio may be set between these functions. For example, a mode that prioritizes text-to-speech and a mode that prioritizes AI conversation may be provided. For example, in the mode that prioritizes text-to-speech, text-to-speech continues automatically when no conversational text is input from the user. When a conversational text is input from the user, the conversation automatically shifts to a conversation with the AI, and text-to-speech is suppressed.

[0369] In the second embodiment, as in the first embodiment, a priority / order of precedence may be set for the multiple cameras of the device 1. That is, the contents of surrounding objects, etc. recognized from the image of a camera with a higher priority will be preferentially reflected in the conversation of the conversational AI.

[0370] In the second embodiment, similarly to the first embodiment, control may be applied according to whether communication between the device 1 and the server 2 is possible or not.

[0371] As a variant, it is also possible to use an image from only one camera equipped in the device 1. For example, only the camera C2 (in-camera) of the smartphone 1A may be used. The image taken by the in-camera shows the user's face and the surrounding situation (surrounding objects) behind it. The server 2 (conversation AI) recognizes at least one of the facial expression and behavior and the surrounding objects from the camera image. The conversation AI reflects the recognized content in the response conversation.

[0372] As a modified example, in a configuration as shown in FIG. 36B , device 1 recognizes surrounding objects from camera images using image recognition AI 357 and obtains text representing the surrounding objects. Then, conversation client 353 may add a prompt with text representing the surrounding objects as an attached prompt to the prompt of the input conversational sentence. The request data sent from device 1 to server 2 may be, for example, a modified request b4' shown at the bottom of FIG. 38 , which includes an attached prompt 3803 in addition to a prompt 3801 of the input conversational sentence. The conversation client may create a prompt that combines the prompt 3801 of the input conversational sentence and the attached prompt 3803. The conversation AI on the server 2 generates a response conversational sentence from request b4' including such a prompt.

[0373] As a variant example, the conversation AI may generate, as candidates for the input conversation, an output conversation (first conversation) when the surrounding objects recognized from the camera image are not reflected as a topic, and an output conversation (second conversation) when the surrounding objects recognized from the camera image are reflected as a topic. The conversation AI selects and outputs an output conversation from the candidates based on a predetermined judgment / evaluation. The flow of the conversation changes depending on the selection. Facial expressions may be used as the predetermined judgment / evaluation. For example, if the facial expression is happy / entertaining, the conversation on the topic of the surrounding objects may be developed, and if the facial expression is angry / sad, the conversation on the topic of the surrounding objects may be stopped or switched to another topic.

[0374] Although the embodiments of the present disclosure have been specifically described above, they are not limited to the above-described embodiments and can be modified in various ways without departing from the spirit of the present disclosure. Except for essential components, components can be added, deleted, or replaced in each embodiment. Unless otherwise specified, each component may be singular or plural. A combination of each embodiment and its variations is also possible.

[0375] 1...Device (image reading device), 2...Server device, 3...Cloud computing system, 9...Communication network, 20...AI, U1...User

Claims

1. An image reading system comprising a device carried or worn by a user, which reads out aloud objects from images captured by a camera, wherein the device automatically and repeatedly captures images from the camera's video at predetermined times, analyzes the captured images to obtain information including text representing the objects in the images, determines the objects and text to be read out based on the obtained information at a predetermined judgment, and automatically and repeatedly reads out the text representing the determined objects from the device at predetermined times.

2. An image reading system according to claim 1, wherein the repeated image capture at the predetermined timing is continuously performed at a predetermined first time interval even without any operation by the user, and the repeated voice reading at the predetermined timing is continuously performed at a predetermined second time interval even without any operation by the user.

3. An image reading system according to claim 1, which determines changes in the object between multiple captured images on a time axis, and determines the object to be read aloud so as to give priority to objects with relatively large changes.

4. An image reading system as claimed in claim 1, wherein the system determines the relative distance or speed between the device and an object in the image based on analysis of the image, and determines the object to be read aloud so as to give priority to objects whose relative distance is within a threshold or whose relative speed is above a threshold.

5. An image reading system as claimed in claim 1, which determines the possibility of contact between the user and the object based on analysis of the image and taking into account the direction of movement of the object, determines a direction in which to move the user to an empty space where the possibility of contact with the object is low in order to avoid said contact, and guides the user to move in a direction that will move the user to said empty space.

6. An image reading system as claimed in claim 4, wherein the smaller the relative distance or the greater the relative speed, the closer the object is to the user, and for an object that is closer, the volume of the voice reading is increased, or the tone is changed, or an alert sound is added, depending on the closer the object is to the user.

7. An image reading system according to claim 1, wherein a specific object is set in advance as the object to be read aloud, and the object to be read aloud is determined so as to give priority to the specific object, and the specific object includes a crosswalk, a traffic light, and a road sign.

8. An image reading system as claimed in claim 7, wherein the state of the signal of the traffic light is determined as the specific object based on analysis of the image, and when it is determined that the signal has changed, a voice reading representing the change in the signal is performed from the device.

9. An image reading system as claimed in claim 1, wherein when the object is repeatedly read aloud at a predetermined timing, if there is no significant change in the same object on the time axis, the image reading system is controlled so that the same object is not read aloud.

10. An image reading system as claimed in claim 1, wherein when the object is repeatedly read aloud at a predetermined timing, if there is no significant change in the same object on the time axis, the system is controlled so that a reading aloud indicating that there is no change in the same object is performed.

11. An image reading system as claimed in claim 1, wherein, when the user inputs a reading instruction into the device, the following is executed as a priority in accordance with the reading instruction, apart from the automatic image capture and aloud reading at the predetermined timing; the device captures an image from the video of the camera, analyzes the captured image to obtain information including text representing an object in the image, determines the object and text to be read aloud based on the obtained information, and reads aloud the text representing the determined object from the device; and when executing the priority, the device always reads aloud the text representing the object even if there is no significant change in the object.

12. An image reading system as claimed in claim 1, wherein when reading aloud automatically at a predetermined timing or when reading aloud when the user inputs a reading instruction into the device, in the case of a second or subsequent reading aloud of the same object, the system controls the reading aloud to use text with a different expression than the text representing the object when reading aloud the first time.

13. An image reading system according to claim 1, wherein the camera of the device has a camera that can capture 360-degree images, and images of the front, back, left, and right are obtained from the image captured by the camera that captures 360 degrees, based on the orientation of the user.

14. An image reading system according to claim 1, wherein the camera of the device has a plurality of cameras with different shooting directions, and the device automatically captures images repeatedly from the images of the plurality of cameras at predetermined timings.

15. An image reading system according to claim 14, wherein the plurality of cameras include a front camera that captures images in front of the user and a rear camera that captures images behind the user, based on the user's orientation.

16. An image reading system as described in claim 14, wherein the plurality of cameras include a front camera that captures images in front, a rear camera that captures images behind, a left camera that captures images to the left, and a right camera that captures images to the right, based on the orientation of the user.

17. An image reading system according to claim 1, wherein the camera of the device has a plurality of cameras with different shooting directions, and a predetermined timing for capturing the image and a predetermined timing for performing voice reading based on the image from the camera are set for each of the plurality of cameras.

18. An image reading system as claimed in claim 17, wherein a first cycle is set as the predetermined timing for the capture and voice reading in a first camera among the plurality of cameras, and a second cycle longer than the first cycle is set as the predetermined timing for the capture and voice reading in a second camera, and when the object or a significant change in the object is detected from the image of the second camera, the second cycle is set in the first camera and the first cycle is set in the second camera.

19. An image reading system as claimed in claim 15, wherein a first cycle is set for the front camera as the predetermined timing for the capture and voice reading, and a second cycle the same as the first cycle is set for the rear camera as the predetermined timing for the capture and voice reading, and the image reading system is controlled so that the capture and voice reading by the front camera and the capture and voice reading by the rear camera are processed at alternating timings.

20. An image reading system as claimed in claim 1, wherein the camera of the device has a plurality of cameras with different shooting directions, the device automatically captures images repeatedly from the images of the plurality of cameras at predetermined timings, determines changes in the object between the plurality of captured images on the time axis, and determines the object to be read aloud so as to give priority to objects with relatively large changes, and sets a threshold for determining the magnitude of the change for each of the plurality of cameras.

21. An image reading system according to claim 1, wherein the camera of the device has a plurality of cameras with different shooting directions, the device automatically captures images repeatedly from the images of the plurality of cameras at predetermined timings, and a different tone of voice is used for each of the plurality of cameras when performing the voice reading based on the image of that camera.

22. An image reading system according to claim 1, wherein the device has a three-dimensional audio output device for performing the audio reading, and when the object is read aloud, three-dimensional audio is output from the three-dimensional audio output device so that the audio sounds to the user as if it is coming from the position of the object in three-dimensional space.

23. The image reading system according to claim 1, wherein the device is at least one of a smartphone, a smart watch, a tablet terminal, smart glasses, and a head-up display.

24. An image reading system according to claim 1, further comprising a server device connected for communication with the device, the server device performing a process of analyzing the captured image and obtaining information including text representing objects in the image.

25. An image reading system as described in claim 1, wherein the device has the server device perform the analysis when communication with the server device is possible, and when communication with the server device is not possible, performs a simplified analysis within the device as the analysis, and the simplified analysis is an analysis performed at a longer cycle than the cycle of the analysis performed by the server device, or an analysis using images from a fewer number of cameras than the number of cameras used in the analysis performed by the server device.

26. An image reading apparatus in an image reading system that includes a device carried or worn by a user and that reads out aloud objects from images captured by a camera, wherein the device automatically and repeatedly captures images from the camera's video at predetermined times, the image reading system analyzes the captured images to obtain information including text that represents objects in the images, the image reading system determines the objects and text to be read out based on the obtained information and at a predetermined judgment, and the device automatically and repeatedly reads out aloud the text that represents the determined objects at predetermined times.

27. An image reading method in an image reading system that includes a device that is carried or worn by a user and that reads out aloud objects from images captured by a camera, the image reading method comprising the steps of: the device automatically and repeatedly capturing images from the camera's video at predetermined times; the image reading system analyzing the captured images to obtain information including text that represents objects in the images; the image reading system determining, based on the obtained information and using predetermined judgment, the object and text to be read out; and the device automatically and repeatedly reading out aloud the text that represents the determined object at predetermined times.

28. A response output system comprising a device carried or worn by a user, which responds to the user, wherein the device comprises a plurality of cameras with different shooting ranges, automatically and repeatedly captures a plurality of images from the plurality of cameras at set timings, creates an input conversation sentence based on input by the user, inputs the input conversation sentence and the plurality of images into an interface, and the interface recognizes the user's surrounding situation based on image recognition processing of the plurality of images, and outputs an output conversation sentence generated in accordance with the input conversation sentence.

29. A response output system according to claim 28, wherein the interface inserts the generated output conversation sentence, which talks about a surrounding object detected from the surrounding situation in an image taken by at least one camera among the plurality of images, into the conversation flow.

30. A response output system according to claim 28, wherein the interface inserts the generated output conversation sentence, which is about the recognized facial expression of the user from an image captured by at least one camera among the plurality of images, into the conversation flow.

31. A response output system as described in claim 28, wherein the interface outputs the generated output conversation that talks about a surrounding object detected from the surrounding situation in an image taken by at least one camera among the plurality of images, and also controls whether to output the output conversation that talks about the surrounding object or whether to change to an output conversation that talks about another surrounding object, depending on the user's facial expression recognized from the image taken by at least one camera among the plurality of images, thereby inserting the output conversation into the flow of conversation.

Citation Information

Patent Citations

  • Voice dialog system

    JP2005037662A

  • Robot control device

    JP2008158697A

  • Emotional answer generation device and emotional answer generation program

    JP2009134008A

  • Portable suspicious individual detecting apparatus, suspicious individual detecting method, and program

    JP2010081480A

  • Environment information transmitting device

    JP2013017555A