Image reading system, device, and method, and response output system
The image reading system with multiple cameras and AI analysis addresses the limitations of existing technologies by automatically capturing and interpreting surroundings, providing continuous and hands-free object recognition and voice output for enhanced usability and safety.
Patent Information
- Application Number
- JP2024196060
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-29
- Filing Date
- 2024-11-08
- Publication Date
- 2025-10-14
AI Technical Summary
Existing image reading technologies, such as the 'Seeing AI' app, are not optimized for use while a user is walking, requiring frequent manual input operations and struggling to effectively capture and interpret complex surroundings due to changing environments.
An image reading system with multiple cameras mounted on a wearable device that automatically captures images periodically, analyzes them using AI, determines relevant objects, and reads them aloud at predetermined intervals without user input, prioritizing objects based on proximity and significance.
Enables continuous and hands-free interpretation of surroundings, enhancing usability for visually impaired individuals by providing real-time object recognition and voice output, improving safety and convenience during walking, nighttime navigation, and vehicle operation.
Smart Images

Figure 2025155703000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technology for converting images and videos into text and character strings and reading them aloud. [Background technology]
[0002] There is a technology (sometimes referred to as image reading or voice reading) that uses AI (artificial intelligence) to recognize and analyze images and videos (including still images and videos) captured by a camera, convert them into text and character strings (in other words, create sentences), and read them aloud. One such technology is the "Seeing AI" app for iPhones provided by Microsoft, as shown in Non-Patent Document 1. The "Seeing AI" app uses AI to analyze photographic images captured by a camera, convert them into text, and read them aloud. Such technology is effective, for example, in supporting the visually impaired.
[0003] Further, examples of prior art include Japanese Patent Application Laid-Open No. 2000-325389 (Patent Document 1), Japanese Patent Application Laid-Open No. 2006-251596 (Patent Document 2), and Japanese Patent Application Laid-Open No. 2017-77446 (Patent Document 3). [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2000-325389 [Patent Document 2] Japanese Patent Application Laid-Open No. 2006-251596 [Patent Document 3] Japanese Patent Application Publication No. 2017-77446 [Non-patent literature]
[0005] [Non-Patent Document 1] <URL:https: / / www.microsoft.com / ja-jp / ai / seeing-ai>,“Seeing AI” Summary of the Invention [Problem to be solved by the invention]
[0006] For example, the "Seeing AI" app converts images captured by the user into text. The "Seeing AI" app also has a "Scene" function that allows users to take photos and describe the captured scene.
[0007] However, such technology does not sufficiently consider how to make it more suitable for use while the user is walking, and there is room for improvement. For example, the user's input operations, such as a capture operation, are time-consuming. As the user's surroundings change over time as the user walks, the user must input operations multiple times in order to grasp the surroundings over time through audible reading of camera images. Furthermore, as the user walks, the surroundings may be complex, with various objects in front, behind, left, right, and so on. Therefore, grasping the complex surroundings through audible reading of camera images using a single camera is difficult.
[0008] Prior art examples such as Patent Document 1 disclose that, for example, a single camera on a smartphone performs image reading or the like in response to a user's capture operation. However, the prior art examples do not provide specific embodiments on how to best support visually impaired persons using images captured by the camera. Furthermore, the prior art examples do not consider a preferred embodiment for when a person wears the camera. Furthermore, the prior art examples do not sufficiently consider user usability or preferred functions and processes tailored to usage scenarios such as supporting visually impaired persons.
[0009] An object of the present disclosure is to provide a technology that can be more suitably used with respect to the image reading-out technology. [Means for solving the problem]
[0010] A representative embodiment of the present disclosure has the following configuration: One embodiment is an image reading system including a device carried or worn by a user, which reads out aloud an object from an image captured by a camera, wherein the device automatically and repeatedly captures images from the camera's video at predetermined times, analyzes the captured images to obtain information including text representing the object in the image, determines the object and text to be read out based on the obtained information at a predetermined judgment, and automatically and repeatedly reads out the text representing the determined object from the device at predetermined times. [Effects of the Invention]
[0011] According to a representative embodiment of the present disclosure, it is possible to provide a technology that can be used more preferably with respect to the image reading technology. Problems, configurations, effects, etc. other than those described above will be described in the description of the embodiment. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a diagram showing the configuration of an image reading-out system according to a first embodiment. [Figure 2] FIG. 10 is a diagram showing another configuration of the image reading-out system according to the first embodiment. [Figure 3] FIG. 10 is a diagram showing another configuration of the image reading-out system according to the first embodiment. [Figure 4] FIG. 2 is a diagram showing the processing flow of the image reading-out system and method according to the first embodiment. [Figure 5A] 1 is a diagram showing an example of the configuration of an image reading-out device according to a first embodiment. [Figure 5B] 1 is a diagram showing an example of the configuration of an image reading-out device according to a first embodiment. [Figure 6A] FIG. 1 is a diagram showing an example of use of the image reading-out device according to the first embodiment. [Figure 6B] FIG. 1 is a diagram showing an example of use of the image reading-out device according to the first embodiment. [Figure 7A] 1 is a diagram showing an example of the configuration of an image reading-out device according to the first embodiment. [Figure 7B] FIG. 2 is a diagram showing an example of the configuration of a server device in the first embodiment. [Figure 8] FIG. 2 is a diagram showing an example of a configuration of predetermined timings for image capture, voice reading, etc. in the first embodiment. [Figure 9A] FIG. 2 is a diagram showing an example of the configuration of a plurality of cameras in the first embodiment. [Figure 9B] FIG. 2 is a diagram showing an example of turning on / off a plurality of cameras in the first embodiment. [Figure 9C] FIG. 3 is a diagram showing an example of control of a plurality of cameras in the first embodiment. [Figure 10] FIG. 2 is a diagram showing an example of an object in the first embodiment. [Figure 11] FIG. 2 is a diagram showing an example of a relationship between a user and an object in the first embodiment. [Figure 12] FIG. 10 is a supplementary explanatory diagram regarding priority given to the direction of movement, etc., in the first embodiment. [Figure 13] FIG. 3 is a diagram showing control modes in the first embodiment. [Figure 14] FIG. 2 is a diagram showing an example of a setting information table in the first embodiment. [Figure 15A] FIG. 2 is a diagram showing an example of image recognition and voice reading in the first embodiment. [Figure 15B] FIG. 2 is a diagram showing an example of image recognition and voice reading in the first embodiment. [Figure 16A] FIG. 2 is a diagram showing an example of image recognition and voice reading in the first embodiment. [Figure 16B] FIG. 2 is a diagram showing an example of image recognition and voice reading in the first embodiment. [Figure 17] FIG. 3 is a diagram showing an example of a processing sequence between a device and a server apparatus in the first embodiment. [Figure 18] FIG. 3 is a diagram showing basic control relating to periodic control in the first embodiment. [Figure 19]FIG. 10 is a diagram showing control not to perform voice reading when there is no change in the first embodiment. [Figure 20] FIG. 10 is a diagram showing control for performing a voice readout indicating that there is no change when there is no change in the first embodiment. [Figure 21] FIG. 10 is a diagram showing an example of control of a read-out instruction again in the first embodiment. [Figure 22] 3A to 3C are diagrams showing examples of voice reading using different expressions in the first embodiment. [Figure 23] 4A to 4C are diagrams showing examples of control of front and rear cameras in the first embodiment. [Figure 24] FIG. 10 is a diagram showing an example of control in which threshold values for determination are made different before and after in the first embodiment. [Figure 25A] FIG. 2 is a diagram showing an example of support for a visually impaired person in the first embodiment. [Figure 25B] FIG. 2 is a diagram showing an example of assistance for nighttime roads in the first embodiment. [Figure 25C] FIG. 10 is a diagram showing an example of how to handle "walking while using a smartphone" in the first embodiment. [Figure 25D] FIG. 2 is a diagram showing an example of vehicle assistance in the first embodiment. [Figure 26] FIG. 10 is a diagram showing an example of control depending on whether server communication is possible or not in the first embodiment. [Figure 27] FIG. 2 is a diagram showing an example of voice reading according to the priority of an object in the first embodiment. [Figure 28] FIG. 3 is a diagram showing examples of risk factors and safe space guides in the first embodiment. [Figure 29] FIG. 2 is a diagram showing an example of voice reading that gives priority to a specific object in the first embodiment. [Figure 30] FIG. 2 is a diagram showing an example of an image taken by a 360-degree camera in the first embodiment. [Figure 31] 5A to 5C are diagrams showing an example of control in which different audio output modes are provided depending on the camera shooting direction in the first embodiment. [Figure 32] FIG. 3 is a diagram showing an example of control of three-dimensional audio output in the first embodiment. [Figure 33] FIG. 10 is a diagram showing the configuration of a system according to a second embodiment. [Figure 34A] FIG. 10 is a diagram showing the configuration of a device in the second embodiment. [Figure 34B] FIG. 10 is a diagram showing the configuration of a server in the second embodiment. [Figure 35] FIG. 10 is a diagram showing a processing flow in the second embodiment. [Figure 36A] FIG. 10 is a diagram showing an example of the system configuration in the second embodiment. [Figure 36B] FIG. 10 is a diagram showing an example of the system configuration in the second embodiment. [Figure 36C] FIG. 10 is a diagram showing an example of the system configuration in the second embodiment. [Figure 36D] FIG. 10 is a diagram showing an example of the system configuration in the second embodiment. [Figure 37] FIG. 10 is a diagram showing a more detailed example of the configuration of the system in the second embodiment. [Figure 38] FIG. 11 is a diagram showing an example of the configuration of request data in the second embodiment. [Figure 39] FIG. 10 is a diagram showing an example of a surrounding situation in the second embodiment. [Figure 40] FIG. 10 is a diagram showing an example of a plurality of images in the second embodiment. [Figure 41] FIG. 10 is a diagram showing the surrounding situation in a first specific example in the second embodiment. [Figure 42] FIG. 10 shows a first example of the conversation flow in the first specific example in the second embodiment. [Figure 43] FIG. 10 shows a second example of the conversation flow in the first specific example in the second embodiment. [Figure 44] FIG. 10 is a diagram showing the flow of conversation in a second specific example in the second embodiment. [Figure 45] FIG. 10 is a diagram showing an example of a screen display of a device in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the drawings, identical parts are generally designated by the same reference numerals, and repeated explanations will be omitted. In the drawings, the representation of components may not represent their actual positions, sizes, shapes, ranges, etc., in order to facilitate understanding of the invention.
[0014] For the sake of explanation, when describing processing by a program, the program, function, processing unit, etc. may be described as the main body, but the main hardware body for these is a processor, or a controller, device, computer, system, etc. that is configured with the processor, etc. A computer executes processing according to a program read into memory using resources such as memory and communication interfaces as appropriate through the processor. This realizes predetermined functions, processing units, etc. A processor is configured, for example, with semiconductor devices such as a CPU / MPU or GPU. Processing is not limited to software program processing, but can also be implemented using dedicated circuits. Dedicated circuits such as FPGAs, ASICs, and CPLDs can be used.
[0015] The program may be pre-installed as data on the target computer, or may be distributed as data from a program source to the target computer. The program source may be a program distribution server on a communication network or a non-transitory computer-readable storage medium, such as a memory card or disk. The program may be composed of multiple modules. The computer system may be composed of multiple devices. The computer system may be composed of a client-server system, a cloud computing system, an IoT system, etc. Various data and information may be composed of structures such as, but not limited to, tables and lists. Expressions such as identification information, identifiers, IDs, names, and numbers are interchangeable.
[0016] <First Embodiment> The image reading-out system and the like according to the first embodiment will be described with reference to FIG.
[0017] The image reading-out system and method of the first embodiment use a device (also referred to as an image reading-out device) equipped with a camera and an audio output device as a device carried or worn by a user. This device is equipped with at least a camera (i.e., a photographing function) and an audio output device (i.e., an audio output function), as well as other necessary software and hardware. This device includes a wearable device, and may be, for example, a smartphone, a smart watch, a tablet terminal, smart glasses, a head-up display (HMD), a shoulder device, etc. Furthermore, this device may be a hat, clothing, bag, shoes, etc. equipped with a camera, etc. The device may be equipped on a hat, clothing, etc.
[0018] The image reading system of the first embodiment utilizes multiple cameras mounted on a device that is an image reading apparatus. The multiple cameras may be, for example, cameras in two directions (front and back) or cameras in four directions (front, back, left and right). Each camera may be, in particular, a stereo camera or a sensor with a distance measurement function. The number and orientation of cameras mounted on a device are not limited and various possibilities are possible.
[0019] The device captures (or takes a picture of) an image from the input video of each camera. This image (also referred to as a captured image) is a photographic image of the surroundings taken in a specific direction, i.e., the shooting direction of the camera (e.g., forward or backward), relative to the user and the device. The image captures various objects, such as pedestrians and cars, as the surroundings.
[0020] The device automatically captures images periodically (periodically, referred to as a first time interval) from the video input by the camera at a predetermined timing. As a result, the device obtains a group of images that continuously / intermittently capture the user's surroundings in time series. The image reading system detects objects in the surroundings from the group of images through analysis, including recognition by AI (e.g., a machine learning model). For example, the image reading system extracts features associated with objects from the group of images through AI recognition, and detects, as objects, parts of the features that change significantly between images over time, in other words, parts where there is a significant change.
[0021] The image reading system determines whether a detected object is to be read aloud based on a predetermined judgment. The predetermined judgment may be, for example, whether the object is close to the user and device. In addition, if there are multiple objects, the image reading system also determines the priority of the objects to be read aloud, in other words, the output order.
[0022] The image reading system converts the object determined as the target into text representing the object based on the AI's recognition results. The AI (e.g., a machine learning model) responds with information such as the position and size of features associated with the object, as well as text information representing the object, as recognition results.
[0023] The image reading system device reads out the text representing the object (e.g., a pedestrian) to be read aloud from the audio output device in order of priority. The text and audio read aloud might be, for example, "pedestrian" or "there is a pedestrian ahead."
[0024] The image reading system also automatically reads out the object aloud at a predetermined period (second time interval). The image capture period (first time interval) and the read-out period (second time interval) may be different.
[0025] As described above, the device that is an image reading device automatically and intermittently captures images at predetermined times, repeatedly converts the captured images into text through AI recognition, and automatically and repeatedly reads aloud based on the captured images at predetermined times. This image capture and aloud reading operation basically does not require any operational input by the user, such as a capture instruction operation.
[0026] The image reading system stores data and information of the history of the continuous image capture and voice reading in a storage resource, for example, for at least a predetermined period of time, so that it can be referenced later.
[0027] [System Configuration (1)] Fig. 1 shows the configuration of an image reading-out system according to embodiment 1. The image reading-out system of Fig. 1 is a system including a device 1, which is an image reading-out apparatus. The system of Fig. 1 includes the device 1 and a server device 2 on a cloud computing system 3, which are appropriately connected by communication via a communication network 9.
[0028] In other words, device 1 is a user terminal (such as a mobile terminal, wearable terminal, information processing terminal, or computer) that user U1 carries or wears. User U1 may be, for example, visually impaired or non-visually impaired. Specific examples of device 1 include a smartphone 1A, a smart watch 1B, smart glasses (or HMD) 1C, a hat 1D, clothing 1E, a shoulder device 1F, and other wearable devices.
[0029] Note that one user U1 may use multiple devices 1. The multiple devices 1 may communicate with each other.
[0030] The device 1 has a control function 10, a communication function 11, a photographing function 12, an audio output function 13, etc. The control function 10 controls the image reading function. The communication function 11 performs communication processing with the server device 2, which is the AI server 2, via a communication network 9. The photographing function 12 is a function of capturing an image using a camera. The audio output function 13 is a function of outputting audio from an audio output device such as a speaker or earphone jack. In addition, although not shown, the device 1 may also have a display function, an audio input function, etc. The device 1 has at least the photographing function 12 and the audio output function 13. Details of the device 1 will be described later using Figure 7A.
[0031] The server device 2 has the functions of AI 20, or in other words, is an AI server 2. In other words, AI 20 has an analysis function and a recognition function. This AI 20 is software and hardware that recognizes an object (in other words, a feature) from an input image, converts the object (feature) into text representing the object (feature), and outputs information about the object including the text. Details of the server device 2 will be described later using FIG. 7B.
[0032] The device 1 and the server device 2 may function as a client-server system. The device 1 sends a request (e.g., an image) for recognition by the AI 20 to the server device 2 as needed. In response to the request, the server device 2 performs recognition using the AI 20 and sends the recognition result (e.g., text) to the device 1 as a response.
[0033] In this embodiment, the AI 20 of the server device 2 performs conversion to text. The conversion from text to voice is performed by the control function 10 of the device 1 (step S5 in FIG. 4). As a modified example, the AI 20 of the server device 2 may also perform conversion from text to voice after conversion to text, and output the converted voice data to the device 1.
[0034] In addition, in the embodiment described below, it is possible to switch between performing recognition by the AI 20 of the server device 2 and performing recognition within the device 1, depending on the communication status between the device 1 and the server device 2, etc. In this case, an AI function 20b may be provided within the device 1 shown in Fig. 1 as a corresponding component. This AI function 20b is a copy of the function of the AI 20 of the server device 2, or software / hardware that realizes a simple analysis function / recognition function with lower performance than the function of the AI 20.
[0035] [System Configuration (2)] Fig. 2 shows a system configuration of a modified example. The system in Fig. 2 is a system in which device 1a (in other words, an input / output device) having an image capturing function 12 and an audio output function 13, device 1b having a control function 10 and a communication function 11 as a main body, and server device 2 which is an AI server 2 are connected. Device 1a may be, for example, smart glasses 1C. Device 1b may be, for example, a smartphone 1A. The combination of device 1a and device 1b corresponds to device 1 in Fig. 1.
[0036] [System Configuration (3)] Figure 3 shows a system configuration of a modified example. The system in Figure 3 is a system in which the functions of the AI 20 of the AI server 2 in Figure 1 are integrated into a single device 1 such as a smartphone 1A. In this case, the AI 20 (in other words, analysis and recognition functions) is provided within the device 1, and there is basically no need to access the server device 2 (Figure 1) on the communication network 9 when recognizing camera images. The AI 20 in the device 1 may be downloaded and installed from a server device of a business operator on the communication network 9, or may be updated or upgraded via communication as appropriate.
[0037] [AI] In other words, the AI20 of the AI server 2 in FIG. 1 is a large-scale DB (database) for image analysis and text conversion. A multimodal large-scale language model (MLLM) can be applied to this AI20. The MLLM is an LLM that can process multiple types of data and information, such as images, text, and audio. The MLLM applied to the AI20 in this embodiment has the function of describing and expressing characteristic parts / areas / object shapes and positions in images using words / text / character strings / natural language based on learning from a group of images.
[0038] As the MLLM technology applicable to the AI 20, for example, known technologies such as those shown in the following reference information 1 and 2 can be applied.
[0039] [Reference information 1] <https: / / arxiv.org / abs / 2306.13549> ,”A Survey on Multimodal Large Language Models”
[0040] [Reference information 2] <https: / / browse.arxiv.org / pdf / 2306.13549.pdf> ,”A Survey on Multimodal Large Language Models”
[0041] [Usage scenarios] The image reading system according to the first embodiment can be used in a wide range of applications, including assisting a user when walking. This system can be used in a variety of applications. More specific application scenarios include the following:
[0042] 1. Support for the visually impaired: In this application, objects around the user are detected and read aloud to assist the visual ability of a visually impaired user. Figure 25A, which will be described later, shows an example of support for the visually impaired. For example, in the case of a visually impaired person with normal hearing, it is difficult or impossible for the person to recognize the surroundings while walking with their own eyes. However, an image reading system converts the surroundings from camera images into text and reads it aloud. This allows the visually impaired person to recognize objects in the surroundings with their hearing, contributing to road safety and convenience.
[0043] 2. Nighttime Road Assistance: In this application, for example, in dark conditions such as roads at night or in unsafe areas, the system detects whether a suspicious person is following the user from behind and provides a voice readout to warn the user. An example of nighttime road assistance is shown in Figure 25B, which will be described later.
[0044] 3. Support for "walking while using a smartphone": In this application, when the user uses the smartphone 1A or the like while walking or standing still, after ensuring the safety of the user's surroundings, surrounding objects are detected and read aloud as support. Figure 25C described below shows an example of support for "walking while using a smartphone." Note that this application does not encourage so-called "walking while using a smartphone" or "using a smartphone while doing other things," which is using a smartphone or the like in a situation where the safety of the user's surroundings is not ensured, but provides support that contributes to the safety of the user and those around them.
[0045] 4. Vehicle Driving Assistance: For example, when a user is riding a vehicle such as an electric bicycle or electric kick scooter, this application detects surrounding objects and provides voice readings as assistance. This application does not encourage the use of smartphones while driving, which may lead to accidents, but rather provides assistance that contributes to the safety of the user and those around them. Figure 25D, described below, shows an example of vehicle driving assistance.
[0046] Even for users who are not visually impaired, the above-mentioned nighttime road assistance, support for dealing with "walking while using a smartphone," and vehicle driving assistance are effective.
[0047] [Image reading method and processing flow] FIG. 4 shows a processing flow of the image reading method according to the first embodiment. The system of FIG. 1 executes the processing flow of FIG. 4. In step S1, the device 1 of the user U1 turns on the image reading function. In step S2, the device 1 captures an image from the camera at a predetermined timing. In other words, the device 1 takes an image. In step S3, the system analyzes the captured image obtained in step S2. The analysis includes object recognition by the AI 20. As a result of the analysis, the system obtains information including text representing the object (in other words, words, character strings, sentences, descriptions, etc.). In step S4, the system determines the object and text to be read aloud based on the result of step S3 through a predetermined judgment (for example, a judgment of the degree of change, which will be described later). In step S5, the system performs a voice reading (in other words, voice output) of the text representing the object determined in step S4 from the device 1 at a predetermined timing. In step S6, device 1 checks whether to turn off the image reading function. If it is to be turned off (YES), the flow ends. If it is to be kept on (NO), the flow returns to step S1 and the same process is repeated.
[0048] [Device example] FIG. 5A shows an example of the implementation of cameras and the like in smart glasses 1C as an example of device 1 in FIG. 1. In this example, smart glasses 1C are equipped with multiple cameras with four shooting directions: front, back, left, and right. Note that the coordinate system WU in FIG. 2 indicates a coordinate system (X, Y, Z) based on user U1 and device 1. The X-axis / X direction is the left-right direction as seen from user U1, the Y-axis / Y direction is the front-back direction as seen from user U1, and the Z-axis / Z direction is the up-down direction as seen from user U1.
[0049] The smart glasses 1C are equipped with cameras C11 and C12 located on the left and right sides of the front surface, which take images facing forward, and cameras C21 and C22 located on the left and right sides of the rear surface, which take images facing backward. The smart glasses 1C also have cameras C31 and C32 located in front and behind the right side, which take images facing right, and cameras C41 and C42 located in front and behind the left side, which take images facing left. The cameras C11 and C12 take images in the forward direction (+Y direction), in other words, they are front cameras. The cameras C21 and C22 take images in the backward direction (-Y direction), in other words, they are rear cameras. The cameras C31 and C32 take images in the right direction (+X direction), in other words, they are right cameras. The cameras C41 and C42 take images in the left direction (-X direction), in other words, they are left cameras.
[0050] In this example, for example, the front camera functions as a stereo camera (in other words, a sensor with a distance measurement function) using two cameras C11 and C12, but is not limited to this. A single camera may be used as the front camera. Relative distances, which will be described later, can be determined by analyzing images using AI 20, but the distance measurement function of such a stereo camera may also be used.
[0051] A device 1 such as smart glasses 1C may be provided with a remote controller 1c or the like. Alternatively, another smartphone 1A or the like may function as a remote controller for the smart glasses 1C or the like. When a user U1 operates a button or the like provided on the remote controller 1c, a signal such as an instruction is transmitted from the remote controller 1c to the smart glasses 1C. The smart glasses 1C operates in accordance with the signal.
[0052] Although not shown, the smart glasses 1C are also equipped with a display, a microphone, a speaker, a vibrator, a battery, etc. The audio output device is not limited to a speaker, and bone conduction earphones or the like may also be used.
[0053] If the user U1 is not visually impaired, he or she may use the video information displayed on the display surface 1Cd of the smart glasses 1C by looking at it.
[0054] FIG. 5B shows a shoulder-mounted device 1F as another example of the device 1. The shoulder-mounted device 1F is a type of device that is worn around the shoulder or neck. The shoulder-mounted device 1F in FIG. 5B has two rod-shaped housings on the left and right, and a semi-ring-shaped housing on the back side that connects the housings. This shoulder-mounted device 1F is equipped with multiple cameras that can capture images in four directions: front, back, left, and right. The multiple cameras are the same as the multiple cameras in FIG. 5A.
[0055] The above example is an implementation example in which cameras are provided in four directions (front, rear, left, and right), but this is not limiting; cameras that capture images in at least one direction may be used. The left and right cameras (C31, C32, C41, C42) may be omitted, and only front and rear cameras (C11, C12, C21, C22) may be provided. If cameras (C21, C22) that capture images in the rear direction are provided, the situation behind the user can be captured appropriately. If cameras (C31, C32, C41, C42) that capture images in the left and right directions are provided, the situation to the left and right of the user can be captured appropriately. Note that, for example, if the angle of view of the front camera is sufficiently large (for example, close to 180 degrees), the situation to the left and right of the user can also be captured.
[0056] [Device usage: Smartphone] FIG. 6A shows an example of how the device 1 is used when it is a smartphone 1A. FIG. 6A shows how a user U1 holds the smartphone 1 in his / her hand and captures images of the user U1 in the front-to-back directions (±Y directions) using cameras on the front and back (front and back) of the smartphone 1A. In FIG. 6A, the user U1 holds the smartphone 1A in his / her hand so that the housing is nearly vertical. Camera C1 is a front camera (outer camera) provided on the back side (side without a display screen) of the flat housing of the smartphone 1A, with its optical axis for capturing images in the forward direction (+Y). Camera C2 is a rear camera (inner camera) provided on the front side (side with a display screen) of the flat housing of the smartphone 1A, with its optical axis for capturing images in the backward direction (-Y).
[0057] If the user U1 is not visually impaired, the user U1 may use the smartphone 1A by viewing the video information on the display screen 1Ad as needed.
[0058] The user U1 may hold the smartphone 1A in a direction other than the forward direction shown in the figure, in which case the front camera C1 can capture an image in that direction. If the smartphone 1A is equipped with a camera that captures an image in another direction, that camera can also be used.
[0059] [Device usage: hats and clothes] FIG. 6B shows another example of how device 1 is used, in which cameras are mounted on hat 1D and clothing 1E to capture images of user U1 in the front and back directions. For example, hat 1D is equipped with camera C3 on the front side for capturing images in the forward direction (+Y), and camera C4 on the back side for capturing images in the backward direction (-Y). Alternatively, clothing 1E is equipped with camera C5 on the front side for capturing images in the forward direction (+Y), and rear camera C6 on the back side for capturing images in the backward direction (-Y). These cameras may be connected via communication with device 1 (e.g., smartphone 1A) that serves as the main body. Similarly, cameras may be mounted on the left and right sides.
[0060] The device 1 may be attached to the body, clothes, bag, shoes, etc. of the user U1. For example, a camera or a device 1 with a camera may be attached to a pocket / holder of clothes 1E. The user U1 may wear / carry multiple devices 1 such as smartphones 1A. Even in the case shown in FIG. 6B, the device 1 can capture the front and rear directions of the user U1 using front and rear cameras.
[0061] [Device configuration example] FIG. 7A shows an example configuration of device 1. Device 1 includes processor 301, memory 302, nonvolatile memory (i.e., storage) 303, camera (i.e., image capture unit) 304, image processing unit 305, image storage unit 306, position detection unit 311, orientation detection unit 312, biometric information detection unit 313, button / operation unit 314, display / display unit 315, microphone / audio input unit 316, speaker / audio output unit 317, connector / input / output interface 318, battery / power supply unit 319, vibration unit 320, and communication unit 330. Communication unit 330 is equipped with a communication interface and includes, for example, first wireless communication unit 331, second wireless communication unit 332, third wireless communication unit 333, and wired communication unit 334. Each wireless communication unit is equipped with a respective wireless communication interface and is implemented, for example, in accordance with a mobile network, wireless LAN, Bluetooth (registered trademark), or the like. Camera 304 includes multiple cameras, for example, cameras #1 to #8. These cameras correspond to, for example, the eight cameras on the front, back, left and right sides in FIG. 5A.
[0062] Processor 301 is composed of a CPU and the like, and controls the entire device 1 and each unit. Memory 302 stores data and information to be processed by processor 301. Processor 301 reads a program from non-volatile memory 303 into memory 302 and executes processing in accordance with the program. This allows various functions to be realized as execution modules.
[0063] The non-volatile memory 303 stores various data and information, such as programs such as an OS 341, image reading software 342, AI data 343, user information 344, setting information 345, and access destination information 346. The image reading software 342 is data such as a program for realizing the image reading function in the first embodiment. The AI data 343 is data such as a recognition result exchanged with the AI 20 of the server device 2. When the recognition process is performed on the device 1 side, the AI data 343 includes data corresponding to the AI 20. The user information 344 is information about the user U1. The setting information 345 is system setting information and user setting information. The access destination information 346 is information for communicating with the server device 2, etc.
[0064] Camera 304 includes a circuit for controlling image capture by multiple cameras, etc. Image processing unit 305 processes images captured by camera 304. Image storage unit 306 stores data such as images captured by camera 304 and images processed by image processing unit 305.
[0065] The position detection unit 311 detects the position of the device 1 using, for example, GNSS or the like and obtains position information. The orientation detection unit 312 includes an orientation sensor and detects the orientation of the device 1. The biometric information detection unit 313 detects, for example, the fingerprint of the user U1 to perform fingerprint authentication, or detects the face of the user U1 to perform face authentication. The button / operation unit 314 includes a power button, volume button, etc., and accepts operation inputs from the user U1. The display / display unit 315 processes video and images on the display. The microphone / audio input unit 316 inputs and recognizes audio through a microphone. The speaker / audio output unit 317 outputs audio through a speaker, earphone jack, etc. The connector / input / output interface 318 is equipped with an input / output interface for connecting input / output devices. The battery / power supply unit 319 supplies the necessary power to each unit using a battery. The vibration unit 320 generates vibrations as needed.
[0066] [Server configuration example] 7B shows an example configuration of the server 2. The server 2 has a processor 401, a memory 402, a non-volatile memory (in other words, storage) 403, an operation input unit 414, a connector / input / output interface 418, a battery / power supply unit 419, and a communication unit 430. The communication unit 430 is equipped with a communication interface, and includes, for example, a first wireless communication unit 431, a second wireless communication unit 432, a third wireless communication unit 433, and a wired communication unit 434.
[0067] The processor 401 is configured with a CPU / GPU, etc., and controls the entire server device 2 and each unit. The memory 402 stores data and information to be processed by the processor 401. The processor 401 reads a program from the non-volatile memory 403 into the memory 402 and executes processing in accordance with the program. This allows various functions to be realized as execution modules.
[0068] The non-volatile memory 403 stores various data and information, such as programs such as an OS 441, AI software 442, AI data 443, user information 444, setting information 445, and access destination information 446. The AI software 442 is data such as a program for implementing the AI 20 in FIG. 1, which performs analysis processing including image recognition processing. The AI data 443 is data such as models and image data for the recognition processing of the AI 20, and recognition results exchanged between the AI 20 and the device 1. The user information 444 is information about the user U1. The AI setting information 445 is setting information for the AI 20. The access destination information 446 is information for communicating with the device 1, etc.
[0069] An operation input unit 414 receives operation input. An input / output interface for connecting an input / output device is implemented in a connector / input / output interface 418. A battery / power supply unit 419 supplies the necessary power to each unit using a battery.
[0070] [Specified timing for image capture, etc.] FIG. 8 shows an example of the relationship between input video 801 (time point or image frame) from camera 304 of device 1, image capture 802, recognition 803 by AI 20, and voice reading 804, including predetermined timing and cycles.
[0071] The input video 801 has image frames at each time point in the time series. The image frames are captured at a predetermined frame cycle of the camera 304. For example, frame f1 is captured at time point 1. Image capture 802 is automatically performed continuously / intermittently at a preset cycle (first time interval). Images are captured from the image frames of the input video 801 at that cycle. For example, frame 1 at time point 1 is captured as image 1, and frame 3 at time point 3 is captured as image 2. In this example, image capture 802 is performed at a cycle of every two frames. Note that in practice, capture may be performed once every eight frames, for example. The captured images are stored in memory (image storage unit 306). The images may be captured as still images by each camera, or may be captured as one frame of a video during video capture.
[0072] Recognition 803 of AI20 is performed for each captured image. In this example, recognition 803 of AI20 is performed at the same cycle as image capture 802. Voice reading 804 is automatically performed at a preset cycle (second time interval) for objects detected from the results of recognition 803 and determined as targets for voice reading. For example, voice reading 1 is performed for an object detected from image 1 at time point 1, and voice reading 2 is performed for an object detected from image 3 at time point 3. In this example, voice reading 804 is performed at a cycle of four frames. Note that if no object is detected as a result of recognition 803 of the captured image, or if the detected object is not a target for voice reading, voice reading 803 will not be performed even at a timing corresponding to the cycle.
[0073] The device 1 stores a history of the processes such as the image capture 802, recognition 803, and voice reading 804 (including data such as captured images and recognition results) for at least a predetermined period of time in a storage resource (for example, the nonvolatile memory 403). The device 1 may transmit the history data to the server device 2 or the like for storage.
[0074] If the device 1 determines that the image captured this time has a different movement or a significant change in the object compared to the image captured last time, the device 1 controls the device 1 to give priority to reading out the object. A specific example will be described later.
[0075] [Number of cameras and shooting direction] FIG. 9A is a schematic diagram viewed from above in the vertical direction, showing an example of the number and shooting directions of the cameras 304 of the device 1 of the user U1. State A is a first example. The user U1 is walking, for example, on a sidewalk 901. The user U1 carries or wears the device 1 (for example, a smartphone 1A). The first example is an example in which two cameras (C1, C2) on the front and back of the smartphone 1A, as shown in FIG. 6A, are used as front and rear cameras. For clarity, the front and rear cameras (C1, C2) are illustrated separately in the drawing. DC1 is the shooting direction of the front camera (C1), which is the +Y direction corresponding to the front direction of the user U1. DC2 is the shooting direction of the rear camera (C2), which is the -Y direction corresponding to the rear direction of the user U1. In this example, the front camera (C1) and the rear camera (C2) are equipped with wide-angle lenses, and the shooting range / angle of view is close to 180 degrees. The semicircle is a general image of the shooting range / angle of view. The same applies when FIG. 6B is applied.
[0076] State B is a second example. The second example is an example in which four cameras (front, back, left, and right) of smart glasses 1C as shown in FIG. 5A are used. A user U1 is wearing a device 1 (for example, smart glasses 1C). Here, the front camera is C10, the rear camera is C20, the right camera is C30, and the left camera is C40. DC10 is the shooting direction (+Y) of the front camera C10. DC20 is the shooting direction (-Y) of the rear camera C20. DC30 is the shooting direction (+X) of the right camera C30. DC40 is the shooting direction (-X) of the left camera C40. The dashed triangles are an overview image of the shooting range / angle of view of each camera, which is at least 45 degrees or more and may be close to 180 degrees. The same applies when FIG. 5B is applied.
[0077] The camera 304 is not limited to being integrally mounted on the device 1, and may be configured such that an external camera device is connected to the device 1 (for example, a smartphone 1A) that serves as the main body. The camera device may also be attached to a predetermined position on the body (for example, a hat 1D or clothing 1E).
[0078] [Camera On / Off] In other embodiments, only some of the multiple cameras of the device 1 may be set to be used in an on state. FIG. 9B shows an example of on / off of multiple cameras. State A is an example in which only the front camera C10 of the four cameras in the front, back, left, and right directions of the smart glasses 1C is in an on state. State B is an example in which the front camera C10 and the rear camera C20 are in an on state. State C is an example in which the front camera C10, the right camera C30, and the left camera C40 are in an on state. State D is an example in which the right camera C30 and the left camera C40 are in an on state.
[0079] [Example of controlling multiple cameras] Furthermore, when using multiple cameras, such as front and rear cameras, on device 1, it is also possible, as an embodiment, to simultaneously and in parallel perform the processes from image capture from each camera to voice reading (e.g., FIG. 8). However, in this case, the processing load is high and delays may be a concern. Therefore, as an embodiment, the processing of images from each of the multiple cameras may be controlled to have a predetermined timing or frequency, in other words, to be separated in time. For example, the overall schedule may be controlled so that the capture, recognition, and voice reading processes from the front camera (C1) and the capture, recognition, and voice reading processes from the rear camera (C2) have a predetermined timing or frequency. For example, the following control example can be given.
[0080] First, processing parts that require a low load even in parallel processing, such as image capture and image saving, may be performed repeatedly in the same way for both the front and rear cameras. If the parallel processing load for these image capture and image saving operations is high, the timing may be divided and the processing may be performed sequentially. Then, when device 1 reads the captured images from each camera from memory and performs subsequent processing, the timing is divided for each image from each camera. For example, the images from the front and rear cameras may be read and processed alternately.
[0081] Control 1 in Figure 9A shows an example of control of the front and rear cameras corresponding to state A. For example, at time points 1 and 3, the image from the front camera (C1) is processed (recognition and voice reading), and at the next time points 2 and 4, the image from the rear camera (C2) is processed (recognition and voice reading), and so on, alternately. In this example, the cycle (second time interval) of voice reading in processing the image from the front camera (C1) and the cycle (second time interval) of voice reading in processing the image from the rear camera (C2) alternate, and the processing frequency ratio is 1:1.
[0082] Control 2 in Figure 9A shows an example of control of the front, rear, left, and right cameras corresponding to state B. For example, at time point 1, an image from the front camera (C10) is processed, at the next time point 2, an image from the rear camera (C20) is processed, at the next time point 3, an image from the right camera (C30) is processed, and at the next time point 4, an image from the left camera (C40) is processed, and so on, with these four directions of processing being treated as one set and repeated in the same manner thereafter. In this example, the voice read-out period (second time interval) is the same for processing images from each direction, but the timing is shifted by one, resulting in a processing frequency ratio of 1:1:1:1.
[0083] As another example of control, it is possible to normally set the state to repeatedly process only images from the front camera (C1), and when a predetermined condition is met, change to a state to repeatedly process only images from the rear camera (C2). Alternatively, it is possible to normally set the state to repeatedly process only images from the front camera (C1), and when a predetermined condition is met, change to a state to alternately process images from the front camera (C1) and images from the rear camera (C2).
[0084] FIG. 9C shows an example of control for changing the processing ratio of each camera in the case of a device 1 equipped with at least front and rear cameras (C1, C2) similar to state A in FIG. 9A. Control A repeats processing of only images from the front camera (C1). Control B mainly processes images from the front camera (C1) while also inserting processing of images from the rear camera (C2) at a frequency of 3:1. Control C alternately repeats processing of images from the front camera (C1) and images from the rear camera (C2) at a frequency of 1:1 (similar to control 1 in FIG. 9A). Control D mainly processes images from the rear camera (C2) while also inserting processing of images from the front camera (C1) at a frequency of 3:1. Control E repeats processing of only images from the rear camera (C2).
[0085] The image reading system may be configured to switch between such multiple controls at appropriate times based on a user's operation input or a predetermined decision.
[0086] From the perspective of computational power and other factors, it may be difficult to achieve real-time, highly accurate voice reading for all of the surrounding conditions in front, behind, left, and right of the user. Furthermore, even if all of the surrounding conditions in each direction are read out simultaneously, the user may become confused. In such cases, the camera to be used, processing ratio, control mode, etc., as described above, can be set / selected depending on which direction recognition and reading is prioritized for the user and usage scenario.
[0087] For example, if rear support is not required depending on the user, Control A may be used. If only rear support is required depending on the user, Control E may be used. Also, it is possible to apply Control C under normal circumstances to recognize the front and rear in a balanced manner, switch to Control B when focusing on an object in front, and switch to Control D when focusing on an object in the rear (see FIG. 23 below).
[0088] [User Surroundings] FIG. 10 is an explanatory diagram of the surrounding situation when a user U1 wearing a device 1 (e.g., smart glasses 1C) is standing still or walking, for example, on a sidewalk 901, and is a schematic diagram viewed vertically from above. The directions are front, back, left, and right as seen from the user U1. In this example, it is assumed that the user U1 is walking forward (+Y direction) on the sidewalk 901. The white arrow indicates the direction of walking / movement. In this example, there is a bicycle lane 902 and a roadway 903 on the right side of the sidewalk 901.
[0089] In the illustrated example, examples of objects to be detected include a pedestrian 1001, a bicycle 1002, and a car (automobile) 1003, all of whom are other people. The pedestrian 1001 is walking on the sidewalk 901 in the -Y direction shown in the figure, and is approaching the front of the user U1. From the perspective of the user U1, the pedestrian 1001 is in front. The bicycle 1002 is traveling in the bicycle lane 902 in the -Y direction, and is approaching the user U1 diagonally in front. From the perspective of the user U1, the bicycle 1002 is diagonally in front and to the right. The car 1003 is traveling on the roadway 903 in the +Y direction. From the perspective of the user U1, the car 1003 is on the right side. Such objects are the targets of voice reading.
[0090] [Positional relationship between the user and the object] FIG. 11 shows an example of the positional relationship between the user U1 and an object. An explanation will be given using an example of a pedestrian 1001 (FIG. 10) that is facing forward from the user U1. State A is when the user U1 is stationary and the pedestrian 1001 is stationary. State B is when the user U1 is stationary and the pedestrian 1001 is walking. State C is when the user U1 is walking and the pedestrian 1001 is stationary. State D is when the user U1 is walking and the pedestrian 1001 is walking.
[0091] In either state, the relative positional relationship, speed, and other relationships between the user U1 and the pedestrian 1001 can be considered. The dashed circle represents a predetermined distance range 1100 centered on the positions of the user U1 and the device 1. The radius 1101 of the range 1100 corresponds to the threshold value of the relative distance between the user U1 and the object. In the example of FIG. 11, the threshold value of the relative distance is the same in each direction. In each of states A to D, there is a relative distance DA between the user U1 and the pedestrian 1001. In state A, the relative speed between the user U1 and the pedestrian 1001 is 0. In state B, when the speed of the pedestrian 1001 is v2(-v2), the relative speed is v2(-v2). In state C, when the speed of the user U1 is v1(+v1), the relative speed is v1(+v1). In state D, if the velocity of the pedestrian 1001 is v2 (-v2) and the velocity of the user U1 is v1 (+v1), the relative velocity is +v1-(-v2)=v1+v2.
[0092] In this embodiment, the above-described relative distance and relative speed can basically be determined based on the analysis of the camera image. However, without being limited to this, the position, orientation, speed, etc. of the user U1, the position, orientation, speed, etc. of the object, and the relative distance and speed, etc. between the user U1 and the object may be determined using various sensor functions provided in the device 1.
[0093] Furthermore, the above-described relative distance and speed can be used for predetermined control. For example, when the relative distance between the user U1 and the object is within a threshold, the object may be subject to voice reading or an alert. For example, when the relative distance DA is within a predetermined distance range 1100, the object may be subject to voice reading. Furthermore, the predetermined distance range 1100 (corresponding threshold) may be set in multiple stages. For example, when the relative distance DA is within a first range, the object may be subject to voice reading, and when it is within a smaller second range, the object may be subject to voice reading with an additional alert. Similarly, when the relative speed between the user U1 and the object is equal to or greater than a threshold, the object may be subject to voice reading or an alert.
[0094] In addition, in this embodiment, the relative relationship between the user and the object is grasped by handling the amount of change in the object between captured images (described later). This makes it possible to similarly process various situations according to combinations of the user being still or moving and the object being still or moving, as shown in Figure 11.
[0095] [Relationship between user movement direction and camera direction] 12 is a supplementary diagram illustrating the relationship between the moving direction of the user U1 and the shooting direction of the camera 403. The moving direction of the user U1 may or may not match the shooting direction of the camera.
[0096] State A is a case where the device 1 of user U1 has front and rear cameras. User U1 is moving forward (+Y direction). User U1's head is facing forward (+Y). The direction of user U1's movement, the direction of his head, and the shooting direction (DC1) of the front camera (e.g., C1) are the same. In state A, if recognition is set to prioritize the direction of user U1's movement, control can be performed to prioritize the capture and recognition of images from the front camera (C1). Note that, for example, if it is desired to recognize an object to the left (-X) of user U1, images from front and rear cameras with a sufficiently large angle of view may be used (used together). Alternatively, images from the left and right cameras of smart glasses 1C may be used.
[0097] As another example, in state B, user U1 is moving to the left (-X direction). User U1's head is facing forward (+Y). The direction of user U1's movement does not match the shooting direction of the front camera (C1). The direction of user U1's head matches the shooting direction of the front camera (C1). In state B, if recognition is set to prioritize the direction of user U1's movement (for example, leftward), images from front and rear cameras with a sufficiently large angle of view may be used (or used together). Alternatively, images from the left and right cameras of the smart glasses 1C may be used. Furthermore, if recognition is set to prioritize the direction of the head in state B, images from the front camera (C1) may be used.
[0098] As another example, state C is a case where the smart glasses 1C of user U1 have cameras on the front, back, left, and right. User U1 is moving, for example, forward (+Y), with his head facing, for example, diagonally forward to the left. The direction of user U1's head (diagonally forward left) and the shooting direction of the front camera (C10) are the same. User U1's movement direction (+Y) and the shooting direction of the front camera (C10) (diagonally forward left) do not match. In addition, the shooting direction of the right camera (C30) is diagonally forward to the right. User U1's movement direction (+Y) and the shooting direction of the right camera (C30) (diagonally forward right) also do not match. There is no camera that directly captures the forward direction (+Y), which is the movement direction, but there are cameras with adjacent shooting directions as close as possible to the forward direction: the front camera (C10) and the right camera (C30). In state C, when the setting is to prioritize recognition in the direction of movement of user U1 (for example, forward), it is sufficient to use both an image from the front camera (c10) and an image from the right camera (C30), which have a sufficiently wide angle of view. For example, it is also possible to use an image that combines the right part of the image from the front camera (C10) and the left part of the image from the right camera (C30). Also, in state C, when it is desired to prioritize recognition in the direction of the head, it is sufficient to use the image from the front camera (C10).
[0099] As another example, state D is a case where the front and rear cameras C5, C6 are located on the trunk of user U1, for example, on clothing 1E (FIG. 6B). User U1 is moving, for example, in the forward direction (+Y). User U1's head is facing diagonally forward to the left, for example, while the trunk and clothing 1E are facing forward (+Y). The movement direction (+Y) of user U1 and the shooting direction of the front camera (C5) match. The direction of user U1's head (diagonally forward to the left) do not match the shooting direction of the front camera (C5). The same applies when the smartphone 1C held by user U1 is facing forward. In state D, if the setting is to prioritize recognition in the movement direction of user U1 (for example, forward), the image from the front camera (C5) can be used.
[0100] As in the example of state C above, when it is desired to recognize a direction different from the camera's shooting direction (for example, the direction of movement), this can be achieved by using an image from a camera with a sufficiently large angle of view, or by combining images from multiple cameras with different shooting directions. Even in the case of the setting that prioritizes the direction of movement, which will be described later, this can be achieved by using images from each camera, as in the example above.
[0101] The device 1 may use various sensors to detect the movement direction of the user U1 in three-dimensional space, the direction of the head, the direction of the trunk, and the direction of each camera. For example, the movement direction of the user U1 can be detected by a sensor of the posture detection unit 312 (FIG. 7A). The device 1 may use each detected direction for control. Furthermore, if the device 1 (for example, smart glasses 1C) has a gaze detection function, the gaze direction detected by the gaze detection function may be used for control.
[0102] [Control Mode] FIG. 13 shows a table summarizing several control modes as an explanatory diagram of the control modes. The image reading system having the device 1 and the server device 2 of FIG. 1 may implement only one specific mode as a control mode, or may implement multiple modes and be able to switch between modes as needed. For example, table 1301 shows control modes related to the camera shooting direction. As examples of control modes, the first mode (M1) is a mode in which shooting and recognition are performed only in the forward direction based on the user U1 and the device 1. The second mode (M2) is a mode in which shooting and recognition are performed in the forward and backward directions. The third mode (M3) is a mode in which shooting and recognition are performed in the forward, backward, left, and right directions. For example, for the four cameras in the forward, backward, left, and right directions of the smart glasses 1C of FIG. 5A, which camera to use is set and controlled depending on the control mode.
[0103] Table 1302 also shows control modes related to the use and cooperation of the server device 2. As examples of control modes, the first mode (MA) is a mode in which device 1 performs analysis and recognition processing on its own without communicating with the server device 2. The second mode (MB) is a mode in which communication is performed with the server device 2, and analysis and recognition processing is performed by the server device 2. The third mode (MC) is a mode in which communication is performed with the server device 2, and analysis and recognition processing is performed by the server device 2 when communication with the server device 2 is possible (or good), and analysis and recognition processing (for example, simple analysis and recognition, which will be described later) is performed by the device 1 when communication with the server device 2 is impossible (or has deteriorated).
[0104] [Settings information] Figure 14 shows an example of a table of setting information for this system. This setting can be made for each user. It is also possible for each user to use multiple settings. User settings for the image reading function can be made on the graphical user interface (GUI) screen provided by device 1 or server device 2 of this system.
[0105] This table has the following setting items: #1 "Camera used", #2 "Reading camera when multiple cameras are used", #3 "Camera image capture interval", #4 "Reading interval", #5 "Determination criteria for reading from the main camera", #6 "Determination criteria for reading from the sub-camera", #7 "Reading target priority", #8 "No change reading function", #9 "Reading function with different expressions", #10 "Safe space guide function", #11 "History retention period", #12 "Other detailed settings", etc.
[0106] In the #1 "Camera Used" item, the camera 403 equipped in the device 1 can be set to be used for shooting by turning it on or off. For example, the front, back, left, and right cameras of the smart glasses 1C can be set to be on or off. Alternatively, when using a 360-degree camera, which will be described later, it can be set to be on or off. On is the enabled state, and off is the disabled state.
[0107] #2 In the "Reading camera when using multiple cameras" item, when using multiple cameras for voice reading, it is possible to set which camera's image to use for voice reading. The priority / priority of voice reading may be set for each camera of multiple cameras.
[0108] The setting value "Even" is a control that executes multiple voice readings corresponding to the multiple cameras used evenly. An example of equal control is to set an equal processing ratio for the front, rear, left, and right cameras as described above (Figure 9A), and perform voice readings sequentially in a predetermined order.
[0109] The setting value "Front direction basic (main)" basically reads out the image from the front camera. Reads out the image from other directions only when necessary (for example, when certain conditions are met) by interrupting, etc. For example, in control B of Figure 9C, the front camera is the main camera and the rear camera is the sub-camera. The setting value "Rear direction basic (main)" basically reads out the image from the rear camera. Reads out the image from other directions only when necessary (for example, when certain conditions are met) by interrupting, etc. For example, in control D of Figure 9C, the rear camera is the main camera and the front camera is the sub-camera. A specific example is shown in Figure 23, which will be described later.
[0110] The setting value "movement direction basic" is a control that basically executes voice reading from the image of the camera associated with the movement direction of user U1 (see, for example, FIG. 12 described above). Voice reading from other directions is performed only when necessary (for example, when a predetermined condition is met) by interrupting, etc. Explaining using the example of FIG. 12, in state A, the movement direction is forward (+Y), and voice reading from the image of the front camera in the same direction is basically executed. In state B, the movement direction is left (-X), and if there is a left camera that matches the movement direction, voice reading from the image of the left camera is basically executed. If there is no camera that matches the movement direction, processing can be performed using the image of another camera that is as close as possible to the movement direction. In state C, the movement direction is forward (+Y), and there is no single camera that matches the movement direction, but processing can be performed using the image of the front camera or the right camera that is as close as possible to the movement direction.
[0111] The #3 "Camera Image Capture Interval" item allows you to set the first time interval (cycle, frequency, etc.) for capturing images to be used for voice reading from a time-series group of camera images. It may also be possible to set the time interval for the camera used primarily (i.e., the main camera) and the time interval for other cameras used (i.e., the sub-camera). This capture interval may be set numerically, or may be set by selecting from several levels, for example.
[0112] The #4 "Reading Interval" item allows you to set a second time interval (cycle, frequency, etc.) for reading aloud based on the captured image. This reading interval may be set numerically, or may be set by selecting from several levels. For example, it may be set to high (every second), medium (every minute), or low (every few minutes). It may also be possible to set a time interval for the camera used primarily (in other words, the main camera) and a time interval for other cameras used (in other words, the sub-camera). The system processing load can be adjusted according to the settings of #3 and #34.
[0113] #5 "Decision Criteria for Reading from Main Camera" item allows you to set the decision criteria for reading aloud for the main camera according to the setting of the reading camera in #2. Setting value A. "Objects in the direction of movement" is a control that performs voice reading of objects detected in the direction of movement / travel of user U1 as much as possible. Note that as another control, it is also possible to apply a control that performs voice reading of all recognized and detected objects as much as possible.
[0114] Setting value B. "Approaching object" is a control that prioritizes an object approaching the user U1 as the target for voice reading (described later). It is also possible to set the relative distance and speed thresholds (Figure 11) used to determine an approaching object with this control.
[0115] Setting value C. "Significant Change" is a control that detects characteristic parts with significant changes between images and prioritizes objects corresponding to those significant changes for voice reading (described later). A significant change may be, for example, a change or difference between images in terms of features or characteristic points that is greater than a threshold.
[0116] Setting value D. "Specific object" is a control that detects a specific object and makes it the target of voice reading. The specific object can be set by the user (for example, by selecting from candidates) on another screen that transitions. Examples of specific objects include crosswalks, traffic lights, and traffic signs, which will be described later.
[0117] Setting value E. "Danger detection" is a control that detects dangerous situations / objects in the surroundings and makes them the subject of voice reading (described later).
[0118] The item #6 "Sub-camera reading interrupt timing" allows you to set the timing (second time interval) of the voice reading interrupt for the sub-camera according to the reading camera setting in #2. Setting value A. "Every specified time" allows you to set the interrupt cycle for the sub-camera. This setting may also be a setting for the ratio to the main camera. For example, in the example of control B in Figure 9C, the main camera (front camera) reads three times and the sub-camera (rear camera) reads once, which is a setting of 3:1, or 1 / 4.
[0119] Setting value B. "When an approaching object is detected" interrupts the voice reading when an object approaching user U1 is detected. Setting value C. "When a significant change is detected" interrupts the voice reading when an object with a significant change between images is detected. Setting value D. "When a specific object is detected" interrupts the voice reading when a specific object is detected. Setting value E. "When a danger is detected" interrupts the voice reading when a dangerous condition / object is detected in the surrounding situation.
[0120] The #7 "Reading Object Priority" item allows you to set the priority for the objects and conditions of voice reading. This priority is particularly the priority for voice reading of various objects that can be recognized and detected from an image. For example, there are objects and conditions similar to the settings A to E of #5 above, and these may be prioritized or sorted, or the priority may be set to several levels, such as high, medium, or low, for each object or condition. In addition, default settings for the system may be provided, and for example, "danger" may be set as the highest priority, followed by "approaching objects."
[0121] #8 "No change reading function" allows you to set whether or not to read out loud the state of no change when it is determined that there is not much change in the object between images (described below).
[0122] #9 "Reading aloud using different expressions" allows you to set whether or not to read aloud the text representing the object using different expressions when repeatedly reading aloud the same detected object (see below).
[0123] #10 "Safe Space Guide Function" allows you to set whether to provide a voice readout to guide the user to move to a safe space to avoid contact with the object, for example, when the object is approaching the user, based on the relative positions of the user and the object (described below).
[0124] #11 "History retention period" allows the user to set the period for which the voice reading history data is retained. The history retention period may be set separately for the device 1 and the server device 2. The options for the setting value may be, for example, forever, one year, three months, or no retention.
[0125] #12 "Other detailed settings" are detailed settings other than #1 to #11 above. For example, image resolution can be included. Image resolution may be set to, for example, high (600 dpi), medium (350 dpi), low (72 dpi), etc.
[0126] [Recorder function] One of the functions of this system is a recorder function related to image reading. According to user settings (#11 in FIG. 14), historical data of processing by the image reading function (including image capture 802, recognition 803, and voice reading 804 in FIG. 8) can be recorded and saved in a storage resource (device 1 or server device 2). This recorder function (in other words, a walking recorder) functions like a drive recorder in an automobile. That is, this function records the surrounding circumstances while the user is walking as a history, allowing the history to be checked later. The saved historical data can be output later (screen display or audio output). This function allows the surrounding circumstances at the time to be checked as evidence. The historical data may include the date and time, captured images, text of the recognition results, etc.
[0127] [Image object recognition and voice reading] FIG. 15A is a schematic explanatory diagram showing an example of object recognition and voice reading in an image captured by the camera 403 of the device 1. FIG. 15A schematically shows an example of recognition by the AI20 from a captured image and an example of changes in the object in the image. In particular, this example shows an example of an object approaching the user U1 (i.e., the device 1). The upper part is an image 1501 of a state at a first time point, and the lower part is an image 1502 of a state at a second time point. Assume that the object OB1 in the image is a pedestrian who is another person walking toward the user. The AI20 recognizes such objects / situations / changes as features in the image and converts them into text representing the object / situation / change. The AI20 expresses the result of its guess as to what the object OB1 is in text.
[0128] First, when a pedestrian (pedestrian 1001 in FIG. 10) is recognized / detected as object OB1 from image 1501, the object OB1 may be converted into text representing the object OB1, such as "pedestrian" (or "person"), and the "pedestrian" may be read aloud. Within image 1501, region r1 is a rectangular image region including object OB1 (OB1a). Point p1 is an example of position coordinates representing object OB1 / region r1. AI 20 of server device 2 outputs, as a response, information such as the position coordinates, size, and shape of the region of object OB1, and text representing object OB1.
[0129] Furthermore, if it is determined between images that the object OB1 is approaching the user, text describing the situation in which the object OB1 is approaching may be obtained and read aloud, for example, "A pedestrian is approaching."
[0130] In image 1502, region r2 is a rectangular image region that includes object OB1 (OB1b). Point p2 is an example of position coordinates that represent object OB1 / region r1. For example, in image 1502, the position (point p2) of region r2, which is a pedestrian that is object OB1, has moved and changed in size more significantly compared to the state of the position (point p1) of region r1 in image 1501. In the example of FIG. 15A, the amount of change in size is large.
[0131] The device 1 or the server device 2 may determine whether the object OB1 has approached the user based on the amount of change in size of the region (r1, r2) of the object OB1 between images. Alternatively, the device 1 or the server device 2 may determine the relative distance and speed of the object from the user based on the images, and determine that the object has approached when the relative distance falls within a predetermined range (threshold), as shown in FIG.
[0132] Furthermore, the determination of approach is not limited to a binary determination of whether or not there is approach, but may be a multi-value determination of the degree of approach.
[0133] In the example of FIG. 15A, the word "pedestrian" or the situation of a pedestrian approaching may be read out loud at the time of image 1501, or the word "pedestrian" or the situation of a pedestrian approaching may be read out loud at the time of image 1502.
[0134] The present system performs control so that reading out is performed with priority given to objects that are closer to the user. In other words, the present system performs control so that reading out is performed with priority given to objects that are closer to the user.
[0135] FIG. 15B shows another example. In image 1511, the detected object OB2 (OB2a) is assumed to be a pedestrian coming from the left side of the intersection and moving to the right. In image 1512, the detected object OB2 (OB2b) is assumed to be a pedestrian who has reached the center of the intersection. Compared to image 1511, in image 1512, the size of region r3 of object OB2 remains almost the same as region r4, but the position coordinates have changed significantly from p3 to p4. The device 1 or the server apparatus 2 may determine the movement status of object OB2 based on the amount of change in the position coordinates of region (r3, r4) of object OB2 between images.
[0136] From the change in the feature size increasing as in Figure 15A, it can be understood that the object is approaching the user in the forward or backward direction. From the change in the feature position coordinates as in Figure 15B, it can be understood that the object is moving, for example, left or right, crossing in front of the user.
[0137] Using the amount of change in the size, position, etc. of features corresponding to the object between the images as described above, the device 1 or the server device 2 can grasp the positional relationship between the user and the object. Based on this understanding, predetermined control is possible. Examples of predetermined control include determining an object as a target for voice reading when the amount of change in the object is large, or prioritizing an object as a target for voice reading when the degree of approach is large.
[0138] When controlling device 1 to give priority to reading out an object approaching the user's position, device 1 may determine the priority of reading out depending on the relative distance from the user, in other words, depending on the distance range as shown in Fig. 11. For example, multiple distance thresholds are set, such as 1 m, 3 m, 5 m, and 10 m. For example, if there are both an object that is 10 m away and an object that is 5 m away around the user at the same time, the closer object is given a higher priority for reading out and is read out first.
[0139] Alternatively, the system may control the volume of the voice reading to be louder for objects that are closer to the user (or objects that are closer relative to the user).Also, the system may control the volume of the voice reading to be louder for objects that are closer relative to the user (Figure 11).
[0140] Furthermore, the present system may apply not only control of changing the volume but also control of changing the tone according to the degree of approach. For example, when the degree of approach exceeds a threshold, the tone of the voice reading may be changed from a first tone to a second tone.
[0141] The system may also be controlled to add and output an alert sound or vibration depending on the degree of approach. The cycle of the alert sound output may be shortened depending on the degree of approach. For example, when the degree of approach exceeds a threshold, an alert sound or vibration may be added to the reading voice, or the reading voice may be changed to an alert sound or vibration.
[0142] In addition, if a non-visually impaired user is using smart glasses 1C or the like, a display of some kind that accompanies the voice reading, such as a display that indicates that an object is approaching, may be added to the display surface of the device 1.
[0143] The degree of approach may also be regarded as a degree of danger, which will be described later. That is, the greater the degree of approach, the greater the degree of danger, and voice reading or alert output may be applied according to the degree of danger.
[0144] The movement or change of the object (corresponding feature) is not limited to the above example, but can also be read aloud when a new object appears and is detected within the user's field of view (in the corresponding image), or when the direction of movement of the object changes (see below).In addition, when the recognition of the object by the AI 20 changes, it can also be read aloud.
[0145] [Noticeable changes in subject matter between images] Based on the captured images, the device 1 or the server device 2 identifies and detects portions of the image in time series where there has been a significant change in the surrounding situation / object, determines the portions as targets for voice reading, and controls the device to prioritize voice reading. Based on the recognition results between images, the device 1 or the server device 2 identifies and detects portions of the image where there has been a significant change as targets based on the amount of change in features between the surrounding situation at a first time point (an earlier time point) and the surrounding situation at a second time point (a later time point). For example, a portion moving differently is detected as an object with a significant change. The device 1 or the server device 2 then controls the device to prioritize the object corresponding to the portion with the significant change as a target for voice reading over other portions (i.e., portions with little change).
[0146] As a simple example, the voice reading of an object that has undergone a significant change may be a voice output of only words such as the name of the object. For example, words such as "person," "pedestrian," "bicycle," and "car" may be used. When using short text such as words, it is possible to read out a large number of objects per unit time. More specifically, this voice reading may be a voice output of a written description that describes the situation of the object. For example, written descriptions such as "A person is approaching," "A bicycle is approaching from behind," and "A car is crossing." When written descriptions are used, it is easier for the user to understand the surrounding situation.
[0147] Furthermore, when reading out an object that has undergone a significant change, the system may output a voice based on text describing how the object changed before and after the change. For example, the voice output may be something like, "The stopped car has started moving toward you," or "The traffic light has changed from red to green."
[0148] If the system is set to read out loud only objects that have undergone significant changes and not those that have not undergone significant changes, the same object that has not undergone significant changes will not be read out repeatedly. This reduces the amount of audio information and annoys the user, and may make it easier for them to understand the surrounding situation.
[0149] FIG. 16A illustrates an example of significant changes in an object between images captured by the camera 304. First, in image 1601 captured at a first time point, no particularly noticeable object (especially a person) is detected on the sidewalk ahead of the user. Next, in image 1602 captured at a second time point, a pedestrian appears as object OB3 on the sidewalk ahead of the user. This object OB3 is, for example, a pedestrian appearing in the image as if coming out of the road on the left onto the sidewalk where the user is located. The device 1 or the server device 2 detects such changes and differences between images and, based on the changes and differences, detects the object as having a significant change. The device 1 or the server device 2 then determines the detected object with a significant change as a priority target for voice reading. In the example of FIG. 16A, the text and voice reading of the recognition result for object OB3 may be, for example, "pedestrian" or "pedestrian coming from the left."
[0150] 16B shows another example. In this example, the direction of movement of an object changes significantly between images. In image 1611 at a first time point, the direction of movement of a pedestrian, represented by object OB4 (OB4a), is to the right (+X). In image 1612 at a second time point, the direction of movement of a pedestrian, represented by object OB4 (OB4b), has changed to a forward direction (-Y) coming toward the user.
[0151] The system also determines the amount of change in the direction of movement of the object corresponding to the features of the recognition result. The device 1 or the server device 2 detects such changes and differences between the images and detects the object OB4 as having a significant change based on the changes and differences. The device 1 or the server device 2 then determines the detected object OB4 having a significant change as the object to be prioritized for voice reading. In the example of FIG. 16B, the text and voice reading of the recognition result for the object OB4 may be, for example, "pedestrian" or "a pedestrian coming from the left is turning toward you."
[0152] In the above control example, if the server device 2, rather than the device 1, makes the judgment through analysis processing, the response information from the server device 2 to the device 1 may include the judgment result (for example, information about the distance to the target object) in addition to the text of the recognition result.
[0153] [Processing Sequence] 17 shows an example of a processing sequence of device 1 and server device 2 in the system of FIG. 1. In step S101, device 1 turns on the text-to-speech function. For example, the text-to-speech function may be turned on automatically according to a setting when device 1 is started, or the user may perform an operation input to turn on the text-to-speech function. Alternatively, device 1 may determine the state of device 1, such as its posture, and automatically turn on the text-to-speech function.
[0154] In step S102, the device 1 reads the setting information of the user U1 (FIG. 14) from the non-volatile memory 303 or the like. Furthermore, in step S102, the device 1 may determine a rule (also referred to as an analysis rule) for analysis (the processing of steps S3 and S4 in FIG. 4) or a control mode corresponding to the analysis rule, based on the read user setting information. The analysis rule is explained below, for example, in the case of the setting of #5 in FIG. 14 ("Determination condition for reading out from the main camera"). For example, the determination condition may be "B. approaching object" as the first priority, "D. specific object" as the second priority, and "A. object in the direction of movement" as the third priority, and the analysis rule may be such that analysis (the processing of steps S3 and S4 in FIG. 4) is performed based on these priorities / priorities. Furthermore, the control mode may be selected from the control modes shown in FIG. 13, for example.
[0155] In step S103, in response to the start of use of the image reading function, device 1 turns on the video input of camera 403 and inputs the video (time-series image frames) of camera 403.
[0156] In step S104, the device 1 starts communication for cooperation with the server device 2. The device 1 establishes a communication connection with the server device 2 using the communication unit 330 (FIG. 7A). The device 1 transmits the user setting information (or information obtained by processing the same) of step S102 to the server device 2 over the established communication. At this time, the device 1 may also transmit user information of the user U1, device information of the device 1, and other information to the server device 2. Furthermore, if the device 1 has determined an analysis rule or a control mode in step S102, the device 1 may transmit information on the analysis rule or control mode to the server device 2 in step S104.
[0157] In step S105, the server device 2 uses the communication unit 430 (FIG. 7B) to communicate with the device 1 for cooperation and receives information such as user setting information, analysis rules, and control modes from the device 1. Based on this information, the server 2 determines the content of the analysis process to be performed by the server device 2 (the processes of steps S3 and S4 in FIG. 4). In step S105, if the information received from the device 1 does not include an analysis rule or a control mode, the server device 2, rather than the device 1, may determine the analysis rule and the control mode based on the user setting information from the device 1.
[0158] In step S106, the device 1 monitors whether or not there is an operation input by the user U1, and determines whether or not there is a stop instruction, etc. A stop instruction is an instruction to end or pause the image reading function (particularly the series of processes from image capture to voice reading). If there is a stop instruction, etc., the device 1 transmits the stop instruction, etc. to the server device 2, and stops the processing in the device 1. In step S107, the server device 2 stops the corresponding processing in the server device 2 in response to the stop instruction, etc. from the device 1. If there is no stop instruction, etc., the processing from step S108 onwards is similarly repeated as a loop. Note that if the processing is paused in accordance with the stop instruction, and then a start instruction is input based on the operation input by the user U1, the device 1 and the server device 2 resume the processing.
[0159] In step S108, the device 1 waits for a predetermined time according to the user settings. In this embodiment, the device 1 captures images at a predetermined cycle based on the setting of #3 in FIG. 14 (first time interval), and performs voice reading at a predetermined cycle based on the setting of #4 (second time interval). Step S108 is a standby time for the image capture. Capture is possible at each cycle of the main camera and the sub camera.
[0160] In step S109, device 1 captures and acquires images from the input video of camera 403 and stores them in memory. When multiple cameras are used, the timing of image capture for each camera may be controlled in a predetermined manner, for example, sequential capture as shown in FIG. 9A. Image processing unit 305 in FIG. 7A controls such image capture and stores the acquired images in image storage unit 306. Note that if the device is equipped with predetermined hardware and does not pose a problem with the processing load, simultaneous capture may be performed by multiple cameras. Furthermore, when capturing images, image information such as the date and time of capture, image identification information, and camera identification information is also added to each image data.
[0161] In step S110, the device 1 checks and determines the communication status with the server device 2. The device 1 determines whether communication with the server device 2 is possible or impossible, or whether the communication status is good or bad. If the communication status with the server device 2 is bad or impossible, the device 1 temporarily cuts off communication with the server device 2, or temporarily suspends transmission of requests and the like while maintaining the communication connection. Examples of a bad or impossible communication status include congestion on the communication network 9 and a heavy processing load on the AI 20 of the server device 2. When the device 1 cuts off or stops communication with the server device 2, it stores in memory the history of processing up to that point.
[0162] In step S111, the server device 2 temporarily disconnects communication with the device 1 as necessary, based on the communication status between the device 1 and the server device 2 in step S110, or temporarily suspends transmission of responses and the like while maintaining the communication connection. When the server device 2 disconnects / stops communication with the device 1, it stores a history of processing up to that point in memory or a database. This history is retained for a predetermined period or more. The system sets a control flag depending on the communication possible / not possible state in steps S110 and S111, and performs processing from step S112 onwards depending on the flag.
[0163] If communication is impossible / poor in step S110, the system controls the server device 2 not to perform analysis processing, including recognition processing of the AI 20, and controls the device 1 to perform analysis processing in step S112. If the device 1 is equipped with a function for performing analysis processing on its own, it can perform the analysis processing on its own in step S112. If the device 1 is not equipped with a function for performing analysis processing on its own, it cannot perform the analysis processing on its own in step S112, and both the device 1 and the server device 2 stop the analysis processing. In this case, for example, the device 1 notifies the user U1 of the current status and waits for communication with the server device 2 to be restored.
[0164] In step S112, if a function for performing simple analysis processing is implemented on the device 1 side, the device 1 may perform the simple analysis processing. Normally, the server device 2, which has abundant computational resources, performs the analysis processing, but if the analysis processing on the server device 2 side is unavailable due to a communication failure, the device 1 performs the simple analysis processing (described later). Generally, the device 1 side has more limited computational resources than the server device 2 side on the cloud computing system 3. Therefore, the simple analysis processing on the device 1 side refers to processing that is simpler than the analysis processing on the server device 2 side. If the device 1 side has abundant computational resources, it is also possible to perform the analysis processing only on the device 1, as shown in FIG. 3.
[0165] If communication is possible / good in step S110, the system controls in step S113 to perform analysis processing, including recognition processing by the AI 20, on the server device 2 side, rather than on the device 1 side. In step S113, the device 1 transmits an analysis request and an image (the captured image and additional image information in step S109) to the server device 2 for analysis processing on the server device 2 side. Note that when transmitting the request, supplemental information may be attached in addition to the image. The supplemental information is information that can be used in the analysis processing, and may include information that can be detected by sensors of the device 1 (for example, the position detection unit 311 and the orientation detection unit 312 in FIG. 7A), such as position, orientation, direction, and acceleration.
[0166] In step S114, the server device 2 receives the analysis request and image from the device 1 and stores them in memory. In response to the request, the server device 2 starts analysis processing, including recognition processing by the AI 20. First, the server device 2 recognizes the image using the AI 20 to extract features related to the object from the image. In other words, the object is detected from the image. Furthermore, the server device 2 converts the object associated with the features through the recognition into text representing the object. As an output of the AI 20, feature information including text information associated with the features, position coordinates, size, etc. is obtained. The above-mentioned processing for each captured image is repeated in a similar manner at a predetermined cycle.
[0167] Next, in step S115, the server device 2 detects changes (in other words, change points, etc.) between the image at the current time point and the image at the previous time point for the group of captured images in time series. The change points correspond to changes in the position coordinates, size, shape, direction, etc. of the object (in other words, image area) associated with the feature, for example.
[0168] Next, in step S116, the server device 2 determines the distance, speed, and direction of change (in other words, the direction of movement) of the object in the image and between the images, depending on the analysis rule, control mode, etc. to be applied. The server device 2 may determine the distance, speed, and direction of change by analyzing the image. The server device 2 may also determine significant changes, the degree of approach, the degree of danger, and the safe space, etc., related to the object. Note that, in this embodiment, the server device 2 makes this determination in step S116, but such a determination may also be made on the device 1 side after obtaining information about the object from the server device 2 side.
[0169] In step S117, the server device 2 determines the image, object / surrounding situation, text, etc. to be read aloud based on the results of the analysis processing and determination processing up to step S116. The server device 2 selects and determines the object, etc. to be read aloud based on the characteristics, change points, distance, speed, change direction, etc. from the image in accordance with the analysis rules, control mode, etc., in accordance with the conditions of the analysis rules, for example, in accordance with the priority of #7 in Fig. 14. The text associated with the determined object is also determined.
[0170] In step S118, the server device 2 transmits a response to the device 1 as the analysis result, including information such as text about the object to be read aloud determined in step S117. When transmitting the response from the server device 2 to the device 1, the server device 2 may also add processing result information (e.g., priority / priority order for reading aloud multiple objects) of the analysis process or determination process performed by the server device 2, depending on the analysis rule / control mode, etc. In step S119, the device 1 receives the response from the server device 2, and performs reading aloud by converting the text to speech based on the text of the object to be read aloud and the additional information, and outputting the speech from a speech output device such as a speaker. If the target text has priority / priority order attached, the device 1 controls the reading aloud accordingly. For example, the device 1 reads aloud multiple texts in the order of priority. Furthermore, for example, the device 1 changes the volume of the reading aloud depending on the degree of proximity. Furthermore, for example, if an alert output instruction is attached, the device 1 outputs an alert sound from a speech output device.
[0171] The voice reading in step S119 is set to be performed at a predetermined cycle (second time interval). Therefore, the system controls the processes from step S112 to step S119 to be performed at the cycle. If there is no object to be read out at the periodic timing of the voice reading, for example, if there is no significant change, the response in step S118 will be a response indicating that there is no object to be read out (no change) and that there is no target text.
[0172] 17 is a case where a series of processes from image capture to voice reading are automatically executed at periodic timing, but if a read-aloud instruction (a read-aloud re-instruction, described later) is input by user U1 at any timing, the following control may be performed. That is, when device 1 receives the read-aloud instruction and the communication state with server device 2 is available / good, device 1 captures an image and transmits the image and a request corresponding to the read-aloud instruction to server device 2. In response to the request, server device 2 similarly performs an analysis process and transmits a response including the text of the analysis result to device 1. Device 1 performs voice reading based on the text of the response.
[0173] In this embodiment, as in the above example, a series of processes from image capture to voice reading is automatically executed at a predetermined cycle from multiple cameras of device 1. There is basically no need for a user to perform instruction operations such as image capture, which reduces the effort and is highly convenient.
[0174] This system is also capable of reading out all objects recognized in an image, as long as time and computational resources allow. However, reading out a large number of objects at once may be difficult for the user to understand and may confuse them. Therefore, in this embodiment, the object to be read out is determined according to priority, proximity, etc., during the automatic image reading cycle, and the most important objects are read out first. This allows for easy and convenient support for the user. The user can use different user setting information, control modes, etc. depending on the usage scenario, app, etc.
[0175] Furthermore, if no object is detected in the time series or if there is little change in the object, the number of voice readouts may decrease or the same voice readout may be repeated. In such cases, the user may feel uneasy. Therefore, this embodiment also supports voice readout in response to a user's input of a readout instruction at any time, and, as will be described later, has a function to read out that there is no change in the surrounding situation when there is no change. This allows for easy-to-understand and suitable support for the user.
[0176] 17, in step S118, the server device 2 may convert the text representing the object into voice data for reading aloud, and transmit a response including the voice data to the device 1. In this case, the device 1 does not need to convert the text into voice.
[0177] [Basic Control] FIG. 18 is an explanatory diagram regarding the basic control of automatic image capture and voice reading in time series. Voice reading cycle timing 1801 shows, for example, time points 1 to 7. Images 1802 at each time point are example images at the timing of image capture corresponding to the time point in the voice reading cycle. Voice reading 1803 shows an example of the audio content of the voice reading at that time point. Repetition 1804 indicates the number of times the voice reading is repeated in time series, etc.
[0178] At time point 1, it is assumed that no particular object is detected in the image. At time point 1, no voice reading is performed (illustrated as "none"). At time point 2, as in FIG. 16A, a pedestrian 1811 is detected as an object 1811 as an example of a significant change. At time point 3, the pedestrian 1811 is on the sidewalk ahead of the user U1. At time points 4 and 5, the state is the same as at time point 3, with almost no change in the pedestrian 1811. At time points 6 and 7, the pedestrian 1811 is moving closer to the user U1. At each time point after time point 2, a "pedestrian" is detected in the image as the object 1811. Therefore, in the control example of FIG. 18, at each time point after time point 2, a voice reading such as "pedestrian" (or specifically "pedestrian approaching") is output, and the same voice reading "pedestrian" is output multiple times. As for the repetition, time point 2 is the first repetition, and time point 7 is the sixth repetition.
[0179] As in the above example, the user U1 can recognize that there is a "pedestrian" ahead from the voice reading at each point in time, without performing any particular operation input.
[0180] The device 1 stores data on the history of processes, including image capture and voice reading, related to the above-described control in a storage resource. The device 1 stores the history data in the memory of its own device or the server device 2 up to an upper limit, such as a predetermined time, a predetermined number of times, or a predetermined data amount. If the history data exceeds the storage limit, the device 1 or the server device 2 overwrites and erases the data, starting with the oldest.
[0181] [Control not to read aloud if there is no change] Next, FIG. 19 is an explanatory diagram regarding control that does not perform voice reading when there is no change in the object, based on the basic control of FIG. 18. The example of "pedestrian" as object 1811 is the same as in FIG. 18. In the transition from time point 3 to time point 4, in other words, between the images, there is almost no change in "pedestrian" as object 1811. Therefore, in this control example, at time point 4, device 1 does not perform voice reading (shown as "none"). Similarly, in the transition from time point 4 to time point 5, there is no change, so voice reading is not performed.
[0182] The image reading system may determine that there is no change in the object between images if the change in the object (for example, the change in the position coordinates or size of features) is less than a threshold value based on the recognition results of AI20.
[0183] In the transition from time point 5 to time point 6, the "pedestrian" as the object 1811 approaches the user U1. The system determines that there is a change in the "pedestrian" as the object 1811 from the amount of change between images. Therefore, at time point 6, for the object 1811 with the change, a voice readout is output, for example, "pedestrian" (or specifically, "pedestrian approaching"). Considering the number of repetitions from time point 2, time point 6 is the third time. Similarly, at time point 7, it is determined that there is a change, so a fourth voice readout is performed.
[0184] As in the above example, if there is no change in the object, there is no voice reading, so the volume and frequency of voices heard by the user U1 are reduced.
[0185] [Control to read out a message saying there is no change if there is no change] In Figures 18 and 19, for example, from time point 3 to time point 5, there is almost no change in the objects surrounding user U1. If the object remains unchanged for a long period of time, a voice readout may be performed periodically to inform the user that there is no change, as in the control example of Figure 20. For example, the system determines whether the degree of change (e.g., amount of change) in the object remains below a threshold for a predetermined period of time or longer. The time count may be performed in units of a predetermined period of time or images. The predetermined period of time for this determination is configurable. If the state continues for a predetermined period of time or longer, the system determines to perform a voice readout indicating that there is no change. Device 1 reads out a voice such as "There is no change" to inform the user that there is no change.
[0186] In the example of FIG. 20, at time points 2 and 3, the "pedestrian" as the object 1811 in front of the user U1 is detected as having changed, and therefore "pedestrian" is read out as described above. At time point 4, it is determined that there is no change in the "pedestrian" as the object 1811. If this state of no change continues for, for example, one cycle of the read-out cycle, it is decided that a read-out indicating that there is no change is to be performed. Therefore, at time point 4, "no change" is read out. Similarly, at the next time point 5, it is determined that there is no change in the "pedestrian," and this state of no change continues for one cycle, so "no change" is read out.
[0187] In another example, if the specified time during which no change continues is set to two cycles of the voice reading period, at time point 4, no voice reading will be performed, and at time point 5, the state of no change has continued for two cycles, so the voice reading will say "No change."
[0188] As in the above example, if there is no change in the object, a voice readout is made to the effect that there is no change, so that the user U1 is less anxious when there is no voice readout.
[0189] [Read aloud instruction function] FIG. 21 is an explanatory diagram showing a specific example of the read-again instruction function. User U1 can input an instruction to read aloud again (in other words, a forced read-again instruction) to device 1 at any timing. The user U1 can input an instruction to read aloud again by voice using a microphone or the like (microphone / voice input unit 316 in FIG. 7A), or by pressing a predetermined button on the remote control 1c in FIG. 5A, for example. For example, pressing a predetermined button on the remote control 1c indicates an instruction to read aloud again.
[0190] When this read-again instruction is input, device 1 always executes a voice readout of the object at that time and outputs some kind of audio. In response to the operation input of the read-again instruction, device 1 captures an image from the input video, causes AI 20 to recognize the image, and performs a voice readout of the object / surrounding situation. Even if there is no change in the object at the time of input of the read-again instruction, device 1 performs a voice readout that conveys the object or surrounding situation.
[0191] The images at each time point in FIG. 21 are generally similar to those in FIG. 18, etc. At time points 2 and 3, a "pedestrian" is detected as the object 1811 as a significant change, and the device 1 automatically outputs, for example, "pedestrian" as a read-out. From time point 4 onwards, it is assumed that there is almost no change in the "pedestrian" as the object 1811. At time point 4, it is determined that there is no change in the "pedestrian" as the object 1811, so the device 1 does not perform read-out unless there is a particular action from the user U1. At time point 5, it is similarly determined that there is no change, so the device 1 does not perform read-out unless there is a particular action from the user U1.
[0192] On the other hand, for example, assume that the user U1 inputs an instruction to read aloud again just before time point 5. In this case, the device 1 follows the instruction to read aloud again and performs a voice readout of the object 1811 at time point 5. The device 1 outputs, for example, "pedestrian" as the voice readout. Alternatively, as a modified example, similar to FIG. 20 , since there is no change in the object, the device 1 may output "No change" or the like.
[0193] At time point 6, similarly, it is determined that there is no change in the object 1811, so the device 1 does not perform voice reading unless there is a particular action from the user U1. Assume that the user U1 again inputs a read-aloud instruction at time point 7. In this case, the device 1 similarly follows the read-aloud instruction and performs voice reading at time point 7, outputting, for example, "pedestrian" (or "no change").
[0194] In the above example, the timing at which the instruction to read aloud again is input is almost the same as the periodic timing of the voice reading, such as time point 5 or time point 7, but these timings may be different. In the above example, at time point 5 or time point 7, device 1 basically transmits a request to server device 2 in response to the instruction to read aloud again, receives a response from server device 2, and performs voice reading using the response. Even if there is no change in the object compared to the results of the previous recognition and voice reading, device 1 performs voice reading again based on the current recognition.
[0195] As a variant, when a reading instruction is input again (especially when it is different from the periodic timing of the voice reading), the device 1 may omit communication with the server device 2 and perform the voice reading using the results of the previous voice reading (e.g., the same text) that have been stored as history up to that point.
[0196] As in the above example, in addition to the automatic voice reading by this system, user U1 can check the surrounding situation by having voice reading performed instantly at the timing of his / her choice. Furthermore, the above-mentioned read aloud instruction function can be effectively used in combination with the voice reading function with different expressions shown in Figure 22 below.
[0197] [Voice reading function with different expressions (1)] As in the above control example, the readout is repeated multiple times in a time series by automatic periodic readout or by readout in response to a user U1's instruction to read it again. When the surrounding circumstances / object are almost the same at each point in time and there are no significant changes, the readout mode as in each control example is possible, and the same audio output may be repeated, but as another example, the following may also be used. In another example, when a similar surrounding circumstances / object is read out multiple times, control may be exercised so that different text is read out each time, explaining it from a different perspective or using different expressions.
[0198] FIG. 22 is an explanatory diagram showing a specific example of the function of text-to-speech with different expressions. The situation of the object in the image in FIG. 22 is the same as that in FIG. 16B. In FIG. 22, a "pedestrian" coming out from the left side of user U1 is detected as object 2201 in image 2200. Assume that there is almost no change in the situation of object 2201 in the time-series images. Assume that the function of text-to-speech with different expressions is set to ON in device 1. For example, at time point 1, device 1 detects a significant change in the response from server device 2 and outputs, for example, "A pedestrian is coming from the left" as the first automatic text-to-speech.
[0199] Next, time point 2 is assumed to be the timing of a predetermined cycle of voice reading, or the timing when user U1 performs an operation input for a read-out instruction again as in FIG. 21. At this time point 2, device 1 executes a second voice reading for "pedestrian," which is object 2201. When repeatedly reading out the same object 2201 that has not changed, device 1 controls the reading out to be performed using a different expression from the previous time (first time). Based on the response from server apparatus 2, device 1 outputs, for example, "There is a person diagonally ahead of you to the left," in the second voice reading.
[0200] Next, time point 3 is assumed to be the timing of a predetermined cycle of voice reading or the timing when user U1 performs an operation input for a read-out instruction again. At this time point 3, device 1 executes a third voice reading for object 2201. Device 1 controls to perform voice reading using an expression different from the first and second times. Based on the response from server device 2, device 1 outputs, for example, "A person is coming out of the road on the left" in the third voice reading.
[0201] To realize such a function, the device 1 may cause the AI 20 to perform recognition again at each timing of reading aloud, and acquire text in a different expression from the AI 20. Alternatively, the device 1 may acquire and store text in multiple expressions collectively from the AI 20 when the AI 20 recognizes the text in one request-response, and select one expression from the acquired multiple expressions and output it aloud at each reading aloud.
[0202] Another specific example is as follows. Suppose that image recognition detects a situation in which a man walking in front of the user pushing a bicycle approaches the user. At a first point in time, the first voice readout is, for example, "Man and bicycle." At a second point in time, the second voice readout is expressed differently, for example, "A man is walking pushing a bicycle." At a third point in time, the third voice readout is expressed even differently, for example, "A man pushing a bicycle is approaching."
[0203] As in the example above, there can be multiple ways to describe a similar situation / object. This feature reads out different sentences at different times, making it easier for the user to understand the situation.
[0204] [Voice reading function with different expressions (2)] Furthermore, the following control is also possible as a modified example of the function of reading aloud using different expressions. The system may perform reading aloud by adding or synthesizing the current recognition result to the content read aloud of the previous recognition result at each timing of reading aloud, such as automatically, or in response to an instruction to read aloud again. In other words, if the system determines at each timing that there has been a change in the object, it may perform reading aloud of both the audio representing the previous object and the audio representing the part with the current change.
[0205] As an example, at a first time point, in response to a first re-reading instruction, the AI20 recognizes a "pedestrian" as an object in front of the user U1 and reads out "pedestrian." Next, at a second time point, in response to a second re-reading instruction, the AI20 recognizes a "bicycle" (or a "person riding a bicycle") as an object separate from the previous "pedestrian." In this case, the device 1 outputs a combined voice reading of both the previous "pedestrian" and the current "bicycle," such as "pedestrian and bicycle." Next, at a third time point, in response to a third re-reading instruction, the AI20 recognizes a "dog" as another object separate from the previous "pedestrian" and "bicycle." The device 1 outputs a combined voice reading of these objects, such as "person, bicycle, and dog." In this example, the nouns of multiple objects are connected in parallel and read aloud, but this is not limited to this. It is also possible to make a sentence that expresses the situation of one or more objects in more detail, for example, "A person riding a bicycle is walking a dog."
[0206] [Voice reading function with different expressions (3)] The following modification is also possible. The system may read out a different representation of the object in response to updating / correcting the recognition by the AI20 at each timing of voice reading, such as an automatic cycle or a re-read instruction. For example, at a first time point, in response to a first re-read instruction, a "pedestrian" is detected in front of the user U1, and, for example, "pedestrian" is output as voice. At a second time point, in response to a second re-read instruction, the AI20 re-recognizes the previous object, "pedestrian," and updates / corrects the recognition, and more accurately detects it as a "bicyclist," and, for example, outputs "bicyclist." Furthermore, at a third time point, in response to a third re-read instruction, the AI20 re-recognizes the previous object, "bicyclist," and, more accurately, detects it as a "bicyclist with a dog," and, for example, outputs "bicyclist with a dog."
[0207] Furthermore, when device 1 performs a voice readout based on the result of correcting the previous content to the current content as described above, the current voice readout may be a voice readout that notifies the user of the correction. For example, in the second voice readout, the voice output may be, "Correction: This is not a pedestrian, but a person riding a bicycle."
[0208] [Example of front and rear camera control] FIG. 23 shows an example of control of the front and rear cameras, assuming that multiple cameras are used as in FIG. 9A, etc. FIG. 23 particularly shows an example of control that switches between control in which the front camera is the main camera and control in which the rear camera is the main camera. This control enables a function such as interruption. For example, the front and rear cameras are used in accordance with the control mode (M1) in FIG. 13 described above. In FIG. 23, as an example of a situation, assume that a user U1 is walking on a sidewalk 901, and there is a pedestrian 2301 in front of him and a bicycle 2302 behind him. Assume that the device 1 of the user U1, for example, a smartphone 1A, has a front camera C1 and a rear camera C2, as in FIG. 9A.
[0209] Under normal circumstances, this system performs control 1 shown in the figure, for example, by frequently prioritizing the capture, recognition, and reading of images in the forward direction (+Y) using the front camera C1. Control 1 corresponds to control B in FIG. 9C, and the processing ratio between the front and rear cameras is set to 3:1. During control 1, changes in the pedestrian 2301 can be detected with high accuracy from the image taken by the front camera C1, and a voice reading can be made, for example, saying "A pedestrian is approaching."
[0210] The system automatically switches from Control 1 to Control 2 as a result of a predetermined judgment or the like. For example, during Control 1, at a periodic timing such as time point 4, an object located behind the user U1 is detected from the image captured by the rear camera C2. For example, a bicycle 2302 moving toward the user U1 can be detected as a significant change. For example, when an object with a significant change in the rear direction is detected from the image captured by the rear camera C2, the device 1 switches to Control 2.
[0211] Control 2 shown in the figure prioritizes the frequent capture, recognition, and reading of images in the rear direction (-Y) using rear camera C2. Control 2 corresponds to control D in FIG. 9C, and the processing ratio between the front and rear cameras is set to 1:3. With control 2, changes in the bicycle 2302 can be detected with high accuracy from the image captured by rear camera C2, and a voice reading can be performed, for example, saying "A bicycle is approaching."
[0212] The above control example of switching from Control 1 to Control 2 allows the user U1 to have the function and effect of interrupting the recognition of the rearward direction with the recognition of the forward direction. Automatic switching from Control 2 to Control 1 can also be achieved in a similar manner. When in Control 2, if an object with a significant change is detected in the image from the front camera, the system switches to Control 1. In other words, this function allows the recognition and voice reading of the sub-camera to be realized as an interruption to the recognition and voice reading of the main camera. Similar control is possible when using front, rear, left, and right cameras.
[0213] As a modified example, when objects are detected both in the forward and backward directions, the system may compare the degree of change and priority of the objects in front and behind the vehicle to determine which of the two is more important, in other words, which object requires more attention, and select between Control 1 and Control 2 based on the determination result. In the illustrated example, when a pedestrian 2301 and a bicycle 2302 are detected, it is determined that the bicycle 2302 has a higher priority, for example, based on the relative distance and speed between them. In that case, the system switches to Control 2 or maintains Control 2 and gives priority to reading out the information about the bicycle 2302.
[0214] As a modified example, a control that treats the front and rear directions equally, such as the state A in Fig. 9A or the control C in Fig. 9C, may be interposed between control 1 and control 2. For example, the present system normally maintains control C, and transitions to control B when an object with a significant change is detected in the image from the front camera, and transitions to control D when an object with a significant change is detected in the image from the rear camera.
[0215] In a modified example, control 1 in FIG. 23 may be replaced with control A in FIG. 9C described above, and control 2 may be replaced with control E described above, and switching may be performed between control A and control E. However, in this case, for example, when control A is in use, there is no detection in the image from the rear camera, so the predetermined judgment for switching needs to be based on a different condition. An example of such a condition is when, when control A is in use, a state in which no object with a significant change is detected in the image from the front camera continues for more than a predetermined time, and then switching to control E is performed. Alternatively, the front and rear controls may be switched simply in response to a button operation input or a voice input by user U1.
[0216] [Controlling the front and back thresholds to be different] FIG. 24 shows a control example for determining significant changes in an object (when the object moves differently from previous images) using front and rear cameras capturing images in the front and rear directions relative to user U1. In particular, the control example in FIG. 24 illustrates a case in which different thresholds are used for the images and recognition from the front and rear cameras of device 1. This control example is based on the control of the relative distance and speed between user U1 and the object, as described above (FIG. 11). In FIG. 24, for images captured by front camera C1 capturing images in front of user U1, a distance threshold TD1 and a speed threshold TV1 are used as the thresholds for determination. On the other hand, for images captured by rear camera C2 capturing images behind user U1, a distance threshold TD2 and a speed threshold TV2 are used as the thresholds for determination. Situation A in FIG. 24 is an explanatory diagram for determining relative distance, and situation B is an explanatory diagram for determining relative speed.
[0217] For example, similar to Control 1 in FIG. 23, assuming that the front camera C1 is the main camera and the rear camera C2 is the sub - camera, and imaging, detection, and voice reading are performed more frequently with priority given to the front over the rear. In this case, from the perspective of frequency, objects in the rear are less likely to be detected than those in the front. Therefore, the thresholds (TD2, TV2) for determining voice reading of objects from the image of the rear camera C2 may be set as different thresholds to make them easier to detect than the thresholds (TD1, TV1) for determining voice reading of objects from the image of the front camera C1.
[0218] In the example of Situation A, for rear - priority, the rear distance threshold TD2 is set to a larger value than the front distance threshold TD1 (TD1 < TD2). In this example, TD2 is about twice TD1. In Situation A, assume that a pedestrian 2401 is moving in the - Y direction from the front of user U1, and another pedestrian 2402 is moving in the + Y direction from the rear of user U1. On the front side, based on the analysis of the image of the front camera C1, when the distance D2401 from the pedestrian 2401 is within the distance threshold TD1, the pedestrian 2401 is determined to be the object of voice reading. On the rear side, based on the analysis of the image of the rear camera C2, when the distance D2402 from the pedestrian 2402 is within the distance threshold TD2, the pedestrian 2402 is determined to be the object of voice reading. For example, even if the distances D2401 and D2402 are similar, due to the different distance thresholds, the rear pedestrian 2402 is detected preferentially first, and voice reading (e.g., "A pedestrian is approaching from behind") is performed.
[0219] In situation B, it is assumed that user U1 is stationary and has a speed v1=0, a pedestrian 2401 in front is moving toward user U1 in the -Y direction at a speed v2A, and a pedestrian 2402 behind is moving toward user U1 in the +Y direction at a speed v2B. In the example of situation B, a rear speed threshold TV2 is set to a smaller value than a front speed threshold TV1 to give priority to the rear (TV1>TV2). On the front side, based on an analysis of an image from the front camera C1, if the relative speed of the pedestrian 2401 with respect to speed v1A is equal to or greater than the speed threshold TV1, the pedestrian 2401 is determined to be a target for voice reading. On the rear side, based on an analysis of an image from the rear camera C2, if the relative speed of the pedestrian 2402 with respect to speed v2B is equal to or greater than the speed threshold TV2, the pedestrian 2402 is determined to be a target for voice reading. For example, even if the speeds v1A and v1B are approximately the same, the speed thresholds are different, so the pedestrian 2402 behind is detected first and then read out. Similar control is possible using the relative speeds even when the user U1 is moving.
[0220] Furthermore, when recognition and voice reading are performed equally using the front and rear cameras, as in state A in Fig. 9A, the threshold control as shown in Fig. 24 can be similarly applied. For example, if user U1 is not visually impaired, taking into consideration that it is easy to see the front but difficult to see the rear, priority may be given to the rear, and the threshold for the rear side may be set to be more sensitive so that objects at the rear can be detected with higher sensitivity than those at the front.
[0221] Furthermore, the above-mentioned determination of relative distance and determination of relative speed may be used alone or both may be used in combination. In the case of combined use, for example, an object may be determined as a target for voice reading if it satisfies an AND (logical product) condition using both thresholds. Similar threshold control can be applied to each direction when front, rear, left, and right cameras are used. The above-mentioned control thresholds can be set in various ways depending on the user, app, usage scenario, etc. Various thresholds can be variably set using the user setting function, and a mode corresponding to each setting can be selected and used.
[0222] [Support for the visually impaired] FIG. 25A is an explanatory diagram of a usage scenario in which a visually impaired person is supported. For example, this is a YZ plane view of a situation in which a visually impaired user U1 is standing still or walking on a sidewalk 2500. For example, as in FIG. 6B and other figures, a device 1 such as a hat 1D may have cameras (in this example, a front camera C3 and a rear camera C4) that capture images in the front and rear directions. It is preferable that the front camera C3 and other cameras have a capture range that can capture images up to the feet of the user U1. The user U1 walks while tapping the road surface, such as the sidewalk 2500, with a white cane, for example. In Situation 1, there is an obstacle 2501 on the road surface in front of the user U1. Even in this case, the obstacle 2501 is captured based on, for example, an image from the front camera C3, recognized as an object by the AI 20, and converted into text (e.g., "obstacle" / "situation with an obstacle ahead"). The device 1 then reads out the object aloud (e.g., "There is an obstacle ahead"). This will help visually impaired people walk.
[0223] In situation 2, a person 2502 is approaching the user U1 from behind. Even in this case, the person 2402 is captured, for example, based on an image from the rear camera C4, recognized as an object by the AI 20, and converted into text (e.g., "person" / "a situation in which a person is approaching from behind"). The device 1 then reads out the object aloud (e.g., "A person is approaching from behind"). In particular, the system may be controlled to increase the degree of approach or danger depending on the distance and speed of the approaching object and to output an additional alert. For example, a voice or alert sound saying "Please be careful" is output. This allows for more effective support for the walking of visually impaired people.
[0224] For example, in the case of a visually impaired user who cannot see clearly both in front and behind, it is desirable to control multiple cameras so that both the front and rear are read aloud. In this case, for example, control with an equal ratio between the front and rear as shown in Fig. 9A above may be applied, or control that switches between the front and rear controls may be applied as shown in Fig. 23.
[0225] [Nighttime Road Support] FIG. 25B is an explanatory diagram of night road assistance as one usage scenario. For example, on a night road 2510, images from the front and rear cameras of a smartphone 1A or the like are used as a device 1 of a user U1. The system also uses a rear camera C2 to detect whether a suspicious person (person) 2511 is following the user U1 from behind, and performs a voice readout to warn the user (e.g., "Someone is approaching from behind," "Be careful"). In addition to the voice readout, an alert sound may be output, the device 1 may vibrate, or the volume of the voice readout may be increased above normal. In addition, the brightness of the light emitted from the screen of the smartphone 1A or the like may be temporarily increased. For example, increasing the volume of the voice readout can have the effect of intimidating the suspicious person 2511.
[0226] The system may determine the degree of suspiciousness of a suspicious person (person) 2511, which is an object detected behind the user U1, and change the manner of voice reading (for example, volume, tone, alert, etc.) depending on the degree of suspiciousness. Methods for determining the degree of suspiciousness include, for example, judging situations such as a person detected behind following the user U1 while maintaining a certain distance between them, or appearing to be hiding.
[0227] Furthermore, when the device 1 is applied to assistance for road users at night, the camera 403 of the device 1 may be an infrared camera that can capture images even in dark environments.
[0228] [Features for "walking while using a smartphone"] FIG. 25C is an explanatory diagram of a function for supporting "walking while using a smartphone" as one usage scenario. In a generally safe situation, such as a sidewalk, user U1 is walking while searching for a destination while holding, for example, a smartphone 1A as device 1 and viewing a map using a map app. Images captured by the front and rear cameras of smartphone 1A are used to provide user U1 with a voice readout of objects from the surrounding image. For example, various objects listed on the map, particularly facilities and intersections along the route to the destination, may be detected and read out loud. To ensure safety while "walking while using a smartphone," objects approaching the user (such as people, bicycles, and cars) are also detected and read out loud in the same manner as in the control described above. Approaching objects are given priority over facilities and other objects. In the illustrated example, "intersection A" and "convenience store A" are detected in the direction of user U1's travel and read out loud. Furthermore, when a bicycle 2521 approaches user U1, the bicycle is detected and a voice readout is provided, such as "A bicycle is coming from the right. Be careful." In this way, safety can be increased while walking while using a smartphone, and it can also be linked to map apps and other applications.
[0229] [Vehicle driving assistance] FIG. 25D is an explanatory diagram of vehicle driving assistance as one usage scenario. For example, a user U1 is riding an electric bicycle 2531. While driving the electric bicycle 2531, the camera 403 of a device 1, such as smart glasses 1C or a smartphone 1A worn by the user U1, captures images of the surroundings, detects objects, and performs voice reading. For example, if there are objects (people, bicycles, cars, etc.) approaching the user U1 on the electric bicycle 2531 in directions such as the front, back, left, and right, voice reading of those objects is performed preferentially. This can assist driving and contribute to traffic safety.
[0230] [Example of control depending on whether server communication is possible or not] FIG. 26 shows an example of control depending on whether communication with the server device 2 is possible, and shows an example of control related to step 110 in FIG. 17 described above. State A is an example of control when communication with the server device 2 is possible, in which analysis processing is performed on the server device 2 side. The device 1 (e.g., smart glasses 1C) applies, for example, control A1 to the four cameras (C10, C20, C30, C40) in the front, rear, left, and right directions. Control A1 sets equal ratios for the four directions and sequentially transmits requests to the server device 2 to perform analysis processing including recognition by AI20. For example, at a first point in time, request 1 is transmitted to analyze an image from the front camera C10. At a second point in time, request 2 is transmitted to analyze an image from the rear camera C20. At a third point in time, request 3 is transmitted to analyze an image from the right camera C30. At a fourth point in time, request 4 is transmitted to analyze an image from the left camera C40. From a fifth point in time onward, the same process is repeated, starting with the front camera C10. The server device 2 performs analysis processing in response to each request and transmits each response. In response to each response, the device 1 performs voice reading if it is time for voice reading and there is a target. As mentioned above, it is also possible to set a ratio and prioritize a camera in a specific direction as the main camera.
[0231] State B is an example of control in a state where communication with the server device 2 is not possible, and is a case where analysis processing is performed on the device 1 side. Here, it is assumed that the device 1 has fewer computational resources than the server device 2, in other words, lower computational power, and the analysis processing takes time. Therefore, here, the analysis processing on the device 1 side is simpler than that of the server device 2. The simple analysis processing here is analysis processing that keeps the processing load on the device 1 low. In order to reduce the processing load, control B1 sets the processing ratio so that only some of the cameras in the four directions are used (on state). For example, only the front and rear cameras are used, the left and right cameras are turned off, and the front and rear are set to an equal ratio.
[0232] For example, at a first point in time, the device 1 analyzes the image captured by the front camera C10. At a second point in time, the device 1 pauses the analysis process or continues the analysis process of the image captured by the front camera C10 at the first point in time. At a third point in time, the device 1 analyzes the image captured by the rear camera C20. At a fourth point in time, the device 1 pauses the analysis process or continues the analysis process of the image captured by the rear camera C20 at the third point in time. This process is repeated from the fifth point in time onwards. At each point in time when the simple analysis process is performed, the device 1 performs voice reading if it is time for voice reading and there is a target. In the case of control B1, the cycle for image capture and voice reading is longer than, for example, control 1 of FIG. 9A, resulting in a lower processing load.
[0233] In this way, even when communication with the server device 2 is not possible, the user U1 can be assisted by simple analysis on the device 1 side. Similarly, control can be performed to limit analysis processing to cameras in a specific direction. Control B2 is an example of limiting analysis processing to only the front camera C10. At the first and third time points, the device 1 analyzes images from the front camera C10. Furthermore, if the analysis is set to be performed only at the first time point, the processing load can be further reduced.
[0234] As another example of control, control may be performed according to a multi-value state, not limited to the binary state of whether or not communication with the server device 2 is possible. For example, the present system may grasp the communication speed or the like as the performance of communication between the device 1 and the server device 2, and apply different control depending on the communication speed or the like. For example, normal control may be applied when the communication speed is equal to or greater than a first threshold, low-speed control may be applied when the communication speed is less than the first threshold, and even lower-speed control may be applied when the communication speed is less than a second threshold.
[0235] [Multiple Object Priority] 27 shows an example of control for performing voice reading according to the priority of objects when multiple objects are detected in the surroundings of user U1. Assume that user U1 is walking forward (+Y) on the left side of a road 2700. A car 2701 detected ahead of user U1 is stopped. Pedestrians 2702 and 2703 are also detected ahead, walking in the -Y direction toward user U1. A bicycle 2704 is also detected behind, traveling in the +Y direction toward user U1. Assume that device 1 detects these four objects at approximately the same time.
[0236] The device 1 determines the priority of these four objects and performs voice reading in order of priority. Examples of priority determination are as follows. First, for example, since the car 2701 is stopped, it is assigned a low priority based on its relative distance and speed from the user U1. Since the pedestrian 2702 is approaching the user U1, it is assigned a high priority based on the above-mentioned relative distance, speed, and degree of approach. Furthermore, for the moving pedestrians 2702, 2703, and bicycle 2604, their movement paths and the possibility of contact with the user U1 are determined. The pedestrian 2703 is approaching the user U1 in the Y direction, but its predicted path k3 is different from the path k1 of the user U1. In other words, the pedestrian 2703 is assigned a low priority because it is unlikely to come into contact with the user U1. The pedestrian 2702 is approaching the user U1 in the Y direction, and the predicted path k2 is close to the path k1 of the user U1's movement (for example, the distance in the X direction is small). In other words, the pedestrian 2702 is likely to come into contact with the user U1, and therefore is given a high priority.
[0237] The bicycle 2704 is approaching the user U1 in the Y direction, and is given a high priority due to its relative distance and speed to the user U1. In addition, the bicycle 2704's predicted route k4 is close to the route k1 of the user U1's movement. In other words, the bicycle 2704 is given a high priority because it may come into contact with the user U1.
[0238] Furthermore, the device 1 or the server device 2 compares the pedestrian 2702, which is assigned a higher priority, with the bicycle 2704, and assigns priorities by, for example, comparing their relative speeds. For example, if the bicycle 2704 has a higher relative speed than the pedestrian 2702, the bicycle 2704 is assigned the first priority and the pedestrian 2702 is assigned the second priority. Based on these priorities, the device 1 reads out the information in the following order at the timing of the read-out: bicycle 2704, pedestrian 2702, and car 2701. For example, the voice output may be, first, "A bicycle is coming from behind," second, "A pedestrian is coming from the front," and third, "A car is stopped in front."
[0239] In this way, the voice reading is performed in order of priority taking into consideration the proximity, route, etc., thereby increasing the safety of the user U1 and those around him.
[0240] [Guide to the control of risk factors according to their level of risk and safe spaces] 28 shows an example of control according to the risk level of a risk factor and an example of control for guiding a safe space. The device 1 and the server device 2 may set a predetermined object as a risk factor, or may determine the object as a risk factor by taking into account the relative distance and speed of the object. The risk factor here refers to an object that may come into contact with the user U1 and pose a risk to the user U1.
[0241] As an example of setting a predetermined object as a risk factor in advance, for example, the roadway 903 and the car 1003 running on the roadway 903 in Fig. 10 may be set as risk factors. In another example, a tool capable of killing or injuring may be set as a risk factor.
[0242] In the example of FIG. 25, it is assumed that user U1 is walking forward (in the +Y direction) on the left side of road 2800. Device 1 or server device 2 estimates and calculates a route k10 in the moving direction and traveling direction of user U1 based on an image from camera 403 or a sensor. It is assumed that a bicycle 2801 and a pedestrian 2802 are detected as objects ahead of user U1. The bicycle 2801 is traveling in the -Y direction toward user U1. The pedestrian 2802 is walking in the +Y direction, in the same direction as user U1. Device 1 or server device 2 estimates and calculates a route k11 of bicycle 2801. It is determined that route k11 of bicycle 2801 is close to route k10 of user U1, and there is a possibility of contact if things continue as they are.
[0243] The device 1 or the server device 2 determines and sets the bicycle 2801 as a risk factor based on the determination. The device 1 or the server device 2 may also determine and set a risk level for the risk factor. For example, the risk level may be calculated based on the relative distance or speed between the user U1 and the bicycle 2801, or based on the proximity between the route k10 of the user U1 and the route k11 of the bicycle 2801. Even if the route of the user U1 and the route of the object are in different directions, the object may be set as a risk factor if the routes are expected to intersect.
[0244] The risk level of a risk factor may be a single value, but may also be determined using multiple levels / multiple values, such as high / medium / low or first / second / third. In this example, the risk level of the bicycle 2801 is set to medium based on the relative distance D1 or relative speed V1 between the user U1 and the bicycle 2801. The device 1 controls the bicycle 2801, which is a risk factor, to be read out with a high priority based on the risk level.
[0245] For example, at time point 1, bicycle 2801 is detected and a voice readout equivalent to a basic notification such as "A bicycle is approaching from the front" is performed. Next, at time point 2, the system determines that bicycle 2801 is a risk factor and sets the risk level to medium based on the fact that the relative speed V1 between user U1 and bicycle 2801 is above a threshold, that the route is close, etc. The system determines that bicycle 2801, which is a risk factor, is a high priority target for voice readout in accordance with the risk level being medium. Furthermore, the system adds an alert to the voice readout of bicycle 2801 in accordance with the risk level being medium. The alert may be a voice output such as "Be careful," for example, or a predetermined alert sound. The alert sound may be a sound that repeats at a predetermined interval (e.g., beep beep...) or may generate a vibration.
[0246] Furthermore, when risk factors are read out aloud, the manner of the read out aloud / alert may be changed depending on the level of risk. For example, the higher the level of risk, the louder the volume may be. The tone of the voice / alert sound may be changed depending on the level of risk. The higher the level of risk, the shorter the period of the alert sound or vibration may be.
[0247] Furthermore, this system has a function to guide user U1 to a safe space by voice guidance to prevent or avoid danger factors. This will be explained using the same Figure 28. The device 1 and server device 2 predict the route k11 of the bicycle 2801 relative to the route k10 of user U1, and perform normal control until the bicycle 2801 approaches within a certain distance D11 from user U1 on the route k10. For example, at time point 1, the voice reads, "A bicycle is coming from the front." In normal control, the system basically maintains the user U1's position and route k10, and waits in the hope that the bicycle 2801 will change the route k11 taking user U1 into consideration.
[0248] If the bicycle 2801 enters within a predetermined distance D11 on the route k10 without changing the route k11, the system determines that there is a possibility of contact and performs control related to the guide of the safe space. Based on an analysis of the image from the camera 403 (for example, the front camera), the system calculates the safe space and the direction of movement to the safe space (in other words, the direction of escape) to avoid contact with the bicycle 2801 coming toward the user U1 on the route k10 of the user U1.
[0249] In the illustrated example, the system first determines whether it is possible to move to the left (-X) or right (+X) from the current position on the road 2800. In this example, there is no space to the left due to a wall or the like, and there is an empty space 2810 (in other words, a space with no objects) to the right. The system determines that movement to the left is not possible, but that movement to the right is possible. In addition, at this time, the system also determines the distance to surrounding objects, for example, for the empty space 2810 to the right. In particular, the system compares the path k11 of the bicycle 2801 with the paths of other objects to determine whether those paths intersect with the empty space 2810. As a result of this determination, the system determines that the empty space 2810 is safe, or in other words, sets the empty space 2810 as a safe space.
[0250] The device 1 guides the user U1 to move to the right, using the empty space 2810 to the right (+X) of the user U1 as a safe space. For example, at time point 3, the voice output corresponding to the guidance is "Please move to the right" or "Please move to the right."
[0251] As described above, the safety space is selected as a space that avoids the object, taking into consideration the path of the object approaching the user U1, thereby preventing or avoiding contact between the user U1 and the object, such as the bicycle 2801.
[0252] [Control to prioritize specific objects when reading aloud] Figure 29 shows an example of control for giving priority to specific objects when reading aloud. In this system, the objects to be read aloud may be set in advance to be limited to specific objects. This setting may be a system or application setting, or may be a setting for each user. For example, specific objects related to assistance for the visually impaired, traffic safety, accident prevention, etc. may be set as specific objects.
[0253] In the illustrated example, a user U1 is walking forward (+Y) on a sidewalk 2900 with a roadway on the right side. Objects detected ahead of the user U1 include a traffic light 2901, a crosswalk 2902, a pedestrian 2903, a crosswalk 2904, and a road sign 2905. In this system, it is assumed that traffic lights, crosswalks, and road signs are set in advance as specific objects. This system detects the specific objects, such as the traffic light 2901 and the crosswalk 2902, from the camera image with priority, and gives voice reading priority to the detected traffic light 2901 and the crosswalk 2902.
[0254] For example, when the device 1 first detects the crosswalk 2902, it reads out "There is a crosswalk on your right." Next, when the device 1 first detects the traffic light 2901 (for example, when the traffic light is green), it reads out "The traffic light is green." The system also detects changes in specific objects and reads out the changes. For example, a change from a green light to a red light is detected as a change in the status of the traffic light 2901. When the traffic light changes, the device 1 reads out the change in the traffic light. For example, it outputs "The traffic light has turned red."
[0255] The system also places importance on traffic safety and calculates the safety and risk when user U1 crosses the crosswalk 2902. For example, when the traffic light 2901 is red, the system may determine that the risk is high and output an alert according to the risk level. For example, the system may output a message saying, "The traffic light has turned red. Please be careful."
[0256] Similarly, other specific objects, such as road signs 2904, can be read out with high priority. In the illustrated example, when a road sign 2904 indicating a road closure for pedestrians is detected, a message such as "There is a road closure sign for pedestrians" is read out with priority. Other examples of specific objects include intersections and tactile paving blocks.
[0257] As described above, by giving priority to specific objects when reading them out loud, efficient support can be achieved according to the usage scenario, such as support for the visually impaired.
[0258] [Control example using a 360-degree camera] FIG. 30 shows a control example when a 360-degree camera (in other words, a celestial camera) is used as the camera 403. The illustrated image 3001 is a schematic diagram of an image captured by the 360-degree camera. In image 3001, the center point of the circular / ring-shaped area corresponds to the position of user U1, i.e., the position of the 360-degree camera. The circumferential direction of the circular / ring-shaped area corresponds to the front-back, left-right, and right-left directions (in other words, azimuth angles). The radial direction of the circular / ring-shaped area corresponds to the up-down direction (in other words, elevation and depression angles). Objects can be detected within a 360-degree celestial spherical image corresponding to such a circular / ring-shaped area. For example, if there is a pedestrian diagonally to the left in front of user U1, the pedestrian will appear as pedestrian image g1 in image 3001.
[0259] For example, user U1 in FIG. 6B may be provided with a 360-degree camera with a vertically upward optical axis on a hat 1D or the like. Alternatively, hat 1D may be provided with a 360-degree camera with a forward-facing optical axis on the front side and a 360-degree camera with a backward-facing optical axis on the rear side. Such an image 3001 may be obtained using a single 360-degree camera, or may be obtained by synthesizing such images 3001 using multiple cameras. Alternatively, a panoramic image 3002 such as that shown in the lower part of FIG. 30 may be used.
[0260] [Example of control that changes the audio output mode depending on the camera direction] FIG. 31 shows a control example for changing the manner (e.g., tone) of audio output for voice reading using front and rear cameras in the front and rear directions relative to user U1. For example, suppose there is an object 3101 detected in the space in front of user U1 based on an image captured by front camera C1 of device 1 (smartphone 1A), and an object 3102 detected in the space behind user U1 based on an image captured by rear camera C2. Device 1 changes the manner of audio output so that the voice reading for object 3101 based on the image captured by front camera C1 and the voice reading for object 3102 based on the image captured by rear camera C2 can be easily distinguished. For example, the voice tone or volume may be changed. In the illustrated example, the voice reading for object 3101 in the front is controlled to be output in a male voice (first tone, relatively low voice), and the voice reading for object 3102 in the rear is controlled to be output in a female voice (second tone, relatively high voice).
[0261] This makes it easier for the user U1 to recognize the front and rear positions. Also, with this control, even if multiple voice readings for multiple objects are performed almost simultaneously, it becomes easier to distinguish between them based on the difference in tone.
[0262] [Example of control using 3D audio output] FIG. 32 shows a control example using 3D audio output. In the illustrated example, user U1 is wearing smart glasses 1C and is near an intersection. Assume that there are two objects 3201 detected from the front direction (+Y) and 3202 detected from the right direction (+X) of user U1. In this control example, device 1 is equipped with an audio output device with a 3D audio output function. This 3D audio output function is a function that sets the source of audio in a 3D space and outputs audio from a speaker or the like so that the audio sounds as if it is coming from that source.
[0263] In the illustrated example, the device 1 determines the position L1 of the object 3201 detected in front of the user U1 and the position L2 of the object 3202 detected to the right of the user U1 based on the distance determination described above. These positions may be approximate. The device 1 then controls the 3D audio output so that the audio reading for the object 3201 sounds as if it is coming from the front position L1, and controls the 3D audio output so that the audio reading for the object 3202 sounds as if it is coming from the right position L2. This makes it easier for the user U1 to recognize the direction and position of the object. Furthermore, with this control, even if multiple audio readings of multiple objects are performed almost simultaneously, the user U1 can distinguish between them using the 3D audio output.
[0264] The following reference information is an example of technology related to 3D audio output functions. [Reference information] SRS-RA3000 Active Speaker / Neck Speaker Sony (sony.jp)
[0265] [Effects, etc.] As described above, according to the first embodiment, the image reading function can provide more suitable support to users who are visually impaired or non-visually impaired. For example, it can improve the safety of visually impaired people when walking and assist them in using transportation systems and facilities. Even for non-visually impaired people, it can improve safety measures and crime prevention on roads at night, and improve safety when driving while "walking while using a smartphone."
[0266] <Embodiment 2> The response output system (in other words, an AI conversation system) of embodiment 2 will be described with reference to Figure 33 and subsequent figures. The basic configuration of embodiment 2 is the same as or common to embodiment 1, and the following describes the components of embodiment 2 that are different from embodiment 1.
[0267] The following issues can be cited: Conventionally, there are AIs (conversational AIs / chat AIs, etc.) that generate conversations (in other words, chats) based on the content that a user inputs into a response input / output interface, but these do not take into consideration the surrounding situation when the user uses the conversational AI while walking around town, etc. Even when the user captures the surrounding situation with a camera and inputs it into the conversational AI, the capturing operation is necessary, which is cumbersome.
[0268] Therefore, in the second embodiment, when a user uses a conversational AI in a situation such as walking around town, multiple images capturing the surrounding environment are automatically captured at various times using multiple cameras on the device and input to the conversational AI. The conversational AI receives multiple images along with the conversational text input by the user, recognizes the surrounding environment from the multiple images, and generates and outputs a response conversational text that reflects the recognized surrounding environment in the topic of conversation. As a result, in the second embodiment, the surrounding environment is reflected during a conversation with the AI, without requiring complicated operations, enabling a convenient conversation.
[0269] In the first embodiment described above, a system that assists a user in walking, etc., based on input of images from multiple cameras of the device 1 was described. Since the second embodiment also uses input of images from multiple cameras, many of the components overlap with the first embodiment. The following mainly describes the components added in the second embodiment. The functions of the second embodiment share a common configuration with the first embodiment in that an AI analyzes time-series captured images from multiple cameras of the device 1 as input, recognizes the user's surroundings, and generates and outputs response text, etc., according to the surroundings. The functions of the second embodiment differ from the first embodiment in that a conversation takes place between the user and the AI, and the AI generates conversational sentences using text and images. The AI conversation function described in the second embodiment can be implemented in addition to the text-to-speech function described in the first embodiment, and can also be used in combination. That is, the function of the first embodiment can assist the user in walking by text-to-speech reading of the surroundings, while the function of the second embodiment can enable a conversation between the user and the AI while reflecting the surroundings.
[0270] In the second embodiment, the device 1 has a function for conducting a conversation between the user and the AI (sometimes referred to as an AI conversation function). This conversation may be implemented as a user interface in a specific application. In the second embodiment, the user's device 1 has an interface for conducting a conversation with the AI (sometimes referred to as a conversational AI) using text, audio, and images. The interface includes an interface for inputting and sending text, audio, and images of requests (e.g., instructions, questions, prompts, etc.) to the AI, and an interface for receiving and outputting text, audio, and images of responses (e.g., answers, etc.) from the AI. The conversational AI is implemented, for example, by the AI server 2 (particularly the MLLM), but may also be implemented as a function within the device 1 (particularly the local LLM). Note that the slash ( / ) in "text / audio / image" indicates OR. In other words, this expression indicates one or more of text, audio, and images.
[0271] In the AI conversation function of the second embodiment, the device 1 sends and inputs, as a request, data of multiple images (sometimes referred to as camera images) captured by multiple cameras of the device 1, along with text / audio / images of conversational text based on user input, to the conversational AI of the AI server 2. The conversational AI of the AI server 2 generates text / audio / images of a response to the conversational text from the input text / audio / images. At this time, the conversational AI of the AI server 2 recognizes the surrounding circumstances of the user and the device 1 by performing image recognition processing on the multiple input images (camera images). The conversational AI of the AI server 2 then controls the response conversational text to reflect the recognized surrounding circumstances. In other words, the user's device 1 controls the conversational AI to obtain a response conversational text that reflects the surrounding circumstances. For example, the conversational AI detects surrounding objects (in other words, target objects) and reflects the detected surrounding objects as topics in the conversation flow. For example, the conversational AI generates a conversational text that talks about the surrounding objects and inserts or adds it to the conversation flow. Examples of surrounding objects may be facilities or landmarks defined on a map, or objects not defined on a map (such as moving objects). This allows the user to have a conversation with the AI that reflects the surrounding situation.
[0272] Furthermore, in the AI conversation function of the second embodiment, the user's face may be included in the camera image. The conversation AI performs image recognition processing on the input camera image to recognize the surrounding situation (surrounding objects other than the user's face, etc.) as well as the user's facial expressions and behavior (for example, facial orientation, movement, gaze direction, etc.). There is a correspondence between facial expressions and emotions. Facial expressions / emotions can be classified using publicly known technology. For example, classifications include joy / anger / sadness / happiness / neutrality.
[0273] The conversational AI then controls the response text to reflect the recognized facial expressions and behaviors, along with the recognized surrounding circumstances. In other words, the user's device controls the conversational AI to receive response text that reflects the user's facial expressions and behaviors. For example, the conversational AI may generate a conversation that focuses on the user's facial expressions and behaviors and insert or add it to the conversation flow. Alternatively, the conversational AI may be controlled to change the content of subsequent conversations depending on the user's facial expressions and behaviors. For example, when providing guidance regarding the surrounding circumstances, such as navigating a map, the conversational AI may control the guide text to change or add depending on the user's facial expressions and behaviors. For example, the conversational AI may control the conversational AI to continue or develop a topic about a surrounding object if the user has a happy expression, or to discontinue the topic or change to another topic if the user has an anxious or sad expression. Furthermore, for example, when the conversational AI is navigating a route to a destination on a map, if the user looks anxious or sad, the conversational AI may be controlled to add a guide conversation to ease the user's anxiety or sadness.
[0274] [System Configuration] FIG. 33 shows a system configuration of the second embodiment. The system configuration of the second embodiment is roughly similar to the system configuration of FIG. 1 and so on. The system of the second embodiment (FIG. 33) corresponds to an example in which the system is applied to an AI chatbot service (in other words, an AI conversation service) on a communication network 9. A business provides the service to each user's device 1 via an AI server 2. FIG. 33 differs from FIG. 1 in that the device 1 is an AI conversation device compatible with the AI conversation function. Furthermore, the AI server 2 has, as the AI 20, a conversation AI 40 compatible with the AI conversation function. The conversation AI 40 generates a response conversation sentence in response to a conversation sentence input from the device 1 and outputs it to the device 1. This system has an AI conversation function, which is a function for having a conversation with the conversation AI 40 using images captured by multiple cameras equipped on the device 1 of the user U1. The configurations of FIG. 2 and FIG. 3 are also applicable to the second embodiment.
[0275] As in the first embodiment, various devices can be applied as the device 1. The various devices 1 described in FIGS. 5A, 5B, 6A, and 6B can also be applied to the second embodiment. For example, a smartphone 1A (FIG. 6A) can be used, and the multiple cameras can be a front camera C2 (in-camera) and a rear camera C1 (out-camera). User U1 converses with conversation AI 40 through images captured by the front camera C2 (in-camera).
[0276] Conversational sentences created by user U1 inputting text and voice into smartphone 1A are input to conversation AI 40, and multiple images captured by multiple cameras of smartphone 1A are also input to conversation AI 40. The multiple images include an image of the surrounding situation (e.g., the actual view ahead) captured by rear camera C1 (outer camera) and an image of user U1's face captured by front camera C2 (inner camera). During the conversation with the AI, each camera image is also continuously input on the time axis. Note that use of rear camera C1 (outer camera) is not limited to use, and cameras with other shooting directions (e.g., left and right cameras) may also be used.
[0277] The conversation AI 40 generates a response conversation (sometimes referred to as an output conversation) based on the text / audio / images, etc. (sometimes referred to as an input conversation) input from the device 1 of the user U1, and outputs it to the device 1. In doing so, the conversation AI 40 uses input image information from multiple cameras. The conversation AI 40 recognizes the surrounding situation, facial expressions, behavior, etc. from the camera images, and controls the response conversation to reflect this. The device 1 of the user U1 outputs the conversation from the conversation AI 40 in a predetermined format, such as text / audio.
[0278] For example, the conversation AI 40 generates conversational sentences that talk about the recognized surrounding object and inserts them into the conversation flow. Also, for example, the conversation AI 40 generates conversational sentences that talk about facial expressions and inserts them into the conversation flow. The conversation AI 40 may devise or change the way it responds to the conversational sentences by referring to the facial expressions, behavior, etc. of the user U1.
[0279] During a conversation, if the conversation AI 40 recognizes, based on camera images, a surrounding situation or surrounding object, such as a bicycle coming from ahead of the user U1, it generates a conversational sentence such as "There's a bicycle coming from ahead," drawing attention to the surrounding object, and inserts it into the conversation flow. Furthermore, if the conversation AI 40 recognizes, for example, a famous landmark as a surrounding situation or surrounding object, it generates a conversational sentence that discusses the surrounding object, such as "The building you see in front of you on the right is the famous Roppongi Hills, right?" and inserts it into the conversation flow. In this way, the conversation AI 40 generates conversation topics based on the recognition of camera images.
[0280] Furthermore, the conversation AI 40 may use, in addition to images of surrounding objects captured by the outer camera and images of the face captured by the inner camera, information detected on the posture (vertical / horizontal and tilted states), position, and direction of the device 1, and the line of sight (line of sight, gaze point, etc.) of the user U1. For example, the conversation AI 40 may generate conversational sentences to prioritize the topic of surrounding objects in the direction the device 1 is facing or in the line of sight of the user U1.
[0281] In the second embodiment, not only the multiple camera image input described in the first embodiment, but also voice input by the user, text of the voice recognition result, text of the character input result from key buttons, etc. are used as input information to the conversation AI. The conversation AI can generate text / voice of a response conversation based on the text / voice of the input conversation.
[0282] Note that technology for conducting AI chat in response to input of text information is already known. For example, Japanese Patent No. 6433614 discloses a system that uses log information and chatbot evaluation information in a normal chatbot system to realize more efficient chat.
[0283] A known AI technology for chatbots can be applied to the technical portion (particularly the conversation AI 40) related to the AI chat of the second embodiment. Alternatively, a dedicated AI for this embodiment may be implemented.
[0284] [Device and Server Configuration] 34A shows an example of the configuration of device 1 in embodiment 2, which differs from Fig. 7A in that it has AI chat software 347 (in other words, an AI conversation application) in non-volatile memory 303. Also, AI data 343 includes not only data related to image reading software 342 but also data related to AI chat software 347.
[0285] Figure 34B shows an example of the configuration of server 2 in embodiment 2, and differs from Figure 7B in that it has AI chat software 447 (in other words, an AI conversation server) in non-volatile memory 403. In addition, AI data 443 includes not only data related to AI software 442 but also data related to AI chat software 447.
[0286] [Processing flow: AI conversation method] The processing flow in the second embodiment is different from that in Fig. 4. The configurations of settings, image acquisition, image analysis, image readout data creation, image readout, etc. shown in Figs. 8 to 32 can also be applied to the second embodiment in the same way.
[0287] Fig. 35 shows a processing flow corresponding to the method (AI chat method, AI conversation method, etc.) in embodiment 2. The flow in Fig. 35 is executed in the system shown in Fig. 33, etc. That is, the processor 301 of the device 1 and the processor 401 of the server 2 execute processing according to the flow in Fig. 35.
[0288] In step S201, the device 1 of the user U1 turns on (enables) the AI conversation function when an instruction from the user U1 or the state of the device 1 matches a specific condition.
[0289] The actual state of AI20 (conversation AI40) and AI function 20b is that, in the configuration shown in Figures 34A and 34B, processor 301 and processor 401 load AI chat software 347 on device 1 side and AI chat software 447 on server 2 side, either alone or together with image reading software 342 and image reading software 442, and the processing is executed (in other words, as an execution module, instance).
[0290] The AI 20 (particularly the conversational AI 40) is composed of a generative AI, a large-scale language model (LLM), etc. The LLM may be a multimodal LLM (MLLM). Note that in this specification, the actual AI may also be referred to as an AI process circuitry.
[0291] In step S202, the device 1 receives a voice input or a text information input from the user U1. In the case of a voice input, the device 1 may recognize the voice input from a microphone or the like and convert it into text (characters). The device 1 sends the voice or the text of the voice recognition result, or the text of the text information input result to the AI 20 or the AI function 20b in response to an instruction (operation input) from the user U1 or the lapse of a predetermined time after the input.
[0292] In step S203, the device 1 captures and acquires multiple images by capturing images at predetermined timings from input videos of multiple cameras with different shooting ranges / shooting directions, at the same timing as in step S202 or at similar timings with a predetermined relationship. The AI function 20b imports the captured camera images (multiple images, image information). The image import at this time is performed by controlling the shooting timings of the multiple cameras according to settings, as shown in the previous embodiment 1 (FIG. 8, etc.).
[0293] For example, in a smartphone 1A (FIGS. 6A and 9A), if there is a camera C2 (in-camera) on the front side of the housing that takes pictures of the rearward direction and a camera C1 (out-camera) on the back side of the housing that takes pictures of the forward direction, while a conversation with AI is taking place through an image input including a facial image from the in-camera, an image input including a forward image from the out-camera is also taken at a predetermined timing. As in the first embodiment, the processing ratio between the image input from camera C1 and the image input from camera C2 can be set (for example, state A in FIG. 9A). The same applies when using front, rear, left, and right cameras.
[0294] The AI 20 or the AI function 20b determines and outputs the content of the response conversation based on the input text or voice of the user U1 and the camera image (multiple images). In doing so, the AI 20 or the AI function 20b uses an image of the surrounding situation captured by the outer camera or an image of the surrounding situation other than the face captured by the inner camera, and controls the response conversation so that the recognized surrounding situation is reflected. The AI 20 or the AI function 20b may also analyze the facial expression and behavior of the user U1 from the facial image captured by the inner camera, and control the response conversation so that the analyzed expression and the like are reflected.
[0295] Furthermore, the AI 20 or the AI function 20b may detect the posture (length, width, tilt, etc.), position, direction, facial orientation, movement, line of sight, etc. of the device 1. Based on the detection information, the system may, for example, estimate a surrounding object close to the line of sight of the user U1 and generate a conversation sentence so as to incorporate the surrounding object into the topic of conversation.
[0296] Furthermore, the AI 20 or the AI function 20b may understand, based on an analysis of the facial image, behaviors of the user U1, such as whether or not the user U1 is looking at the display screen on the front side of the device 1, in which direction the user is looking around, whether the user's gaze is changing frequently, or whether the user's gaze is steady and calm. When the user U1 holds the device 1 in his / her hand, the movement of the hand, the position and posture of the device 1 held in the hand, etc., change from time to time. The state of such changes can also be understood based on the sensor (posture detection unit 312, etc.) of the device 1 and the contents of the camera image.
[0297] The AI function 20b (AI conversation application, in other words, client software) on the device 1 side may perform pre-processing on the device 1 side, that is, sub-processing before the AI 20 on the server 2 side performs main processing. The AI function 20b may format the text of the conversation, or may perform format conversion processing such as converting from voice to text or from image to text. The AI function 20b may also extract points or ranges of interest based on the gaze of the user U1 from the camera image, or may perform image processing such as editing the camera image to reflect points of interest.
[0298] In step S204, the device 1 transmits the input conversational text and camera images (multiple images) as well as preprocessing information to the AI 20 on the server 2 side as a communication request (data of a predetermined communication protocol). The device 1 may transmit the captured information (input conversational text and camera images) to the AI 20 on the server side without the AI function 20b performing any action. As in the first embodiment (FIG. 3), the device 1 may complete the series of processes alone. In this case, the device 1 may perform the processing using the AI function 20b without transmitting information to the server 2 side.
[0299] In step S205, the server 2 of the present system analyzes the input conversation text etc. obtained from the device 1, the camera image, and, if any, the preprocessing information, using the AI 20 (conversation AI 40). In particular, in step S205, the AI 20 analyzes / recognizes the surrounding situation from the camera image (e.g., an image from the outer camera), and as a result, obtains text describing the surrounding situation.
[0300] Also, step S206 may be added. In step S206, the AI 20 analyzes / recognizes the facial expression and behavior of the user from a camera image (for example, an image from the front camera), and as a result, obtains information representing the facial expression and behavior.
[0301] In step S207, AI20 generates one or more candidate response sentences (text / audio / image) from the results of steps S205 and S206, i.e., the analysis results of the input conversation sentence and the camera image (surrounding circumstances, facial expressions, etc.). Note that in step S207, one response sentence may be determined from the beginning without generating candidates. The response sentence (output conversation sentence) is typically text (in other words, a sentence consisting of character strings) generated using text representing the surrounding circumstances and facial expressions, which are the recognition results. As the response conversation sentence, a conversation sentence using audio or images (still images or video) may be generated and output instead of or in addition to text.
[0302] In step S208, the system (e.g., AI20 of server 2) determines the conversation sentence to be output (output conversation sentence) by selecting one from the candidates based on a predetermined judgment (e.g., judgment of the degree of change described in embodiment 1) and timing based on the response sentence candidates resulting from step S207, and transmits the information to device 1.
[0303] In step S209, the device 1 receives information about the output conversation sentence from the AI 20 determined in step S208, and automatically outputs the output conversation sentence to the user U1 in a predetermined output format. The output may be a voice readout (in other words, audio output), a text display on a screen, or an image display. The output format is configurable.
[0304] In step S210, device 1 checks whether to turn the AI conversation function OFF. If it is to be turned OFF (YES), the flow ends. If it is to remain ON (NO), the flow returns to step S201 and repeats the same process.
[0305] As described above, in the second embodiment, in a system having a conversation AI that responds to conversational sentences input by voice or text from user U1, by inputting information from one or more cameras to the conversation AI at an appropriate timing, it is possible to optimize the responses of the conversation AI, thereby realizing a suitable conversation that reflects the surrounding circumstances, facial expressions, etc. of user U1.
[0306] [System configuration example: Software] FIG. 36A and other figures show an example of a system configuration, particularly a software configuration, according to the second embodiment.
[0307] The system of Figure 36A (in other words, an AI conversation system, etc.) is a system in which a user terminal, which is device 1 of user U1, and server 2, which is an AI server, are appropriately connected via communication, similar to Figure 33. Device 1 is provided with app 351, which is a client application for realizing the AI conversation function. Server 2 is provided with AI 352, which is server software for realizing the AI conversation function. App 351 also has a conversation client 353 and an image recognition client 354 as components of a program, etc. AI 352 has a conversation AI 355 and an image recognition AI 356 as components of a program, etc.
[0308] In the system of Fig. 36A, a conversation client 353 creates an input conversation sentence (instructions, prompts) based on the text / voice input by the user and transmits it as a request a1 (for example, a packet on a communication network) to the server 2. In addition, an image recognition client 354 acquires camera images (plurality of images) and transmits data of the camera images (plurality of images) a2 to the server 2 at approximately the same timing as the input conversation sentence (request a1). The data of the request a1 and the camera images a2 may be transmitted together at approximately the same timing, or may be transmitted separately at different timings.
[0309] The conversation AI 355 of the server 2 generates a conversation (output conversation) for response from the input conversation of the request a1 based on a conversation generation model. In doing so, the server 2 also uses the image recognition AI 356. The image recognition AI 356 receives a camera image a2 as input, recognizes the surrounding situation, facial expressions, and behavior based on image recognition processing, and obtains information a3, such as text describing the surrounding situation, or text describing facial expressions and behavior. The conversation AI 355 receives the information a3 along with the input conversation, and generates an output conversation that reflects the surrounding situation, facial expressions, and the like as a topic based on this information. The conversation AI 355 transmits the text / audio of the output conversation to the device 1 as a response a4. The conversation client 353 of the device 1 outputs the output conversation to the user U1 in the form of text / audio / images.
[0310] Regarding the series of steps related to the AI conversation function, a typical example is that user U1 can input a conversation sentence by voice and the AI can output a conversation sentence in response. Another example is that user U1 can input the text of the conversation sentence on the screen and the text of the conversation sentence in response from the AI can be displayed on the screen.
[0311] 36A, as a modification, image recognition AI 356 may be integrated into conversation AI 355. Also, conversation client 353 and image recognition client 354 may be integrated into one.
[0312] The system in FIG. 36B is a variation of the system in FIG. 36A, and its main difference is that it includes an image recognition AI 357 on the device 1 side. The image recognition AI 357 on the device 1 side is realized by a local LLM. The image recognition AI 357 on the device 1 side receives an image from the camera of the device 1 as input, recognizes the surrounding situation or facial expressions and behavior based on image recognition processing, and obtains information such as text describing the surrounding situation or text describing facial expressions and behavior. The image recognition AI 357 on the device 1 transmits the information a3 to the server 2 at approximately the same timing as the input conversational text (request a1).
[0313] The conversation AI 355 of the server 2 receives the input conversation sentence of the request a1 and information a3 such as the surrounding situation and facial expressions, and generates a conversation sentence (output conversation sentence) that reflects the surrounding situation or facial expressions as a topic based on a conversation generation model. The conversation AI 355 transmits the text / audio of the output conversation sentence to the device 1 as a response a4. The conversation client 353 of the device 1 outputs the output conversation sentence to the user U1 as text / audio / images.
[0314] For the image recognition AIs 356 and 357, in particular, known object detection technologies such as CNN (Convolutional Neural Network) can be applied.
[0315] Although there may be a time difference between the processing of the text / audio of the conversation and the processing of the camera image due to computer processing time, these processes may be synchronized as much as possible by setting and controlling the timing of transmission of various data and information. For example, in Figure 36A, after image recognition processing by image recognition AI 356 is completed, the input conversation sentence of request a1 and information a3 may be input to conversation AI 355 to generate a response, and after image recognition processing by image recognition AI 357 is completed, the conversation sentence request a1 and information a3 may be transmitted.
[0316] As another variation, an image recognition AI for recognizing the surrounding situation (surrounding objects, etc.) and an image recognition AI for recognizing facial expressions and behavior may be provided separately. As another variation, an image recognition AI may be provided for each camera in correspondence with multiple cameras of device 1. For example, when smartphone 1A uses two cameras, an in-camera and an out-camera, two separate image recognition AIs may be provided: one for the in-camera image and one for the out-camera image.
[0317] The system of FIG. 36C is a variation of FIG. 36A, and its main difference is that it includes a conversation AI 358 on the device 1 side and an image recognition AI 356 on the server 2 side. The conversation AI 358 on the device 1 side is realized by a local LLM. The image recognition client 354 of the device 1 sends a camera image a2 to the server 2. The image recognition AI 356 of the server 2 obtains information a3 representing the surrounding circumstances, facial expressions, etc. from the camera image a2 based on image recognition processing. The image recognition AI 356 of the server 2 sends the information a3 to the device 1 as a response. The image recognition client 354 transfers the information a3 to the conversation AI 358. The conversation AI 358 receives input conversational text entered by the user and the information a3, and generates conversational text (output conversational text) that reflects the surrounding circumstances, facial expressions, etc. as topics based on a conversation generation model. The conversation AI 358 outputs the output conversational text to the user U1 as text / audio / images.
[0318] The system of Fig. 36D is a variation of Fig. 36A, and has a configuration integrated into device 1 similar to Fig. 3, with the main difference being that device 1 is provided with a conversation AI 358 and an image recognition AI 357. The conversation AI and image recognition AI of embodiment 2 are not required on the server 2 side. When used in conjunction with the text-to-speech function of embodiment 1, the server 2 side only needs to implement the text-to-speech function.
[0319] 36A etc., the server 2 may be equipped with a function related to the voice reading in the first embodiment. For example, the AI 20 in the first embodiment (the AI that recognizes an object from an input image, converts it into text, and outputs it) and each AI in the second embodiment (the conversation AI and the image recognition AI) may be integrated and implemented.
[0320] [AI conversation function: configuration example] Figure 37 shows a more detailed example configuration of the AI conversation function based on Figure 36A. In the system of Figure 37, server 2 has conversation AI 355, which includes image recognition AI 356. Device 1 has, as functional units, setting unit 371, image acquisition unit 372, conversation input unit 373, conversation output unit 374, multiple cameras 375, input device 376, output device 377, etc. Each functional unit is realized by processing by a processor or by implementing a circuit.
[0321] The setting unit 371 is a part for performing system settings and user settings related to the AI conversation function, etc., and also provides a user interface for this purpose. The multiple cameras 375 are cameras with different imaging ranges / imaging directions, such as the aforementioned in-camera and out-camera. From the multiple cameras 375, multiple images b1 (corresponding to camera input images 801 in FIG. 8) including images from each camera are obtained.
[0322] The image acquisition unit 372 is a part that acquires capture images (corresponding to 801 in Figure 8) at predetermined timings based on multiple images b1 from multiple cameras 375, and transfers them to the conversation input unit 373 as multiple images b2 (image information).
[0323] The conversation input unit 373 acquires text / audio / image b3 (user input information) input by the user through the input device 376 and creates an input conversation sentence (instructions, prompts, etc.) to be input to the conversation AI 355 based on the user input information. The conversation input unit 373 transmits / inputs the input conversation sentence and multiple images (image information) to the conversation AI 355. When the user input format is voice, the conversation input unit 373 acquires voice data through a microphone, performs voice recognition processing on the voice data to convert it into text, and creates an input conversation sentence in text format. At this time, the conversation input unit 373 also creates a request b4 to send to the server 2 so that multiple images b2 are attached to the input conversation sentence. The request b4 in FIG. 37 corresponds to the request a1 and camera image a2 in FIG. 36A combined into one. The format for attaching multiple images b2 to the input conversation sentence (the format of the request b4) is not limited, but an example is shown in FIG. 38.
[0324] The conversation input unit 373 transmits a request b4 to the server 2. The server 2 receives the request b4, and the conversation AI 355 extracts the text of the input conversation and multiple images from the request b4. The conversation AI 355 analyzes the input conversation and multiple images to generate text / audio / images that will serve as a response sentence (output conversation sentence) for the conversation. The conversation AI 355 recognizes the surrounding situation, facial expressions, etc. from the multiple images using the image recognition AI 356, and obtains text representing the surrounding situation, facial expressions, etc. The conversation AI 355 receives the text of the input conversation sentence and the text representing the surrounding situation, facial expressions, etc. as input, and generates text of a response conversation sentence (output conversation sentence) that reflects the surrounding situation, facial expressions, etc. in the topic. Note that in this case, the conversation AI 355 may generate a first output conversation sentence for the input conversation sentence and a second conversation sentence for the text representing the surrounding situation, facial expressions, etc., and output them as a set. The server 2 (conversation AI 355) transmits the output conversation sentence to the device 1 as a response b5.
[0325] Device 1 receives response b5, and the conversation output unit 374 extracts the text of the output conversation sentence from response b5 and outputs the output conversation sentence in a predetermined format from the output device 377. If the output format is audio, the conversation output unit 374 creates audio data from the text of the output conversation sentence by speech synthesis processing, and outputs the audio from the speaker (output device 377) based on the audio data (b6).
[0326] When recognizing the surrounding situation from camera images, this system may use detection information such as the device 1's position, orientation, and posture, as well as information from a map database, etc., in combination to improve the accuracy of recognizing the surrounding situation. When the AI on the server 2 performs image recognition processing, the device 1 may transmit related information such as detection information along with the camera images.
[0327] [Request Data] FIG. 38 shows an outline of an example configuration of request b4. Request b4 includes text data 3801 of an input conversation (prompt) and image data 3802 of multiple images (camera images). The image data 3802 of the multiple images includes image data (first camera image data to Nth camera image data) for each camera (corresponding shooting direction). For example, in the case of two front and rear cameras (C1, C2) as shown in FIG. 6A, the first camera image data is image data from the outer camera (C1) that captures the front, and the second camera image data is image data from the inner camera (C2) that captures the rear. Each camera image data also includes information such as the shooting direction and the shooting time. Such data of request b4 is transmitted and received as data such as packets according to a communication protocol.
[0328] Furthermore, the illustrated example of request b4 is an example in which both input conversational text and camera images are generated at roughly the same time. However, there are also cases in which only input conversational text or only camera images are generated at a certain time. In such cases, the data content of request b4 is either text data 3801 or image data 3802. The illustrated example of request b4 is a case in which multiple images captured by multiple cameras at roughly the same time are transmitted as a set, but this is not limited to this. As in the aforementioned FIG. 9A, when capturing images at different times with a predetermined cycle or ratio, a corresponding data set can be created. For example, in the case of control 1 in FIG. 9A, the data set of image data 3802 may be composed of one camera image data per time point, and these may be transmitted sequentially. As another example, camera image data from multiple nearby time points may be combined into a single request b4. For example, image data 3802 may contain four camera image data sets from time points 1 to 4 of control 1.
[0329] [Map and Surroundings] FIG. 39 is a schematic diagram showing an example of the surrounding conditions of user U1 and device 1 when using the AI conversation function of the second embodiment, overlooking the current positions of user U1 and device 1, roads, surrounding objects, etc. on a map. The map has north-south, east-west, and latitude and longitude (not shown), and corresponds to a spatial coordinate system (X, Y, Z). In this example, the Y direction is north. User U1 (similar to FIG. 6A ) holding smartphone 1A as device 1 is at position L1 (current position) on the map. Position L1 is expressed by position coordinates (X1, Y1, Z1) and corresponding latitude, longitude, altitude, etc. At position L1, user U1 and smartphone 1A are facing north (+Y direction). Direction DC1 is the shooting direction of the outer camera (camera C1) of smartphone 1A, and direction DC2 is the shooting direction of the inner camera (camera C2) of smartphone 1A. A route 3901 is an example of a travel route from a current position L1 of a user U1 to a destination 3902 (corresponding to a surrounding object SO3). In this example, the user U1 has a destination 3902 and is traveling toward the destination 3902, but in other examples, there may be no particular destination.
[0330] Examples of surrounding objects include surrounding objects SO1 to SO8. In this example, the surrounding objects are buildings, facilities, landmarks, etc. that exist on a map. The surrounding objects are not limited to these. The surrounding objects handled in the second embodiment are any objects that can be recognized and detected from camera images. Note that the surrounding objects may be pre-registered in a service / database such as a map or navigation system, or may not be pre-registered. In the case of an object such as a facility that is registered in a service such as a map, information about the facility can also be provided by the service, but the function of the second embodiment differs in that the surrounding objects such as the facility are reflected in the conversation with the AI. Furthermore, in the second embodiment, surrounding objects that are not registered in a map, etc. can also be reflected in the conversation with the AI.
[0331] Furthermore, gaze direction 3903 is an example of the gaze direction from user U1 at position L1, and is a direction diagonally forward and to the right of north (+Y direction) at approximately 30 degrees. In the case of gaze direction 3903, for example, surrounding objects SO1 and SO2 are visible in the vicinity of gaze direction 3903 within the field of view of user U1 (in other words, these surrounding objects are captured in the corresponding camera image), but surrounding object SO4, etc. are not visible. Note that with regard to the field of view and gaze direction, it does not matter whether user U1 actually recognizes surrounding objects SO1, etc.
[0332] [Multiple images (camera images)] FIG. 40 shows an example of multiple images (camera images) that device 1 captures using multiple cameras and transmits to server 2. This example shows two images captured by two front and rear cameras (C1, C2) as shown in FIG. 6A. Image 4001 in (A) is an image captured by the out-camera (camera C1) capturing an image in front of user U1, and image 4002 in (B) is an image captured by the in-camera (camera C2) capturing an image behind user U1 at approximately the same time. The example in FIG. 40 corresponds to the example situation in FIG. 39.
[0333] In the forward image 4001, for example, a surrounding object SO1 and a surrounding object SO2 are shown, and these surrounding objects can be detected based on image recognition. For example, a store, which is the surrounding object SO1, can be detected from the characters on a signboard. For example, a building, which is the surrounding object SO2, can be detected from the shape of a tower. Other examples of surrounding objects that may be recognized include trees. An x mark 4003 is an example of a gaze point 4003 corresponding to the line of sight.
[0334] In the rear image 4002, no particularly conspicuous surrounding objects are captured and detected. In the rear image 4002, the user's face (corresponding face area 4004) is captured, so the face, facial expression, behavior, etc. can be determined and detected based on image recognition. As behavior, the direction of the face, movement, etc. can be determined. Furthermore, the gaze direction may be estimated from the state of the eyes. Facial expressions are classified into a plurality of types (emotion types), such as joy / anger / sadness / happiness / neutral, based on known emotion recognition technology, for example.
[0335] In the function of the second embodiment, for example, peripheral objects in a forward image 4001, such as peripheral objects SO1 and SO2, are recognized and detected and reflected as topics in the AI conversation. In particular, in the case of control using gaze direction, among all peripheral objects in the image, peripheral objects closer to the gaze direction (point of gaze) are given priority. For example, in image 4001, the gaze direction (corresponding point of gaze 4003) is to the right, so even if a peripheral object is detected in an area to the left, the peripheral objects SO1 and SO2, which are to the right, are used preferentially. The distance between the point of gaze 4003 and the peripheral objects (representative position coordinates) may be determined. An area may be taken with a predetermined radius centered on the point of gaze 4003, and it may be determined whether any peripheral objects are included in that area.
[0336] [Example (1)] A specific example of using the AI conversation function of embodiment 2 will be described. Figure 41 shows the situation of user U1 on a map corresponding to this specific example. Figures 42 and 43 show a specific example of the flow of a conversation with the conversation AI corresponding to the situation in Figure 41.
[0337] In the specific example of FIG. 41, user U1 uses a map app on smartphone 1A to receive navigation (in other words, guidance, etc.) for a route to a destination. This map app is used in conjunction with or integrated with a conversation AI. For example, in FIG. 36A, conversation AI 355 can be considered as a conversation AI integrated with a map service, and conversation client 353 can be considered as a conversation client integrated with a web browser, etc., that receives map services, etc. The conversation AI navigates the route to the destination through a conversation with user U1. During this navigation conversation, the conversation AI reflects surrounding objects, etc., recognized from camera images as topics of conversation.
[0338] In FIG. 41, initially, user U1 is at position p0 as his current position (present location). User U1's destination 4101 is facility "XXX" on the map, and the surrounding object recognized from the image is SOX. Positions p1 and the like are examples of changing current positions. Route 4110 is the travel route from position p0 to destination 4101, and is an example of a route recommended by navigation.
[0339] In FIG. 42, first, user U1 launches an application (conversation client 353 in FIG. 36A) on the smartphone 1A and inputs, for example, "Hello. I want to go to 'XXX'" by voice (referred to as input conversation sentence 4201). The smartphone 1A transmits this input conversation sentence 4201 along with a camera image to the server 2. The server 2 (conversation AI 355 in FIG. 36A, particularly the map application) generates a conversation sentence in response to the input conversation sentence 4201, such as "Hello, user U1. A route from your current location to the destination 'XXX' has been set. The map will be displayed. It's a 15-minute walk to your destination," and responds to the smartphone 1A (output conversation sentence 4202). The route here corresponds to the route 4110 in FIG. 41. The smartphone 1A outputs the output conversation sentence 4202 by voice and displays a map on the screen.
[0340] User U1 follows the map navigation to head to destination 4101. User U1 proceeds along the presented route (path 4110). The conversational AI outputs a navigation such as "Please go straight ahead (north) for a while" (output response sentence 4203). User U1 proceeds along the navigation.
[0341] For example, user U1 arrives at position p1 just before traffic light A (intersection A). If user U1 is unsure where to turn left, he or she may say, for example, "I wonder if I should turn left at the next traffic light?" (input conversation sentence 4204). At this time, the facial expression is normal. In response to input conversation sentence 4204, the conversation AI outputs a navigation message such as, for example, "No. Just follow the road until there are two more traffic lights ahead." (output response sentence 4205). User U1 continues to go straight according to the navigation and arrives at position p2, for example.
[0342] Assume that user U1 continues driving straight ahead but is still unsure where to turn left. He / she is silent and has an anxious expression. A camera image capturing his / her expression at that time is sent to server 2 (silence, camera image 4206). Based on the camera image, the conversation AI recognizes and classifies user U1's expression as sadness, and generates and outputs a navigation response to address the sadness, such as "You're almost there. You can see the traffic light for turning left." (output response statement 4207). User U1 follows the navigation and turns left at traffic light B (intersection B) (silence, camera image 4208). At this time, user U1 is silent, for example, and has a neutral expression.
[0343] Furthermore, from the fluctuations in the camera image at this time, the system can also recognize the act of turning left. At this time, for example, a pedestrian 4105 at the destination of the left turn is recognized and detected from the camera image as a type of surrounding object (moving body). In this case, the presence of the pedestrian 4105 can be reflected in the conversation (navigation). The conversation AI may provide navigation such as, "Proceed toward the pedestrian at that traffic light."
[0344] Next, the conversation AI outputs a navigation such as "Please go straight to 'YYY'" (output conversation sentence 4209). The conversation AI can provide such navigation when it determines that there is a facility such as "YYY" on the route based on map data or recognition of surrounding objects. User U1 goes straight according to the navigation and arrives at position p3, for example. User U1 utters, for example, "I can see 'YYY'" (input conversation sentence 4210). The user's facial expression at this time is normal.
[0345] Based on the camera image, the conversation AI detects, for example, a bicycle approaching from behind as a surrounding object of user U1 (similar to embodiment 1). To alert user U1, the conversation AI generates and outputs a response such as, for example, "A bicycle is approaching from behind. Please be careful" (output conversation sentence 4211).
[0346] User U1 arrives at position p4 near "YYY" just before traffic light C (intersection C) where he should turn right. The conversation AI outputs a navigation message such as "Turn right at the intersection with "YYY"" (output conversation sentence 4212). User U1 turns right at intersection C according to the navigation message (silence, camera image 4213). At this time, the user is silent and has a normal expression.
[0347] Figure 43 is a continuation of Figure 42. User U1 reaches position p5, for example. The conversation AI recognizes stone monument 4103 as a surrounding object based on the camera image. The conversation AI outputs a response that reflects the recognized stone monument 4103 as a topic, such as "The stone monument you see on the left is engraved with a famous song by Mr. B" (output conversation sentence 4214). In response to output conversation sentence 4214, user U1 utters, for example, "That's true" (input conversation sentence 4215). The expression on the user's face at this time is one of joy.
[0348] Furthermore, depending on the facial expression of the user U1 at the time of the reaction (input conversation sentence 4215), the conversation AI may either develop or discontinue the topic regarding the stone monument 4103. For example, if the facial expression is one of joy, the conversation AI may output an output conversation sentence that develops the same topic.
[0349] The conversation AI continues navigating the route. User U1 reaches position p6. The conversation AI recognizes traffic light D (intersection D) where to turn right based on map data or recognition of surrounding objects. The conversation AI recognizes, for example, the church 4104 at traffic light D where to turn right. The conversation AI outputs navigation such as, "You will see a church with a cross diagonally ahead to your right. Turn right at the traffic light just before it." (output conversation sentence 4216). User U1 utters, for example, "OK" and turns right at traffic light D (input conversation sentence 4217). The user's facial expression at this time is normal.
[0350] User U1 continues straight along the route after turning right and arrives at position p7. The conversation AI recognizes facility "XXX", which is destination 4101, as a surrounding object SOX from the camera image. For example, suppose that a tower, which is part of facility "XXX", is detected diagonally forward and to the left from the camera image. The conversation AI reflects the surrounding object and outputs navigation such as "I can see the tower of destination "XXX" ahead and to the left" (output conversation sentence 4218). In response to output conversation sentence 4218, user U1 utters, for example, "Is that it?" (input conversation sentence 4219). At this time, user U1's line of sight is diagonally forward and to the left, which coincides with the direction in which facility "XXX" is located.
[0351] Based on the camera image and gaze direction, the conversation AI determines that user U1 is looking toward facility "XXX" and outputs navigation such as, "Yes. You will arrive at the entrance in 20 meters." (output conversation sentence 4220). User U1 follows the navigation and arrives at destination 4101 "XXX." User U1 utters, for example, "We've arrived." (input conversation sentence 4221). The expression at this time is one of joy. Upon arriving at destination 4101, the conversation AI outputs, for example, "We've arrived. Route guidance will end." (output conversation sentence 4222), and ends route navigation.
[0352] According to the above specific example, by reflecting the surrounding circumstances of the user U1 in the conversation of the AI, more suitable conversations, such as route navigation and guidance, can be realized.
[0353] [Example (2)] Another specific example will be described. Figure 44 shows an example of the flow of a conversation with an AI in another situation. In the situation of user U1, user U1 has just arrived at Tokyo Station and has an appointment to meet someone, but there is still time until the appointment, and the user does not have a specific purpose. User U1 inputs, for example, "Hello. I'm at Tokyo Station today" into the app on device 1 (input conversation sentence 4401). From the input conversation sentence 4401 and the camera image, the conversation AI recognizes and infers that user U1 is at Tokyo Station, particularly at the Yaesu Central Exit of Tokyo Station. The conversation AI responds, for example, "Hello. It appears you are at the Yaesu Central Exit of Tokyo Station. Is there anything I can help you with?" (output conversation sentence 4402), reflecting the recognition result (particularly "Yaesu Central Exit").
[0354] In response to output conversational text 4402, user U1 utters, for example, "I have plans to meet someone at "AAA," but I still have an hour to go." (input conversational text 4403). From input conversational text 4403 (information "AAA"), the conversational AI can determine that its assumption that user U1 is at Tokyo Station (particularly the Yaesu Central Exit) is correct. Furthermore, from input conversational text 4403, the conversational AI can determine that user U1 is interested in a destination on the map called "AAA" (for example, a specific facility) and plans for the next hour.
[0355] The conversational AI generates a topic related to what it has understood from the input conversational sentence 4403, and outputs a response such as, "To get to "AAA," it's quicker to go to the left and pass through "BBB." There are also souvenir shops on the left." (output conversational sentence 4404). In this example, the conversational AI suggests a route to the destination "AAA" and facilities along the route.
[0356] In response to output conversation sentence 4404, user U1 utters, for example, "Maybe I should buy souvenirs on the way home. Maybe I should have something light to eat." (input conversation sentence 4405). In this example, user U1 reacts negatively to the topic (souvenirs) presented by the conversation AI, and instead replies that he is interested in food. From this reaction (input conversation sentence 4405), the conversation AI can determine that user U1 is negative about souvenirs and positive about food.
[0357] Assume that user U1 is walking in a certain direction in Tokyo Station after input conversation sentence 4405. The conversational AI recognizes surrounding objects from camera images at that time to determine user U1's direction and location. The conversational AI generates topics based on the recognized direction and location. For example, the conversational AI outputs a response such as, "There's a newly opened restaurant called 'DDD' at 'CCC', where you're heading. You like sweets, don't you? Their cakes (¥1,000) are apparently popular." (output conversation sentence 4406). In this example, the conversational AI identifies a restaurant called 'DDD' located in the direction user U1 is heading as a candidate and suggests meals based on information about the restaurant (which can be map data or web search information) and user U1's attributes and preferences.
[0358] The conversational AI and app may also generate links such as URLs from words in the conversation. Users can obtain detailed information about a word by specifying the link (either by clicking on the displayed information or by voice input) in the conversation. For example, the link for "DDD" or the link for "cake" can provide detailed information about the store or food.
[0359] User U1 agrees with output conversation sentence 4406, for example, and utters, "That sounds good. I'll go and check it out" (input conversation sentence 4407).
[0360] As in the above example, conversation with the AI can be realized by providing topics based on the user U1's surroundings and changing the topic depending on the user U1's reaction.
[0361] [Conversation screen example] As a supplement, FIG. 45 shows an example of a GUI for a conversation with an AI on the front screen of the smartphone 1A. The example in FIG. 45 is an example in which the text of the conversation (including the text of the voice recognition result) is displayed. In FIG. 45, a touch panel screen 4500 has an in-camera (camera C2), and as an interface for "AI chat," there are a display field 4501 for the AI character's response conversation text and a display field 4502 for the user's input conversation text. Also, a map 4503 is displayed at the bottom when using a map application (navigation function). Note that the map 4503 may be displayed as a separate window.
[0362] [How to reflect facial expressions / emotions] Any of the following methods can be used to reflect facial expressions / emotions in AI conversations. (1) The AI may change the topic of conversation depending on the facial expression / emotional state. (2) The AI may change the way it phrases a conversation, without changing the topic of the conversation, depending on the facial expression / emotional state.
[0363] As described above, according to the second embodiment, even if the user U1 does not perform an image capturing operation, it is possible to realize a conversation with the AI that reflects the surrounding circumstances of the user U1 and the device 1. Furthermore, even if the user does not input a conversation sentence, the AI can automatically generate a conversation sentence in response to the input of a camera image.
[0364] [Modification of the second embodiment] The following modification of the second embodiment is also possible.
[0365] When both input conversational text and camera images are input at approximately the same time, the ratio of one response / input / output process of a conversation AI (e.g., 355 in FIG. 36A) to the response / input / output process of an image recognition AI (e.g., 356 in FIG. 36A) may not be limited to a 1:1 ratio, but may be a preset ratio. For example, if the processing load of recognizing the surrounding situation from an image by the image recognition AI 356 in FIG. 36A is higher than the processing load of conversation generation by the conversation AI 355, the ratio of the conversation AI processing to the image recognition AI processing may be set to, for example, 2:1 to reduce the overall load. Conversely, if the processing load of recognizing the surrounding situation from an image by the image recognition AI 356 in FIG. 36A is lower than the processing load of conversation generation by the conversation AI 355, the ratio of the conversation AI processing to the image recognition AI processing may be set to, for example, 1:2 to reduce the overall load.
[0366] There are several patterns for the surrounding situation recognized from the camera image. For example, there are cases where a fixed surrounding object is detected in a space such as a facility, and cases where an object requiring caution when walking (a moving object such as a car, bicycle, or other person) is detected as described in the first embodiment. A priority / order of precedence may be set between these surrounding objects. A relatively high priority is set for objects requiring caution. When the conversational AI detects an object requiring caution as a surrounding object during a conversation with a user, it inserts a conversational sentence indicating that caution should be exercised over normal conversational sentences (for example, a topic about the facility).
[0367] Additionally, surrounding objects on the map may be categorized into types / categories, and priorities may be set among the types / categories of surrounding objects according to the user's level of interest, and similarly, the output of dialogue may be controlled. Examples of types / categories include history, nature, food, shopping, etc. The user may be able to set their level of interest.
[0368] When the text-to-speech function of the first embodiment (i.e., the function of reading out surrounding objects and situations) and the AI conversation function of the second embodiment (i.e., the function of reflecting the surrounding situation in the conversation as a topic) are used together, a priority or processing ratio may be set between these functions. For example, a mode that prioritizes text-to-speech and a mode that prioritizes AI conversation may be provided. For example, in the mode that prioritizes text-to-speech, text-to-speech continues automatically when there is no conversation input from the user. When there is conversation input from the user, the system automatically shifts to a conversation with the AI and text-to-speech is suppressed.
[0369] In the second embodiment, as in the first embodiment, a priority / order of precedence may be set for the multiple cameras of the device 1. That is, the contents of surrounding objects, etc. recognized from the image of a camera with a higher priority will be preferentially reflected in the conversation of the conversational AI.
[0370] In the second embodiment, similarly to the first embodiment, control may be applied according to whether communication between the device 1 and the server 2 is possible or not.
[0371] As a variant, it is also possible to use an image from only one camera equipped in the device 1. For example, only the camera C2 (in-camera) of the smartphone 1A may be used. The image taken by the in-camera shows the user's face and the surrounding situation (surrounding objects) behind it. The server 2 (conversational AI) recognizes at least one of the facial expression and behavior and the surrounding objects from the camera image. The conversational AI reflects the recognized content in the response conversation.
[0372] As a modified example, in a configuration as shown in FIG. 36B, device 1 recognizes surrounding objects from camera images using image recognition AI 357 and obtains text representing the surrounding objects. Then, conversation client 353 may add a prompt with text representing the surrounding objects as an auxiliary prompt to the prompt of the input conversational sentence. The request data sent from device 1 to server 2 may be, for example, request b4' of a modified example shown at the bottom of FIG. 38, which includes prompt 3801 of the input conversational sentence and auxiliary prompt 3803. The conversation client may create a prompt that combines prompt 3801 of the input conversational sentence and auxiliary prompt 3803 into one. The conversation AI on server 2 generates a response conversational sentence from request b4' including such a prompt.
[0373] As a variant, the conversation AI may generate, as candidates for the input conversation sentence, an output conversation sentence (first conversation sentence) when the surrounding objects recognized from the camera image are not reflected as the topic, and an output conversation sentence (second conversation sentence) when the surrounding objects recognized from the camera image are reflected as the topic. The conversation AI selects and outputs an output conversation sentence from the candidates based on a predetermined judgment / evaluation. The flow of the conversation changes depending on the selection. Facial expressions may be used as the predetermined judgment / evaluation. For example, if the facial expression is happy / entertaining, the conversation on the topic of the surrounding objects may be developed, and if the facial expression is anger / sadness, the conversation on the topic of the surrounding objects may be stopped or switched to another topic.
[0374] Although the embodiments of the present disclosure have been specifically described above, they are not limited to the above-described embodiments and can be modified in various ways without departing from the spirit of the present disclosure. Except for essential components, components can be added, deleted, or replaced in each embodiment. Unless otherwise specified, each component can be singular or plural. A combination of each embodiment and its variations is also possible. [Explanation of symbols]
[0375] 1...Device (image reading device), 2...Server device, 3...Cloud computing system, 9...Communication network, 20...AI, U1...User.
Claims
1. An image reading system including a device carried or worn by a user, which reads out an object from an image captured by a camera by voice, The device automatically captures images from the camera repeatedly at predetermined times; Analyzing the captured image to obtain information, including text, describing objects in the image; Based on the acquired information, a predetermined judgment is made to determine the object and text to be read aloud; The device automatically and repeatedly reads out the text representing the determined object at a predetermined timing. Image reading system.
2. 2. The image reading system according to claim 1, The repeated image capture at the predetermined timing is continuously performed at a predetermined first time interval even without an operation by the user, The repeated voice reading at the predetermined timing is continuously executed at a predetermined second time interval even without an operation by the user. Image reading system.
3. 2. The image reading system according to claim 1, determining changes in the object between a plurality of captured images on a time axis, and determining the object to be read out in such a way that a priority is given to an object with a relatively large change; Image reading system.
4. 2. The image reading system according to claim 1, determining a relative distance or speed between the device and an object in the image based on the analysis of the image, and determining the object to be read out aloud so as to prioritize an object whose relative distance is within a threshold or whose relative speed is equal to or greater than a threshold; Image reading system.
5. 2. The image reading system according to claim 1, Based on the analysis of the image, the possibility of contact between the user and the object is determined in consideration of the moving direction of the object, and in order to avoid the contact, a direction in which to move the user to an empty space where the possibility of contact with the object is low is determined, and guidance is given to the user to move in a direction that will move the user to the empty space. Image reading system.
6. 5. The image reading system according to claim 4, determining that the closer the object is to the user, the smaller the relative distance or the greater the relative speed; For an object whose degree of approach is large, the volume of the voice reading is increased, or the tone is changed, or an alert sound is given, according to the degree of approach. Image reading system.
7. 2. The image reading system according to claim 1, A specific object is set in advance as the object to be read aloud, determining an object to be read aloud so as to give priority to the specific object; The specific objects include crosswalks, traffic lights, and road signs. Image reading system.
8. 8. The image reading system according to claim 7, determining a signal state of the traffic light as the specific object based on the analysis of the image; If it is determined that the signal has changed, the device reads out a voice representing the change in the signal. Image reading system.
9. 2. The image reading system according to claim 1, When the object is repeatedly read aloud at a predetermined timing, if there is no significant change in the same object on a time axis, control is performed so that the same object is not read aloud. Image reading system.
10. 2. The image reading system according to claim 1, When the object is repeatedly read aloud at a predetermined timing, if there is no significant change in the same object on the time axis, the system controls to perform a read aloud indicating that there is no change in the same object. Image reading system.
11. 2. The image reading system according to claim 1, When the user inputs a reading instruction to the device, the device performs the following operations in accordance with the reading instruction, in addition to the automatic image capture and voice reading at the predetermined timing; The device captures an image from the camera's video feed; Analyzing the captured image to obtain information, including text, describing objects in the image; Based on the acquired information, determine the object and text to be read aloud; and reading aloud from said device text representing the determined object. Furthermore, when the priority execution is performed, even if there is no significant change in the object, the text representing the object is always read aloud. Image reading system.
12. 2. The image reading system according to claim 1, When the automatic voice reading is performed at a predetermined timing, or when the user inputs a reading instruction into the device, When the same object is read aloud for the second or subsequent times, the reading is controlled so that the text is read aloud using a different expression than the text representing the object when read aloud the first time. Image reading system.
13. 2. The image reading system according to claim 1, the camera of the device has a 360-degree camera; Acquire front, rear, left, and right images from the video of the camera capturing 360 degrees, based on the orientation of the user. Image reading system.
14. 2. The image reading system according to claim 1, the camera of the device has a plurality of cameras with different shooting directions, The device automatically captures images repeatedly from the images of the plurality of cameras at predetermined times. Image reading system.
15. 15. The image reading system according to claim 14, The plurality of cameras include a front camera that captures a forward image and a rear camera that captures a rear image based on the orientation of the user. Image reading system.
16. 15. The image reading system according to claim 14, The plurality of cameras include a front camera that captures a front image, a rear camera that captures a rear image, a left camera that captures a left image, and a right camera that captures a right image based on the orientation of the user. Image reading system.
17. 2. The image reading system according to claim 1, the camera of the device has a plurality of cameras with different shooting directions, for each of the plurality of cameras, a predetermined timing for capturing the image and a predetermined timing for performing a voice readout based on the image of the camera are set; Image reading system.
18. 18. The image reading system according to claim 17, a first cycle is set as the predetermined timing for the capture and the voice reading in a first camera among the plurality of cameras, and a second cycle longer than the first cycle is set as the predetermined timing for the capture and the voice reading in a second camera; When the object or a significant change in the object is detected from the image of the second camera, the second cycle is set for the first camera and the first cycle is set for the second camera. Image reading system.
19. 16. The image reading system according to claim 15, A first cycle is set for the front camera as a predetermined timing for the capture and voice reading, and a second cycle, which is the same as the first cycle, is set for the rear camera as a predetermined timing for the capture and voice reading, and control is performed so that the capture and voice reading by the front camera and the capture and voice reading by the rear camera are processed at alternate timings. Image reading system.
20. 2. The image reading system according to claim 1, the camera of the device has a plurality of cameras with different shooting directions, The device automatically captures images repeatedly from the images of the plurality of cameras at predetermined times; determining changes in the object between the plurality of captured images on a time axis, and determining the object to be read out in such a way that a priority is given to an object with a relatively large change; setting a threshold value for determining the magnitude of the change for each of the plurality of cameras; Image reading system.
21. 2. The image reading system according to claim 1, the camera of the device has a plurality of cameras with different shooting directions, The device automatically captures images repeatedly from the images of the plurality of cameras at predetermined times; a voice having a different tone is used for each of the plurality of cameras when performing the voice reading based on the image of the camera; Image reading system.
22. 2. The image reading system according to claim 1, the device has a three-dimensional audio output device for performing the voice reading; When the object is read aloud, three-dimensional sound is output from the three-dimensional sound output device so that the sound sounds to the user as if it is generated from the position of the object in three-dimensional space. Image reading system.
23. 2. The image reading system according to claim 1, The device is at least one of a smartphone, a smartwatch, a tablet terminal, smart glasses, and a head-up display. Image reading system.
24. 2. The image reading system according to claim 1, a server device that is communicatively connected to the device; The server device performs a process of analyzing the captured image to obtain information including text representing an object in the image. Image reading system.
25. 2. The image reading system according to claim 1, When the device is able to communicate with the server device, the device causes the server device to perform the analysis, and when the device is unable to communicate with the server device, the device performs a simplified analysis as the analysis; The simplified analysis is an analysis performed at a longer cycle than the cycle of the analysis performed by the server device, or an analysis using images from a smaller number of cameras than the number of cameras used in the analysis performed by the server device. Image reading system.
26. An image reading-out system including a device that is carried by or worn by a user and reads out aloud an object from an image captured by a camera, the device being an image reading-out device, The device automatically captures images from the camera repeatedly at predetermined times, the image reading system analyzes the captured image to obtain information including text describing objects in the image; The image reading system determines, based on the acquired information, the object and text to be read aloud through a predetermined judgment; the device automatically and repeatedly reads out the text representing the determined object at a predetermined timing; Image reading device.
27. 1. An image reading method in an image reading system including a device carried or worn by a user, which reads out an object from an image captured by a camera by voice, comprising: The device automatically captures images from the camera's video at predetermined intervals; the image reading system analyzing the captured image to obtain information including text describing objects in the image; The image reading system determines an object and text to be read aloud based on the acquired information by a predetermined judgment; a step of automatically and repeatedly reading out a text representing the determined object at a predetermined timing by the device; An image reading method comprising:
28. A response output system comprising a device carried or worn by a user, and responding to the user, The device comprises: Equipped with multiple cameras with different shooting ranges, automatically and repeatedly capturing a plurality of images from the plurality of cameras at set timings; creating an input conversation sentence in a conversation based on the input by the user; inputting the input conversation and the plurality of images into an interface; The interface Recognizing a surrounding situation of the user based on image recognition processing of the plurality of images; outputting an output conversation sentence generated in response to the input conversation sentence; Response output system.
29. 29. The response output system according to claim 28, the interface inserts the generated output dialogue into the conversation flow, the output dialogue talking about an ambient object detected from the ambient situation in an image captured by at least one camera among the plurality of images. Response output system.
30. 29. The response output system according to claim 28, the interface inserts the generated output dialogue into the conversation flow, the output dialogue relating to a facial expression of the user recognized from an image captured by at least one camera among the plurality of images. Response output system.
31. 29. The response output system according to claim 28, The interface outputting the generated output conversation, the output conversation being about a surrounding object detected from the surrounding situation in an image captured by at least one camera among the plurality of images; Furthermore, the output conversation is inserted into the flow of conversation by controlling whether to output an output conversation that talks about the peripheral object or whether to change the output conversation to a conversation that talks about another peripheral object, depending on the facial expression of the user recognized from an image captured by at least one camera among the plurality of images. Response output system.
Citation Information
Patent Citations
Visual sense assisting device
JP2000325389A
Support device for visually handicapped person
JP2006251596A
Device for the visually challenged
JP2017077446A