Text-to-speech system and text-to-speech device
The voice reading system engages children and families by converting images into voice output with background music and sound effects, addressing the lack of playfulness in existing systems.
Patent Information
- Application Number
- JP2021190285
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-02
- Filing Date
- 2021-11-24
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2041-11-24
AI Technical Summary
Existing text-to-speech systems lack engagement for children and their families, as they primarily focus on reading text without incorporating playful elements.
A voice reading system with an imaging device and main device that captures images, analyzes text and graphics, and outputs voice based on the analysis, incorporating background music and sound effects related to the content, allowing for a fun and educational experience.
Provides an engaging and educational experience for children and their families by converting images into voice output, including background music and sound effects, enhancing learning through interaction.
Smart Images

Figure 0007790109000001 
Figure 0007790109000002 
Figure 0007790109000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a text-to-speech system. and Voice reading device Place Regarding. [Background technology]
[0002] Patent Document 1 discloses application software that uses the camera function of a smartphone to recognize characters on printed matter and reads the recognized characters aloud, thereby making it possible to recognize small characters. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2014-127197 Summary of the Invention [Problem to be solved by the invention]
[0004] However, while the application software in Patent Document 1 is useful for users with myopia or presbyopia because it makes printed materials easier to read, it simply reads out text, so it lacks a playful element and is not of interest to children or their families.
[0005] The present invention has been made in view of the above circumstances, and provides a text-to-speech system that can be enjoyed by children and their families. and Voice reading device Place The purpose is to provide. [Means for solving the problem]
[0006] The present application includes multiple means for solving the above-mentioned problems, and as one example, a voice reading system includes an imaging device and a main device having a mounting surface on which an object to be imaged can be placed, the imaging device includes a gripping portion, a window portion for looking into the object to be imaged, an imaging unit capable of imaging the object to be imaged while the object is being looked into through the window portion, and a transmitting portion for transmitting image data obtained by imaging with the imaging unit to the main device, and the main device includes a receiving portion for receiving the image data, an analyzing portion for analyzing the image data received by the receiving portion, and an output portion for outputting voice based on the analysis result of the analyzing portion. [Effects of the Invention]
[0007] The present invention provides enjoyment for children and their families. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a schematic perspective view showing an example of the configuration of a text-to-speech system according to a first embodiment. [Figure 2] 1 is a block diagram showing an example of the configuration of a text-to-speech system according to a first embodiment. [Figure 3] FIG. 10 is a schematic diagram showing an example of the configuration of a BGM list. [Figure 4] FIG. 10 is an explanatory diagram showing an example of a sound. [Figure 5] 10A and 10B are schematic diagrams illustrating an example of a method for correcting an imaging range by a correction unit. [Figure 6] 4 is a flowchart illustrating an example of a processing procedure of the text-to-speech system according to the first embodiment. [Figure 7] FIG. 1 is a block diagram showing an example of the configuration of a text-to-speech device. [Figure 8] FIG. 10 is a block diagram showing an example of the configuration of a text-to-speech system according to a third embodiment. [Figure 9] 11 is a flowchart illustrating an example of a processing procedure of the text-to-speech system according to the third embodiment. [Figure 10]FIG. 10 is a diagram illustrating an example of a configuration of an information processing system according to a fourth embodiment. [Figure 11] FIG. 13 is a diagram illustrating an example of a processing procedure of an information processing system according to a fourth embodiment. [Figure 12] FIG. 10 is a diagram illustrating an example of a processing procedure of an interest analysis function. [Figure 13] FIG. 10 is a diagram showing an example of an analysis result of an interest analysis function. [Figure 14] FIG. 10 is a diagram illustrating an example of a processing procedure of an interest type analysis function. [Figure 15] FIG. 10 is a diagram showing an example of an analysis result of an interest type analysis function. [Figure 16] FIG. 10 is a diagram illustrating an example of a processing procedure of an activity type analysis function. [Figure 17] FIG. 10 is a diagram showing an example of an analysis result of the activity type analysis function. [Figure 18] FIG. 10 is a diagram illustrating an example of a processing procedure of a favorite color analysis function. [Figure 19] FIG. 10 is a diagram showing an example of an analysis result of a favorite color analysis function. [Figure 20] FIG. 10 is a diagram illustrating an example of a configuration of an information processing system according to a fifth embodiment. [Figure 21] FIG. 13 is a diagram illustrating an example of a processing procedure of an information processing system according to a fifth embodiment. [Figure 22] FIG. 13 is a diagram illustrating an example of a configuration of an information processing system according to a sixth embodiment. [Figure 23] FIG. 13 is a diagram showing an example of a processing procedure of the imaging device 50 of the sixth embodiment. [Figure 24] FIG. 10 is a diagram showing an example of a transition of interest analysis results. [Figure 25] This is a diagram showing an example of the results of an interest analysis by age group, region, and time series. DETAILED DESCRIPTION OF THE INVENTION
[0009] (First embodiment) Hereinafter, embodiments of the present invention will be described with reference to the drawings. FIG. 1 is an external perspective view showing an example of the configuration of a text-to-speech system 100 according to a first embodiment. The text-to-speech system 100 includes an imaging device 50 and a main body device 10. The imaging device 50 includes a grip 62 and a window 61 provided on one end of the grip 62. The grip 62 is a portion that a user (e.g., an infant, a child, or their family) holds when holding the imaging device 50 in their hand. The window 61 may be provided with a lens (magnifying glass), transparent resin or glass, or simply formed with an opening, allowing the user to look through the window 61 into an object to be imaged (e.g., a character string such as a sentence, a diagram including a photograph, etc.) written in an object (e.g., a picture book, an illustrated encyclopedia, a children's book, etc.).
[0010] The imaging device 50 also includes an imaging unit 51 capable of capturing an image of an object when the object is viewed through a window 61, a distance sensor 52 for detecting the distance to the object, and a button (shutter button) 63 for receiving an operation to start imaging by the imaging unit 51. The imaging unit 51 and the distance sensor 52 are provided on one side of the grip 62 (the side of the object when the object is viewed through the window 61), and the button 63 is provided on the other side of the grip 62 (the side of the user's face when the object is viewed through the window 61). The imaging unit 51 can be configured with at least one camera. The distance sensor 52 may be any sensor capable of detecting distance. Instead of the distance sensor 52, distance may be measured according to the parallax of multiple cameras.
[0011] Main unit 10 has a placement surface 21 on which an object can be placed. Placement surface 21 is rectangular in plan view and is inclined so that the height decreases from each of a pair of edge portions 22 that are sandwiched between center portion 23 of placement surface 21 toward center portion 23. This allows picture books, illustrated guides, children's books, etc. to be placed on placement surface 21 in a spread-open state, and the pages can be placed so that an angle between the pages is less than 180 degrees, allowing the object to be placed so that the document or diagram is easy to read.
[0012] Main unit 10 is provided with restricting portions 24 that restrict the movement of an object on another pair of edge portions of placing surface 21 along the inclined direction of placing surface 21. Restricting portions 24 are provided in a state of protruding from placing surface 21. This makes it possible to prevent picture books, illustrated guides, children's books, etc. placed on placing surface 21 from slipping off placing surface 21.
[0013] The main body 10 has a storage section 25 formed on the placement surface 21 for storing the imaging device 50. The shape of the storage section 25 in a plan view can be the same as the shape of the imaging device 50 in a plan view. The imaging device 50 may be configured to fit into the storage section 25, or the two may be attracted to each other using a magnet or the like. This prevents the imaging device 50 from being lost and helps infants and children to develop the habit of tidying up after themselves.
[0014] The main device 10 may be provided with an indicator light (e.g., an LED) 26 that displays the status of the main device 10. The indicator light 26 can display the status of the main device 10, such as powered, battery-powered, operating, charging, or abnormality. Although not shown, a touch-operable display panel may also be provided. Required setting operations may be performed via the display panel.
[0015] 2 is a block diagram showing an example of the configuration of the text-to-speech system 100 of the first embodiment. In addition to the above-mentioned imaging unit 51 and distance sensor 52, the imaging device 50 includes a correction unit 53, a memory 54, a processor 55, and a communication unit 56. The processor 55 can control the entire imaging device 50. The memory 54 is configured with a semiconductor memory or the like, and can store image data obtained by capturing an image with the imaging unit 51.
[0016] Communication unit 56 realizes a communication function with main device 10 via home network 1 such as a wireless LAN. Imaging device 50 (e.g., processor 55) can transmit image data (including image data temporarily stored in memory 54) captured by imaging unit 51 to main device 10 via communication unit 56.
[0017] The correction unit 53 can correct the imaging range of the imaging unit 51 so that an imaging target within the field of view through the window unit 61 can be imaged according to the distance detected by the distance sensor 52. The correction unit 53 will be described in detail later.
[0018] The main unit 10 includes a control unit 11, a communication unit 12, an analysis unit 13, a voice synthesis unit 14, a sequence estimation unit 15, a memory unit 16, a microphone 17, a speaker 18, and an emotion index calculation unit 19. The analysis unit 13 includes a character string analysis unit 131 and a graphic analysis unit 132. The memory unit 16 is configured, for example, from a semiconductor memory, and can store a background music list 161 and an audio data list 162. The control unit 11 can be configured with a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), etc.
[0019] The communication unit 12 realizes a function of communicating with the imaging device 50 via the home network 1. The communication unit 12 can receive image data transmitted by the imaging device 50.
[0020] The analysis unit 13 analyzes the image data received via the communication unit 12. Specifically, the analysis unit 13 can perform image recognition on the image data received via the communication unit 12 to analyze whether the imaged object is a character string or a figure. For image recognition, for example, a known method can be used, and processing such as preprocessing, feature extraction, matching, and classification can be performed.
[0021] The character string analysis unit 131 is equipped with an image processing engine and a language processing engine, and can analyze image data and output character strings (text). The process of extracting character strings from image data may use, for example, a known method.
[0022] The image analysis unit 132 is equipped with an image processing engine and can analyze image data to determine what the images (including photographs) contained in the images represent. For example, the objects to be imaged include, but are not limited to, vehicles such as trains and automobiles, animals, insects, and musical instruments.
[0023] The control unit 11 can output a sound via the speaker 18 based on the analysis result of the analysis unit 13 .
[0024] With the above-described configuration, infants and children can look at an object that catches their eye in a picture book, encyclopedia, or children's book through the window 61 of the imaging device 50 and press the button 63 to hear a voice explain the object they are looking at. This allows infants and children to have an experience using voice, which sparks their interest in objects and provides a fun experience. Furthermore, rather than simply playing with objects, by looking at something that infants and children are curious about or interested in, an audio experience is provided, providing a function that leads to new discoveries and providing the "fun of learning." Furthermore, family members (e.g., parents) can have a fun experience together with infants and children.
[0025] Next, first to third methods for preparing audio to be output will be explained.
[0026] In the first method, text (sentences) from picture books, illustrated guides, children's books, etc. are recorded in advance, and the recorded audio is stored in the storage unit 16 as an audio data list 162. The audio data list 162 associates information indicating the text with audio data for each book, such as a picture book, illustrated guide, or children's book. Audio data corresponding to the character string (text) analyzed by the character string analysis unit 131 is output from the speaker 18, enabling "audio reading." When playing back recorded audio, the speaker's attributes may be changed. For example, a male or female voice, a young or elderly voice, or the voice of an anime voice actor may be set according to preference. Such settings may be made using a touch-operable display panel. Alternatively, a parent's voice may be recorded using the microphone 17, and the recorded audio may be played back.
[0027] The second and third methods can be performed by the speech synthesis unit 14. First, the second method synthesizes speech by concatenating pre-recorded speech fragments. Specifically, speech can be synthesized by concatenating recorded characters (e.g., "a" and "ka"), words, and phrases. In this case, speech can be adjusted to sound more natural by adjusting the speaking rate, voice pitch, intonation (tone, intonation), etc. Corpus-based speech synthesis may also be used. Corpus-based speech synthesis is a technique that predicts fundamental frequency, phoneme duration, etc. based on linguistic features of text, such as sentences, phrases, accent phrases, morphemes, phonemes, and accents, and selects and concatenates speech fragments that best match the predicted fundamental frequency, phoneme duration, etc. from a speech database prepared in advance.
[0028] The third method synthesizes speech using speech features of pre-recorded speech. Specifically, the speech synthesis unit 14 includes a trained model that has learned speech features from recorded speech. The speech synthesis unit 14 can convert speech into a speech waveform based on the speech features output by the trained model. The speech features include, for example, Mel-Frequency Cepstrum Coefficients (MFCC), Line Spectral Pairs (LSP), and fundamental frequency.
[0029] The control unit 11 can output BGM (also called background music) via the speaker 18 based on the analysis result of the analysis unit 13. This will be specifically described below.
[0030] The emotion index calculation unit 19 can calculate an emotion index by performing semantic analysis on the character string extracted by the character string analysis unit 131. For example, the emotion index calculation unit 19 can extract words expressing emotions from the character string and calculate an emotion index based on the extracted words. Emotions are divided into positive emotions and negative emotions, and each word expressing an emotion is determined in advance as being positive or negative. Furthermore, a value indicating the intensity of the emotion is determined for each word expressing an emotion. The emotion index calculation unit 19 can calculate an emotion index for the extracted character string based on the value of the intensity of the emotion, whether each extracted word is positive or negative. The emotion index calculation unit may calculate an emotion index for each extracted character string, for example, based on the value of the intensity of the emotion.
[0031] The control unit 11 can use the BGM list 161 stored in the storage unit 16 to output BGM from the speaker 18 according to the emotion index calculated by the emotion index calculation unit 19.
[0032] FIG. 3 is a schematic diagram showing an example of the configuration of the BGM list 161. The BGM list 161 defines a correspondence between an emotion index and BGM. The emotion index can be, for example, positive levels 1 to 3 and negative levels 1 to 3. The higher the level number, the stronger the emotion. Multiple BGMs are associated with each level. For example, as shown in the figure, BGM1a, BGM1b, BGM1c, and BGM1d are associated with positive level 1. The same applies to other emotion indexes. Here, the symbols a to d identify different scenes in a story or text found in a picture book or the like. For example, when the emotion index is positive level 1, the BGM can be changed accordingly when the scene changes. This creates a sense of realism, making the audio experience more enjoyable and exciting.
[0033] Furthermore, the control unit 11 can output a sound (also referred to as a sound) via the speaker 18 based on the analysis result of the analysis unit 13. Specifically, the control unit 11 can output a sound related to the analysis result (indicating what the figure (including a photograph) included in the image represents) analyzed by the figure analysis unit 132 via the speaker 18.
[0034] FIG. 4 is an explanatory diagram showing an example of sound. As shown in the figure, the imaged object (a picture included in the image) can be, for example, a train, a car, an animal, an insect, a musical instrument, etc. If the picture is a train, the sound of the train running is output. If the picture is an animal, the animal's cry is output. If the picture is a musical instrument, the sound of the instrument is output. In this way, when an infant or child opens an illustrated book or the like and looks into a picture that interests them through the window 61, a sound related to the picture they see is played. By looking into an object that interests or has questions about it, infants and children can find out what kind of sound the object makes, which can lead them to new discoveries.
[0035] The order estimation unit 15 can estimate the reading order of character strings based on the arrangement (layout, etc.) of character strings extracted by the character string analysis unit 131. For example, if the document layout is vertical, the reading order of character strings can be from top to bottom, and if the document layout is horizontal, the reading order of character strings can be from left to right. This allows the voice to be read aloud whether the object is written vertically or horizontally.
[0036] Next, the correction unit 53 will be described.
[0037] FIG. 5 is a schematic diagram showing an example of a method for correcting the imaging range by the correction unit 53. The correction unit 53 can correct the imaging range so that the imaging target within the range (field of view) visible within the frame of the window 61 can be captured, depending on the distance between the window 61 and the object, such as a picture book. As shown in the figure, when the position of the window 61 relative to the surface of the picture book is P1, the range visible when looking through the window 61 is defined as S1. Note that the imaging range S1 is rectangular, while the window 61 is circular. Therefore, the field of view is actually circular. However, for convenience, the imaging range S1 is defined as a rectangle in which the circle is an inscribed circle. When the window 61 is moved slightly away from the picture book and the surface of the picture book is viewed at position P2, the distance between the window 61 and the picture book increases, and the range visible when looking through the window 61 increases. Therefore, the imaging range S2 is corrected to be larger than the imaging range S1.
[0038] Next, the operation of the text-to-speech system 100 will be described.
[0039] 6 is a flowchart showing an example of the processing procedure of the text-to-speech system 100 of the first embodiment. The imaging device 50 accepts a shutter button operation (S11), detects the distance to the object (S12), captures an image of the object (S13), and corrects the captured image data according to the distance to the object (S14). The imaging device 50 transmits the captured image data to the main device 10 (S15), and the process ends. Note that the imaging data may be transmitted together with the subsequent analysis range without being corrected.
[0040] Main device 10 receives the image data (S16) and determines whether the image capture target is a character string (S17). If the image capture target is a character string (YES in S17), main device 10 reads out the character string (S18). Main device 10 calculates an emotion index based on the character string (S19) and outputs background music corresponding to the emotion index (S20).
[0041] Main device 10 determines whether the scene has changed based on the character string (S21), and if the scene has changed (YES in S21), changes the background music and outputs it (S22), and ends the process. If the scene has not changed (NO in S21), main device 10 ends the process.
[0042] In step S17, if the imaging target is not a character string (NO in S17), main device 10 determines whether the imaging target is a figure (S23). If the imaging target is a figure (YES in S23), main device 10 outputs a sound related to the content of the figure that is the imaging target (S24), as illustrated in Fig. 4, and ends the process. If the imaging target is not a figure (NO in S23), main device 10 ends the process.
[0043] (Second embodiment) In the first embodiment, the text-to-speech system 100 was configured to include an imaging device 50 and a main body device 10, but in the second embodiment, the functions of the main body device 10 are incorporated into the imaging device 50 to form a text-to-speech device.
[0044] 7 is a block diagram showing an example of the configuration of a text-to-speech device 200. The external shape of the text-to-speech device 200 is similar to that of the imaging device 50 of the first embodiment, and includes a grip, a shutter button, and a window. The text-to-speech device 200 includes a control unit 201, an imaging unit 202, a distance sensor 203, a correction unit 204, an analysis unit 205 having a character string analysis unit 206 and a graphic analysis unit 207, a voice synthesis unit 208, an order estimation unit 209, a storage unit 210 that stores a background music list 211 and a voice data list 212, a microphone 213, a speaker 214, and an emotion index calculation unit 215. The functions of the control unit 201, the imaging unit 202, the distance sensor 203, the correction unit 204, the analysis unit 205 having the character string analysis unit 206 and the graphic analysis unit 207, the voice synthesis unit 208, the order estimation unit 209, the memory unit 210 that stores the BGM list 211 and the voice data list 212, the microphone 213, the speaker 214, and the emotion index calculation unit 215 are the same as in the first embodiment, so description thereof will be omitted.
[0045] (Third embodiment) The third embodiment includes an imaging device 50 and a main body device 30, as in the first embodiment, and further includes a server 300, in which the main functions of the main body device 10 of the first embodiment are incorporated.
[0046] 8 is a block diagram showing an example of the configuration of a text-to-speech system according to the third embodiment. The imaging device 50 has the same configuration as the imaging device 50 according to the first embodiment. The main device 30 includes a control unit 36, a first communication unit 31, a second communication unit 32, a storage unit 33 that stores a background music list 331, a microphone 34, and a speaker 35.
[0047] The first communication unit 31 realizes a communication function with the imaging device 50 via the home network 1.
[0048] The second communication unit 32 realizes a communication function with the server 300 via a communication network 2 such as the Internet. The BGM list 331, microphone 34, and speaker 35 are the same as those in the first embodiment, and therefore description thereof will be omitted.
[0049] The server 300 includes a control unit 301, a communication unit 302, an analysis unit 303 having a character string analysis unit 304 and a graphic analysis unit 305, a voice synthesis unit 306, an order estimation unit 307, a memory unit 308 that stores a voice data list 309, and an emotion index calculation unit 310.
[0050] The communication unit 302 realizes a communication function with the main device 30 via the communication network 2. The functions of the control unit 301, analysis unit 303, character string analysis unit 304, graphic analysis unit 305, speech synthesis unit 306, order estimation unit 307, speech data list 309, and emotion index calculation unit 310 are the same as those in the first embodiment, so description thereof will be omitted.
[0051] 9 is a flowchart showing an example of the processing procedure of the text-to-speech system of the third embodiment. The processing of the image capture device 50 is the same as that of the first embodiment, so a description thereof will be omitted. The main device 30 receives the image data transmitted by the image capture device 50 (S31), and transmits the received image data to the server 300 (S32).
[0052] The server 300 receives the image data (S33) and determines whether the image capture target is a character string (S34). If the image capture target is a character string (YES in S34), the server 300 converts the character string into audio data and transmits the converted audio data to the main device 30 (S35). The main device 30 outputs audio (S36). This allows the character string to be read aloud.
[0053] Server 300 calculates an emotion index based on the character string (S37) and transmits the calculated emotion index to main device 30 (S38). Main device 30 receives the emotion index (S39) and outputs background music according to the received emotion index (S40).
[0054] Server 300 determines whether the scene has changed based on the character string (S41), and if the scene has changed (YES in S41), notifies main device 30 of the scene change (S42). Main device 30 changes the background music and outputs it (S43). If the scene has not changed (NO in S41), server 300 ends the process.
[0055] In step S34, if the imaging target is not a character string (NO in S34), server 300 determines whether or not the imaging target is a drawing (S44). If the imaging target is a drawing (YES in S44), server 300 transmits sound data relating to the content of the drawing to main device 30 (S45) and ends the process. Main device 30 receives the sound data and outputs sound based on the received sound data (S46), and ends the process. If the imaging target is not a drawing (NO in S44), server 300 ends the process.
[0056] The text-to-speech system and text-to-speech device of this embodiment can be used for infants and children to play at home, at facilities such as nurseries and kindergartens, at temporary care facilities (for example, shopping malls, supermarkets, beauty salons, etc.), etc. Furthermore, users are not limited to infants and children, and the elderly, foreigners, and people with disabilities may also use the text-to-speech system and text-to-speech device of this embodiment.
[0057] The imaging device, main unit, and text-to-speech device of the text-to-speech system of this embodiment can be sold or rented to users as toys. Also, the server functions can be provided on the cloud and a usage fee can be collected as payment.
[0058] (Fourth embodiment) In the fourth embodiment, we will explain an information processing device system, information processing device, imaging device, etc. that uses functions similar to those of the above-mentioned voice reading system 100 to allow parents and other guardians to understand the interests of their children, such as young children (especially the latest interests and changes in interests due to growth or environmental changes, etc.).
[0059] In recent years, the number of dual-income households has increased, and more children are being sent to daycare centers, making it difficult for parents to observe their children's daytime behavior and to grasp what their children are beginning to be interested in, or if they even have any. Parents are unable to properly grasp the growth and changes in their children that occur little by little every day.
[0060] Furthermore, approaches to fostering intellectual curiosity in early childhood education are often one-size-fits-all or broad-based approaches that are suited to many children, and do not take into account the changes in each individual's intellectual curiosity, interests, and preferences. In the following embodiment, in order to solve these problems, support for fostering interests tailored to each individual is described, taking into account the real-time state of interests and preferences, as well as the child's growth and environmental changes.
[0061] FIG. 10 is a diagram showing an example of the configuration of an information processing system according to the fourth embodiment. The information processing system includes an imaging device 50, a terminal device 100, and a server 400 as an information processing device. The imaging device 50 has the same configuration as in the first embodiment, but differs in that a speaker 57 is incorporated into the imaging device 50. The imaging device 50 can capture an image of an imaging target in a reading mode or a sound effect playback mode. The reading mode or the sound effect playback mode may be switched manually using a switch or button, or may be automatically switched based on the analysis results of the analysis unit 403 of the server 400. The imaging device 50 is a magnifying glass-like device that infants and children play with, and the terminal device 100 is a terminal carried by a parent or guardian.
[0062] The terminal device 100 has a configuration similar to that of the main device 30 of the third embodiment, but differs in that it includes a display unit 104 and an operation unit 105. The control unit 101, the first communication unit 102, the second communication unit 103, and the storage unit 106 are similar to the control unit 36, the first communication unit 31, the second communication unit 32, and the storage unit 33 of the main device 30 of the third embodiment. The terminal device 100 may be configured as, for example, a smartphone, a tablet terminal, or the like. The display unit 104 may be configured as a liquid crystal display or an organic EL (Electro Luminescence) display. The operation unit 105 may be configured as a touch panel or the like, and may be used to input characters on the display unit 104 and to operate icons, images, characters, or the like displayed on the display unit 104.
[0063] The server 400 includes a control unit 401 that controls the entire server 400, a communication unit 402, an analysis unit 403, an order estimation unit 406, a voice synthesis unit 407, a sound reproduction unit 408, an analysis unit 409, and a storage unit 410. The analysis unit 403 includes a character string analysis unit 404 and a graphic analysis unit 405. The storage unit 410 stores a voice data list 411. The control unit 401, the communication unit 402, the analysis unit 403, the character string analysis unit 404, the graphic analysis unit 405, the order estimation unit 406, the voice synthesis unit 407, and the storage unit 410 are similar to the control unit 301, the communication unit 302, the analysis unit 303, the character string analysis unit 304, the graphic analysis unit 305, the order estimation unit 307, the voice synthesis unit 306, and the storage unit 308 of the server 300 of the third embodiment. The analysis unit 403 may automatically switch between a sound effect playback mode and a reading mode depending on whether the image capture target is a figure or a character string. The sound playback unit 408 has the same sound playback function as the control units 11, 301, etc. Details of the analysis unit 409 will be described later.
[0064] 11 is a diagram showing an example of a processing procedure of the information processing system of the fourth embodiment. The imaging device 50 captures an image of an object to be imaged (S101) and transmits image data obtained by the image capture to the terminal device 100 (S102). The terminal device 100 receives the image data (S103) and transmits the received image data to the server 400 (S104).
[0065] The server 400 receives the image data (S105) and analyzes the imaged object based on the received image data (S106). The server 400 generates a sound related to the content of the figure or a voice reading out the character string, depending on whether the imaged object is a figure or a character string (S107). The sound related to the content of the figure and the voice reading out the character string may be collectively referred to as "voice."
[0066] The server 400 transmits the generated sound or audio to the terminal device 100 (S108). The terminal device 100 receives the sound or audio (S109) and transmits the received sound or audio to the imaging device 50 (S110). The imaging device 50 receives the sound or audio (S111) and outputs the received sound or audio (S112).
[0067] The server 400 records the received image data and analysis results in the storage unit 410 (S113), and updates the number of times the sound is played or the voice is read aloud (S114). Every time a user such as an infant or a child takes an image of an object with the imaging device 50, the process shown in Fig. 11 is repeated, and information such as the image data, analysis results, and the number of times the sound is played or the voice is read aloud can be collected.
[0068] The control unit 401 functions as a collection unit and collects image data obtained by capturing an image of an imaging target via the communication unit 402. The analysis unit 409 analyzes the interests of the user (infant or child) who captured the image of the imaging target based on the collected image data. The control unit 401 functions as a provision unit and can provide the analysis results of the analysis unit 409.
[0069] In this way, by collecting image data of subjects captured by infants and children that interest them as part of a life log, analyzing the infants' and children's daily interests based on the collected life log, and providing feedback on the analysis results to parents, parents can follow up and support their infants and children in accordance with their interests. For example, parents can buy their children goods related to areas that interest them, or take them to places or facilities that interest them.
[0070] Next, the analysis process by the analysis unit 409 will be described in detail. The analysis unit 409 has various functions, such as (1) an interest analysis function, (2) an interest type analysis function, (3) an activity type analysis function, and (4) a favorite color analysis function. The life log to be analyzed can be image data captured and collected by the imaging device 50, and the number of times the reading mode and the sound effect playback mode are used. Each analysis function will be described below.
[0071] FIG. 12 is a diagram showing an example of the processing procedure of the interest analysis function. For convenience, the following description will be given assuming that the processing is performed by the control unit 401. The control unit 401 collects image data (S121) and determines whether the data has been collected for a first predetermined period (S122). The first predetermined period can be, for example, one week, but is not limited to this. If the data has not been collected for the first predetermined period (NO in S122), the control unit 401 continues the processing of step S121.
[0072] If the images have been collected over the first predetermined period (YES in S122), the control unit 401 classifies the imaged objects by category (S123). Specifically, the analysis unit 409 classifies the imaged objects by category. The analysis unit 409 may include a learning model for object detection. Examples of the learning model include Histogram of Oriented Gradients (HOG), Region-based CNN (R-CNN), Fast R-CNN, Region Proposal Network (RPN), You Only Look Once (YOLO), Single Shot Detector (SSD), and Transformer. The objects detected by the analysis unit 409 may be classified by category. The categories may be, for example, trains, cars, airplanes, flowers, animals, food, musical instruments, fish, insects, dolls, or the like.
[0073] The control unit 401 calculates the total number of objects captured for each category (S124). For example, if a child captures 20 images of objects classified as "trains" during a week, the number of "trains" is set to 20. The control unit 401 registers the object with the most number of images as an "interesting" category (S125). For example, the number of images captured for each category is calculated, and the top five categories in descending order of the number of images captured are registered as "interesting" categories. Note that the number of "interesting" categories is not limited to five.
[0074] The control unit 401 compares the number of images taken for each category in the most recent first predetermined period (e.g., last week) and registers the category in which the number of images taken for each category in the current first predetermined period (e.g., this week) is on the rise as a "rapidly rising" category (S126). For example, the difference between the number of images taken for each category in last week and this week can be calculated, and the category in which the calculated difference is equal to or greater than a predetermined difference threshold can be registered as a "rapidly rising" category. Alternatively, the category in which the calculated difference is the largest can be registered as a "rapidly rising" category.
[0075] The control unit 401 registers the field with the most appearances among the number of images taken for each field as "my current fad" (S127). For example, the field with the most number of objects taken in one week can be set as "my current fad."
[0076] The control unit 401 transmits the analysis results ("Interested," "Rapidly Rising," "My Craze") to the terminal device 100, and the terminal device 100 displays the analysis results. In response, the control unit 401 provides the analysis results (S128) and ends the processing. Note that the words "Interested," "Rapidly Rising," and "My Craze" are merely examples, and the present invention is not limited to these words.
[0077] As described above, the analysis unit 409 functions as a first analysis unit and can analyze the field of the imaging subject based on the collected image data. The analysis unit 409 may analyze the interests of infants and children (users) based on the analysis results collected each time a first predetermined period of time passes.
[0078] The analysis unit 409 may analyze the user's "interest" (first index) based on the number of subjects captured for each analyzed field. This allows parents to provide appropriate support and encouragement to their children without missing their child's "beginning of interest."
[0079] The analysis unit 409 may analyze a "rapid rise" (second index) related to the user's interest based on a change in the number of imaging subjects captured for each field for each first predetermined period. The analysis unit 409 may also analyze a "current fad" (third index) related to the user's interest based on the field with the largest number of imaging subjects among the number of imaging subjects captured for each field.
[0080] FIG. 13 is a diagram showing an example of the analysis results of the interest analysis function. The "OO-chan's interest analysis results" screen 501 shown in FIG. 13 can be displayed on the display unit 104 of the terminal device 100. The "OO-chan's interest analysis results" screen 501 has a message display area 502 that displays an error message such as "The date is incorrect," a display area 503 that displays the ratio of subjects photographed this week by category, a display area 505 that displays "interesting" categories, a display area 506 that displays "my current obsession" categories, and a display area 507 that displays "trending" categories. Also displayed is a "view all logs" icon 504 for viewing all logs.
[0081] In the example of FIG. 13, the categories of "interested" are displayed as "trains," "flowers," "cars," "food," and "animals," and the number of photographs taken in each category is displayed as 20, 18, 15, 10, and 9, respectively. In "My Craze," for example, the object with the most photographs (in the example of FIG. 13, an image of a "Bullet Train") is displayed from the category of "trains," which has the most photographs, along with a message such as "My current craze is the Bullet Train!" This allows parents to easily grasp their child's real-time interests and tastes. In "Trending," a message such as "It seems he's recently become interested in 'food'" is displayed. This allows parents to appropriately grasp changes in interests and tastes, which can be useful in daily communication with their children and in their daily lives (such as purchasing activities).
[0082] FIG. 14 is a diagram showing an example of the processing procedure of the interest type analysis function. The control unit 401 compares the "interested" rankings for the most recent first predetermined period (e.g., last week) with the "interested" rankings for the current first predetermined period (e.g., this week) (S131) and determines whether there has been a change in the top rankings (S132). For example, as shown in FIG. 13, if the "interested" rankings are displayed from 1st to 5th, the top rankings can be the 1st and 2nd rankings, but are not limited to this. If the top two rankings from last week were, for example, "animals" and "flowers," and the top two rankings from this week were, for example, "trains" and "flowers," then the top ranking has changed from "animals" to "trains" from last week to this week, and in this case, it can be determined that there has been a change in the top rankings.
[0083] If there is a change in the top rankings (YES in S132), the control unit 401 classifies the user's interest type as the "curious" type (S133), provides the analysis result (interest type) to the terminal device 100 (S134), and terminates the processing.
[0084] If there is no change in the top rankings (NO in S132), the control unit 401 determines whether there is no change in the rankings (S135). If there is no change in the rankings (YES in S135), that is, if there is no change in the rankings from 1st to 5th between last week and this week, the control unit 401 classifies the user's interest type as the "thoughtful doctor" type (S136), and performs the process of step S134.
[0085] If there is a change in the ranking (NO in S135), that is, if there is a change in the lower rankings excluding the top rankings (e.g., rankings from 3rd to 5th), the control unit 401 classifies the user's interest type as the "intermediate" type (S137) and performs the process of step S134. Furthermore, a change in ranking may be determined not only by a change in the top or bottom rankings, but also by a change in the overall rankings. For example, a change may be determined based on how many of the rankings (e.g., the top 20) from the top to a predetermined number (the predetermined number is variable) of rankings in the weekly rankings sorted by the number of detections have changed. For example, if N% or more of the rankings have changed, the user may be determined as "curious," whereas if N% or less of the rankings have not changed, the user may be determined as "thoughtful doctor," and in all other cases, the user may be determined as "intermediate." The value N can be set as appropriate. Note that the terms "curious," "thoughtful doctor," and "intermediate" are merely examples and are not limited to these terms.
[0086] As described above, the analysis unit 409 has the function of an identification unit, and may identify changes in "interested" (first indicator) analyzed every first predetermined period, and analyze the user's type of interests according to the identified changes in "interested."
[0087] Fig. 15 is a diagram showing an example of the analysis results of the interest type analysis function. The "OO-chan's interest analysis results" screen 511 shown in Fig. 15 can be displayed on the display unit 104 of the terminal device 100. The "OO-chan's interest analysis results" screen 511 has, for example, a message display area 502 that displays an error message, a display area 512 that displays the interest type, and a display area 514 that displays this week's log.
[0088] In the example of FIG. 15, the interest type is displayed as "OO-chan is a 'curious' type and is interested in a variety of things." This allows parents to understand their child's interest type and follow up and support their child according to their interest type. By operating the "select photo" icon 513, parents can display a desired photo from among photos of their child recorded on the terminal device 100 or photos of their child uploaded from another smartphone or PC.
[0089] This week's log can display rankings of "interesting" categories, the percentage of subjects photographed this week by category, "trending," etc. By operating the "view details" icon 515, the rankings of "interesting" categories can be displayed in more detail. Also, by operating the "view last week's log" icon 516, last week's log can be displayed instead of or in addition to this week's log.
[0090] This allows parents to accurately grasp their child's interests and can use this information in their daily communication with their children and in their daily lives (such as purchasing activities).
[0091] FIG. 16 is a diagram showing an example of the processing procedure of the activity type analysis function. The control unit 401 records the number of times the sound effect playback function and the reading function are used (S141) and determines whether the recording has been performed over a third predetermined period (S142). The number of times the sound effect playback function and the reading function are used is the number of times they are used in the sound effect playback mode and the reading mode. The sound effect playback mode and the reading mode can be set manually or automatically. The third predetermined period can be, for example, one month, but is not limited to this.
[0092] If recording has not been performed over the third predetermined period (NO in S142), the control unit 401 continues the process of step S141. If recording has been performed over the third predetermined period (YES in S142), the control unit 401 determines whether the proportion of the number of times the sound effect playback function has been used is equal to or greater than M% of the total (S143). The total is the total number of times the sound effect playback function has been used and the reading function has been used. The value of M is variable and can be changed as appropriate. If the proportion of the number of times the sound effect playback function has been used is equal to or greater than M% of the total (YES in S143), the control unit 401 classifies the user's activity type as an "explorer type" (S144), provides the analysis result (activity type) to the terminal device 100 (S145), and ends the process.
[0093] If the percentage of the number of times the sound effect playback function has been used is not M% or more of the total (NO in S143), the control unit 401 determines whether the percentage of the number of times the reading function has been used is M% or more of the total (S146). If the percentage of the number of times the reading function has been used is M% or more of the total (YES in S146), the control unit 401 classifies the user's activity type as a "bookworm type" (S147) and performs the processing of step S145. If the percentage of the number of times the reading function has been used is not M% or more of the total (NO in S146), the control unit 401 classifies the user's activity type as a "very interested type" (S148) and performs the processing of step S145.
[0094] As described above, the analysis unit 403 functions as a third analysis unit and can analyze whether the imaged object is a character string or a figure based on the collected image data. The voice synthesis unit 407 functions as a reading unit and can read out the character string if the imaged object is analyzed to be a character string. The sound playback unit 408 functions as a playback unit and can play back a sound related to the content of the figure if the imaged object is analyzed to be a figure. The analysis unit 409 may analyze the user's activity type based on the number of times the voice synthesis unit 407 reads out the character string and the number of times the sound playback unit 408 plays back the character string over a third predetermined period (e.g., one month).
[0095] Fig. 17 is a diagram showing an example of the analysis results of the activity type analysis function. The "Analysis results of OO-chan's interests" screen 521 shown in Fig. 17 can be displayed on the display unit 104 of the terminal device 100. The "Analysis results of OO-chan's interests" screen 521 has, for example, a message display area 502 that displays error messages, a display area 522 that displays activity types, and a display area 514 that displays this week's log.
[0096] In the example of FIG. 17, the activity type is displayed as "OO-chan is an 'Explorer' type. She likes to play by exploring various things." This allows parents to understand their child's activity type and provide follow-up and support that matches their child's activity type. By operating the "Select Photo" icon 513, parents can display a desired photo from among photos of their child recorded on the terminal device 100 or photos of their child uploaded from another smartphone or PC.
[0097] This week's log is the same as that shown in Figure 15, so a detailed explanation is omitted. In this way, parents can properly grasp their child's activity type, which can be useful for daily communication with their child and for daily life (such as purchasing activities).
[0098] 18 is a diagram showing an example of a processing procedure for the favorite color analysis function. The control unit 401 collects image data (S151) and determines whether the data has been collected for a second predetermined period (S152). The second predetermined period can be, for example, one week, but is not limited to this. If the data has not been collected for the second predetermined period (NO in S152), the control unit 401 continues the processing of step S151.
[0099] If the data has been collected over the second predetermined period (YES in S152), the control unit 401 classifies the colors contained in the imaged object (S153). Specifically, the analysis unit 409 has a learning model with a color analysis function, and can detect the colors of areas within the image and the colors used within the image. The analysis unit 409 can classify the colors by frequently occurring colors.
[0100] The control unit 401 calculates the total number of objects captured for each color (S154). That is, the total number of objects detected for each classified color can be calculated. For example, if a child captures 20 images of objects classified as yellow over the course of a week, the number of "yellow" objects is set to 20. The control unit 401 registers the color captured the most as the "favorite color" (S155).
[0101] The control unit 401 compares the number of images captured for each color with the number of images captured for each color with the most recent second predetermined period (e.g., last week), and registers the color for which the number of images captured for each color with the current second predetermined period (e.g., this week) is on the rise as a "rapidly rising" color (S156). For example, the difference between the number of images captured for each color with the previous week and this week can be calculated, and the color for which the calculated difference is equal to or greater than a predetermined difference threshold can be registered as a "rapidly rising" color. Alternatively, the color for which the calculated difference is the largest can be registered as a "rapidly rising" color.
[0102] The control unit 401 provides the analysis results ("favorite color", "rapidly rising") to the terminal device 100 (S157), and ends the process. Note that the words "favorite color" and "rapidly rising" are merely examples, and the present invention is not limited to these words.
[0103] As described above, the analysis unit 409 functions as a second analysis unit and analyzes the colors contained in the image capture target based on the collected image data. The analysis unit 409 may analyze the user's color interests based on the analysis results collected each time a second predetermined period is reached.
[0104] Fig. 19 is a diagram showing an example of the analysis results of the favorite color analysis function. The "OO-chan's Color Preference Analysis Results" screen 531 shown in Fig. 19 can be displayed on the display unit 104 of the terminal device 100. The "OO-chan's Color Preference Analysis Results" screen 531 has, for example, a message display area 502 that displays an error message, a display area 532 that displays an image of a favorite color, a display area 533 that displays a favorite color, and a display area 534 that displays a "trending" color.
[0105] Display area 532 displays images 1 to 4, each with a phrase such as "OO-chan seems to especially like 'yellow'!" drawn in her favorite color (yellow). Images 1 to 4 may be switched to a different image from the life log every time a certain period of time has passed. The ranking of favorite colors (in the example of FIG. 19, first place is "yellow," second place is "red," and third place is "blue") is also displayed.
[0106] The number of objects detected for each color is displayed in display area 533 in order of favorite color. In the example of Fig. 19, the following numbers are displayed: "yellow" 20, "red" 18, "blue" 15, "orange" 10, and "green" 7. "Green" is a color that has been photographed more frequently this week compared to last week's analysis results, and is labeled "Rapidly Rising!"
[0107] Display area 534 displays a message such as "It seems he has recently become interested in 'green,'" along with images 1 and 2 drawn in green. This allows parents to easily grasp in real time which colors their child has become interested in and which colors they like. Parents can also appropriately grasp changes in colors of interest, which can be useful in daily communication with their children and in daily life (such as purchasing activities). In the above example, the interest types shown in FIG. 15, the activity types shown in FIG. 17, and the preferred color analysis shown in FIG. 19 were explained using separate diagrams for convenience, but these are merely examples, and the interest types, activity types, and preferred color analysis can be displayed simultaneously on the same screen.
[0108] (Fifth embodiment) In the fourth embodiment, the analysis function is provided in the server 400, but the present invention is not limited to this. In the fifth embodiment, a configuration in which the analysis function is provided in a terminal device will be described.
[0109] 20 is a diagram showing an example of the configuration of an information processing system according to the fifth embodiment. The information processing system includes an imaging device 50 and a terminal device 150 as an information processing device. The imaging device 50 is the same as in the fourth embodiment. The terminal device 150 differs from the fourth embodiment in that it includes an analysis unit 155, a character string analysis unit 156, a graphic analysis unit 157, an order estimation unit 158, a voice synthesis unit 159, a sound reproduction unit 160, an analysis unit 161, a voice data list 163, and a computer program 164. The analysis unit 155, the character string analysis unit 156, the graphic analysis unit 157, the order estimation unit 158, the voice synthesis unit 159, the sound reproduction unit 160, the analysis unit 161, and the voice data list 163 are similar to the analysis unit 403, the character string analysis unit 404, the graphic analysis unit 405, the order estimation unit 406, the voice synthesis unit 407, the sound reproduction unit 408, the analysis unit 409, and the voice data list 411 provided in the server 400 of the fourth embodiment. The computer program 164, when executed by the control unit 151, can realize all or part of the functions of the analysis unit 155, the character string analysis unit 156, the graphic analysis unit 157, the order estimation unit 158, the voice synthesis unit 159, the sound reproduction unit 160, and the analysis unit 161.
[0110] 21 is a diagram showing an example of a processing procedure of the information processing system of the fifth embodiment. The imaging device 50 captures an image of an imaging target (S161) and transmits image data obtained by capturing the image to the terminal device 150 (S162). The terminal device 150 receives the image data (S163) and analyzes the imaging target based on the received image data (S164). The terminal device 150 generates a sound related to the content of the figure or a voice that reads out the character string, depending on whether the imaging target is a figure or a character string (S165).
[0111] The terminal device 150 transmits the generated sound or voice to the imaging device 50 (S166). The imaging device 50 receives the sound or voice (S167) and outputs the received sound or voice (S168). The terminal device 150 records the received image data and analysis results in the storage unit 162 (S169) and updates the number of times the sound is played or the voice is read aloud (S170). Every time a user such as an infant or child takes an image of an imaging target with the imaging device 50, the process shown in FIG. 21 is repeated, and information such as the image data, analysis results, and number of times the sound is played or the voice is read aloud can be collected. Note that the same processes as those in FIGS. 12 to 19 are performed in the fifth embodiment, and therefore a description thereof will be omitted.
[0112] The control unit 151 functions as a collection unit and collects image data obtained by capturing an image of an imaging target via the communication unit 152. The analysis unit 161 analyzes the interests of the user (infant or child) who captured the image of the imaging target based on the collected image data. The control unit 151 functions as a provision unit and can provide the analysis results of the analysis unit 161.
[0113] In addition, the computer program 164 running on the terminal device 150 causes the computer to execute processing to collect image data obtained by capturing an image of the object to be captured, analyze the interests of the user who captured the image of the object based on the collected image data, and provide the analysis results.
[0114] In this way, by collecting image data of subjects captured by infants and children that interest them as part of a life log, analyzing the infants' and children's daily interests based on the collected life log, and providing feedback on the analysis results to parents, parents can follow up and support their infants and children in accordance with their interests. For example, parents can buy their children goods related to areas that interest them, or take them to places or facilities that interest them.
[0115] (Sixth embodiment) In the sixth embodiment, a configuration in which an analysis function is provided in an imaging device will be described.
[0116] 22 is a diagram illustrating an example of the configuration of an information processing system according to the sixth embodiment. The information processing system includes an imaging device 50 and a terminal device 100. The terminal device 100 is the same as that of the fourth embodiment. The imaging device 50 differs from that of the fourth embodiment in that it includes an analysis unit 71, a character string analysis unit 72, a graphic analysis unit 73, an order estimation unit 65, a voice synthesis unit 66, a sound reproduction unit 67, an analysis unit 68, and a voice data list 70. The analysis unit 71, the character string analysis unit 72, the graphic analysis unit 73, the order estimation unit 65, the voice synthesis unit 66, the sound reproduction unit 67, the analysis unit 68, and the voice data list 70 are the same as the analysis unit 403, the character string analysis unit 404, the graphic analysis unit 405, the order estimation unit 406, the voice synthesis unit 407, the sound reproduction unit 408, the analysis unit 409, and the voice data list 411, respectively, included in the server 400 according to the fourth embodiment.
[0117] 23 is a diagram showing an example of a processing procedure of the imaging device 50 of the sixth embodiment. The imaging device 50 captures an image of an object (S181) and analyzes the object based on the image data obtained by capturing the image (S182). The imaging device 50 generates a sound related to the content of the image or a voice that reads out the character string, depending on whether the object is a figure or a character string (S183).
[0118] The imaging device 50 outputs the generated sound or voice (S184). The imaging device 50 records the image data and the analysis result in the storage unit 69 (S185), and updates the number of times the sound is played or the voice is read aloud (S186). Each time a user such as an infant or a child takes an image of an object with the imaging device 50, the process shown in FIG. 23 is repeated, and information such as the image data, the analysis result, and the number of times the sound is played or the voice is read aloud can be collected. Note that the same processes as those in FIGS. 12 to 19 are performed in the sixth embodiment, and therefore a description thereof will be omitted.
[0119] The imaging device 50 includes a gripping portion 62, a window portion 61 for looking into the imaging target, an imaging unit 51 capable of imaging the imaging target while the imaging target is being looked into through the window portion 61, an analysis unit 68 that analyzes the interests of the user who captured the imaging target based on image data captured and collected by the imaging unit 51, and a control unit 11 that serves as a providing unit that provides the analysis results of the analysis unit 68.
[0120] In this way, by collecting image data of subjects captured by infants and children that interest them as part of a life log, analyzing the infants' and children's daily interests based on the collected life log, and providing feedback on the analysis results to parents, parents can follow up and support their infants and children in accordance with their interests. For example, parents can buy their children goods related to areas that interest them, or take them to places or facilities that interest them.
[0121] In the fourth to sixth embodiments, the analysis function is provided in either the server 400, the terminal device 150, or the imaging device 50, but the analysis function may be distributed among at least two of the server 400, the terminal device 150, and the imaging device 50.
[0122] FIG. 24 is a diagram showing an example of the transition of interest analysis results. By collecting interest analysis results of users such as infants and children as they grow, it is possible to understand the transition of each individual user's interest analysis results. In the example of FIG. 24, the transition of "My Craze" categories (e.g., categories A1 to A5) and "Rapidly Trending" categories (e.g., categories B1 to B5) are displayed along with the user's age. Categories A1 to A5 and B1 to B5 may include shared categories. By collecting data such as that shown in FIG. 24 from a large number of users, it can be utilized as big data. It is also possible to predict how the interests of individual users will change and transition in the future.
[0123] Figure 25 shows an example of the results of an interest analysis by age group, region, and time series. By collecting the analysis results of a large number of children, differences in interest trends by age group, region, and time series can be understood. This allows for understanding time series trends and future predictions for any group (user group), such as by age group, region, or nursery school. Figure 25A shows the results of an analysis of interest fields by age group A1, A2, A3, etc. For convenience, the interest fields are labeled C1 to C5, but the number of fields is not limited to this. The interest fields for each age group may be represented in a diagram similar to a radar chart. Figure 25B shows the results of an analysis of interest fields by region L1, L2, L3, etc. For convenience, the interest fields are labeled C1 to C5, but the number of fields is not limited to this. The interest fields for each region may be represented in a diagram similar to a radar chart. Figure 25C shows the results of an analysis of interest fields by time elapsed since a predetermined reference point (e.g., a specific age). For convenience, the fields of interest are designated as C1 to C5, but the number of fields is not limited to this. The fields of interest at each point in time may be represented in a diagram such as a radar chart. As described above, the analysis unit 409 may analyze the temporal transition of the interests of a group of users divided by at least one of age and region based on the collected image data (for example, including the time series transition for each group or future predictions).
[0124] The voice reading system of this embodiment comprises an imaging device and a main unit having a mounting surface on which an object to be imaged can be placed, the imaging device comprising a holding portion, a window portion for looking into the object to be imaged, an imaging unit capable of imaging the object to be imaged while looking into the object through the window portion, and a transmitting portion for transmitting image data obtained by imaging with the imaging unit to the main unit, the main unit comprising a receiving portion for receiving the image data, an analyzing portion for analyzing the image data received by the receiving portion, and an output portion for outputting voice based on the analysis results of the analyzing portion.
[0125] In the voice reading system of this embodiment, the placement surface is rectangular in plan view and is inclined so that the height decreases from each of a pair of edge portions that are separated by the center of the placement surface toward the center.
[0126] In the text-to-speech system of this embodiment, the main unit includes restricting sections that restrict movement of the object on another pair of edge sections of the placement surface along the inclined direction of the placement surface.
[0127] In the text-to-speech system of this embodiment, the main body device has a housing portion formed on the placement surface for housing the imaging device.
[0128] In the voice reading system of this embodiment, the imaging device includes a detection unit that detects the distance to the object, and a correction unit that corrects the imaging range of the imaging unit so that the object within the field of view through the window can be imaged in accordance with the distance detected by the detection unit.
[0129] In the voice reading system of this embodiment, the analysis unit analyzes whether the imaged object is a string of characters or a figure based on the image data received by the receiving unit, and if the analysis unit analyzes that the imaged object is a string of characters, the output unit outputs a voice reading out the string of characters.
[0130] In the text-to-speech system of this embodiment, when the analysis unit analyzes that the imaged object is a diagram, the output unit outputs a sound related to the content of the diagram.
[0131] The text-to-speech system of this embodiment includes an emotional index calculation unit that, when the analysis unit analyzes that the imaged object is a character string, performs a semantic analysis of the character string to calculate an emotional index, and the output unit outputs background music according to the emotional index calculated by the emotional index calculation unit.
[0132] The text-to-speech system of this embodiment includes a reading order estimation unit that, when the analysis unit analyzes that the imaged object is a character string, estimates the reading order of the character string based on the arrangement of the character string.
[0133] The voice reading system of this embodiment includes a voice synthesis unit that synthesizes voice using voice features of pre-recorded voice, and when the analysis unit analyzes that the imaged object is a character string, the output unit outputs the voice synthesized by the voice synthesis unit based on the character string.
[0134] The voice reading system of this embodiment includes a voice synthesis unit that synthesizes voice by concatenating pre-recorded voice fragments, and when the analysis unit analyzes that the imaged object is a character string, the output unit outputs the voice synthesized by the voice synthesis unit based on the character string.
[0135] The voice reading device of this embodiment comprises a holding unit, a window unit for looking into the imaged object, an imaging unit capable of imaging the imaged object while the imaged object is being looked into through the window unit, an analysis unit for analyzing image data obtained by imaging with the imaging unit, and an output unit for outputting voice based on the analysis results of the analysis unit.
[0136] The voice reading device of this embodiment includes a detection unit that detects the distance to the object on which the image capture target is written, and a correction unit that corrects the imaging range of the imaging unit so that the image capture target within the field of view through the window unit can be captured in accordance with the distance detected by the detection unit.
[0137] The information processing device of this embodiment includes a collection unit that collects image data obtained by capturing an image of an object to be captured, an analysis unit that analyzes the interests of the user who captured the image of the object to be captured based on the image data collected by the collection unit, and a provision unit that provides the analysis results of the analysis unit.
[0138] The information processing device of this embodiment includes a first analysis unit that analyzes the field of the imaging subject based on the image data collected by the collection unit, and the analysis unit analyzes the user's interests based on the analysis results of the first analysis unit collected each time a first predetermined period of time is reached.
[0139] In the information processing device of this embodiment, the analysis unit analyzes a first index related to the user's interests based on the number of imaging subjects captured for each field analyzed by the first analysis unit.
[0140] The information processing device of this embodiment includes an identification unit that identifies changes in the first indicator analyzed by the analysis unit for each first predetermined period, and the analysis unit analyzes the user's type of interests in accordance with the changes in the first indicator identified by the identification unit.
[0141] In the information processing device of this embodiment, the analysis unit analyzes a second index related to the user's interests based on changes in the number of subjects captured for each field analyzed by the first analysis unit for each first predetermined period.
[0142] In the information processing device of this embodiment, the analysis unit analyzes a third index related to the user's interests based on the field with the largest number of imaging objects among the number of imaging objects imaged for each field analyzed by the first analysis unit.
[0143] The information processing device of this embodiment is equipped with a second analysis unit that analyzes the colors contained in the object to be imaged based on the image data collected by the collection unit, and the analysis unit analyzes the user's interests in colors based on the analysis results of the second analysis unit collected each time a second predetermined period is reached.
[0144] The information processing device of this embodiment includes a third analysis unit that analyzes whether the imaging object is a character string or a figure based on the image data collected by the collection unit, a reading unit that reads out the character string if the third analysis unit analyzes that the imaging object is a character string, and a playback unit that plays sound related to the content of the figure if the third analysis unit analyzes that the imaging object is a figure, and the analysis unit analyzes the activity type of the user based on the number of times the reading unit reads out the character string and the number of times the figure is played back by the playback unit over a third predetermined period.
[0145] The information processing device of this embodiment includes an output unit that outputs a sound based on the analysis result of the third analysis unit.
[0146] In the information processing device of this embodiment, the analysis unit analyzes the temporal changes in the interests of a group of users divided by at least one of age and region based on the image data collected by the collection unit.
[0147] The imaging device of this embodiment includes a gripping unit, a window unit for looking into the imaging target, an imaging unit capable of imaging the imaging target while the imaging target is being looked into through the window unit, an analysis unit that analyzes the interests of a user who has imaged the imaging target based on image data captured and collected by the imaging unit, and a provision unit that provides the analysis results of the analysis unit.
[0148] The computer program of this embodiment causes a computer to execute a process of collecting image data obtained by capturing an image of an object to be captured, analyzing the interests of the user who captured the image of the object based on the collected image data, and providing the analysis results. [Explanation of symbols]
[0149] 1 Home network 2. Communication Network 10 Main unit 11, 201 Control section 12 Communications Department 13, 205 Analysis Department 131, 206 String analysis section 132, 207 Figure analysis section 14, 208 Speech synthesis unit 15, 209 Order estimation part 16, 210 storage section 161, 211 BGM List 162, 212 Audio Data List 17, 213 Mike 18, 214 Speaker 19, 215 Emotion index calculation section 21 Placement surface 22 Marginal area 23 Central part 24 Regulatory Department 25 Storage section 26 Indicator light 30 Main unit 31 First Communications Department 32 Second Communications Department 33 Storage section 331 BGM List 34. Mike 35 speakers 36 Control Unit 50 Imaging device 51, 202 Imaging unit 52, 203 Distance sensor 53, 204 Correction section 54 memory 55 processors 56 Communications Department 57 Speaker 69 Memory section 70 Audio Data List 100, 150 terminal equipment 101, 151 Control unit 102 First Communications Department 103 Second Communications Department 152 Communications Department 104, 153 Display section 105, 154 Operation section 106, 162 Storage section 163 Audio Data List 164 Computer Programs 300 servers 301 Control Unit 302 Communications Department 303 Analysis Department 304 String analysis section 305 Diagram Analysis Department 306 Speech synthesis unit 307 Order estimation part 308 Storage section 309 Audio Data List 310 Emotion index calculation unit 400 servers 401 Control Unit 402 Communications Department 403, 155, 71 Analysis Department 404, 156, 72 String analysis section 405, 157, 73 Diagram analysis section 406, 158, 65 Order estimation part 407, 159, 66 Speech synthesis unit 408, 160, 67 Sound playback section 409, 161, 68 Analysis Department 410 Storage section 411 Audio Data List
Claims
1. The image capturing device includes an imaging device and a main body device having a mounting surface on which an object having an image capturing target written thereon can be placed, The imaging device is A gripping portion; a window for looking into the imaging target; an imaging unit capable of imaging the imaging target in a state where the imaging target is viewed through the window; a transmitting unit that transmits image data obtained by imaging with the imaging unit to the main device; Equipped with The main body device includes: a receiving unit that receives the image data; an analysis unit that analyzes the image data received by the receiving unit; an output unit that outputs a voice based on the analysis result of the analysis unit; an emotion index calculation unit that calculates an emotion index by performing a semantic analysis on the character string when the analysis unit analyzes the image data to determine that the image capture target is a character string; Equipped with The output unit outputting background music according to the emotion index calculated by the emotion index calculation unit; Voice reading system.
2. The placement surface is It has a rectangular shape in plan view, The mounting surface is inclined so that the height decreases from each of a pair of edge portions sandwiching the center portion of the mounting surface toward the center portion. The text-to-speech system according to claim 1 .
3. The main body device includes: and a restricting portion for restricting movement of the object on another pair of edges of the placement surface along the inclined direction of the placement surface. The text-to-speech system according to claim 1 or 2.
4. The main body device includes: a storage section for storing the imaging device is formed on the placement surface; The text-to-speech system according to any one of claims 1 to 3.
5. The imaging device is a detection unit that detects the distance to the object; a correction unit that corrects an imaging range of the imaging unit so that an imaging target within a field of view through the window unit can be imaged in accordance with the distance detected by the detection unit; Equipped with The text-to-speech system according to any one of claims 1 to 4.
6. The analysis unit Analyzing whether the imaged object is a character string or a picture based on the image data received by the receiving unit; The output unit When the analysis unit analyzes that the imaged object is a character string, a voice reading the character string is output. The text-to-speech system according to any one of claims 1 to 5.
7. The output unit When the analysis unit analyzes that the imaging target is a figure, it outputs a sound related to the content of the figure. The text-to-speech system according to claim 6.
8. a reading order estimation unit that, when the analysis unit analyzes that the image capture target is a character string, estimates a reading order of the character string based on an arrangement of the character string; The text-to-speech system according to any one of claims 6 to 7.
9. a speech synthesis unit that synthesizes speech using speech features of pre-recorded speech; The output unit When the analysis unit analyzes that the imaging target is a character string, the voice synthesis unit outputs a synthesized voice based on the character string. The text-to-speech system according to any one of claims 6 to 8.
10. a speech synthesis unit that synthesizes speech by connecting pre-recorded speech fragments; The output unit When the analysis unit analyzes that the imaging target is a character string, the voice synthesis unit outputs a synthesized voice based on the character string. The text-to-speech system according to any one of claims 6 to 8.
11. A gripping portion; a window for looking into the imaging target; an imaging unit capable of imaging the imaging target in a state where the imaging target is viewed through the window; an analysis unit that analyzes image data captured by the imaging unit; an output unit that outputs a voice based on the analysis result of the analysis unit; an emotion index calculation unit that calculates an emotion index by performing a semantic analysis on the character string when the analysis unit analyzes the image data to determine that the image capture target is a character string; Equipped with The output unit outputting background music according to the emotion index calculated by the emotion index calculation unit; Voice reading device.
12. a detection unit that detects the distance to the object on which the image is to be captured; a correction unit that corrects an imaging range of the imaging unit so that an imaging target within a field of view through the window unit can be imaged in accordance with the distance detected by the detection unit; Equipped with The text-to-speech device according to claim 11.
Citation Information
Patent Citations
Picture input device
JP1994141141A
Information providing system and information processing system
JP2005025645A
Divination system
JP2007061367A
Contactless scanner
JP2009141556A
Imaging device, dictionary database system, and main subject image output method and main subject image output program
JP2009260800A