Information processing apparatus
The information processing device addresses security and usability by using location and time inputs to manage image and data output restrictions, effectively protecting sensitive information.
Patent Information
- Application Number
- JP2025148456
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-12-09
AI Technical Summary
Existing information processing devices lack consideration for security levels and user-friendliness in handling face image privacy, failing to adequately protect sensitive information based on context.
An information processing device that incorporates input devices for location and time information, along with a judgment device to determine output restrictions on images and data based on these factors, and processing devices to modify or mask sensitive content as needed.
Enhances security and usability by dynamically adjusting output based on location and time, ensuring privacy protection and user-friendly operation.
Smart Images

Figure 2025179196000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device. [Background technology]
[0002] In an information society, consideration of privacy in images captured by a camera is important. Recently, in consideration of privacy, a device has been proposed that detects a face image from an image captured by a camera, and if the detected face image matches a face image of a specific person registered in advance, does not perform mask processing (abstraction processing) on the face image, but if they do not match, performs mask processing on the face image (see Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-62560 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the above-mentioned Patent Document 1 lacks careful consideration according to security levels, and is not necessarily user-friendly.
[0005] SUMMARY OF THE INVENTION The present invention has been made in view of the above-mentioned problems, and has as its object to provide an information processing device that takes security into consideration and is easy to use. [Means for solving the problem]
[0006] The information processing device of the present invention comprises a first input device (52) for inputting an image, a second input device (52) for inputting at least one of location information and time information, and a judgment device (70) for determining whether to impose restrictions on the output of the image based on the information input by the second input device when an instruction to output the image is received.
[0007] In this case, the second input device can input information about the location where the image was taken or the location where the image is to be output, as well as information about the time when the image was taken or the time when the image is to be output.
[0008] The determination device may determine whether or not output restriction of the image is necessary based on the information and on whether security or privacy is ensured.
[0009] The system may also include a first processing device (70) that imposes a restriction on at least a portion of the image when the determination device determines that a restriction should be imposed on the output of the image. In this case, the first processing device may impose a restriction on at least a portion of the image when the location information indicates that the location has changed from a secure location to an unsecured location. The first processing device may also identify a portion of the image where an output restriction should be imposed based on at least one of the location information and the time information. The system may also include a third input device (53) that inputs information about a subject in the image, and the first processing device may identify a portion of the image where an output restriction should be imposed based on the information input by the third input device. In this case, the first processing device may output the information about the subject input by the third input device together with the image. Furthermore, the first processing device may convert the information about the subject input by the third input device based on a predetermined rule and output the converted information together with the image.
[0010] In the information processing device of the present invention, the determination device can determine whether or not to impose a restriction on output of the image based on information about the facial expression of the person in the image.
[0011] In addition, the information processing device of the present invention may include a fourth input device (52) for inputting at least one of voice data and text data, and the judgment device may determine whether to impose restrictions on the output of at least one of the voice data and the text data based on at least one of the location information and the time information.
[0012] In this case, the device may include a second processing device (70) that outputs text data converted and generated from the voice data when the determination device determines that it is necessary to limit the output of the voice data. Also, the device may include a second processing device that converts a specific noun in at least one of the voice data and the text data into another noun and outputs the converted noun when the determination device determines that it is necessary to limit the output of the voice data.
[0013] The information processing device of the present invention may have a first housing (11) having a display unit (14) that displays an image processed by the first processing device, and a second housing (51) having a determination unit, and the first housing and the second housing may be separate. In this case, the first housing may have a position detection unit (22) that detects position information of the first housing, and the second input device may input the detection result of the position detection unit.
[0014] In order to clearly explain the present invention, the above description has been made in association with the reference numerals in the drawings representing one embodiment, but the present invention is not limited to this, and the configuration of the embodiment described below may be appropriately improved, or at least a part of it may be replaced with other components. Furthermore, components that are not particularly limited in terms of their location may be located in any position that can achieve their function, not limited to the location disclosed in the embodiment. [Effects of the Invention]
[0015] The present invention has an effect of providing an information processing device that takes security into consideration and has improved usability. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a functional block diagram of a personal assistant system 100 according to an embodiment. [Figure 2] 2(a) to 2(d) are flowcharts showing the process of recording the voice input from the voice input unit. [Figure 3] FIG. 10 is a diagram showing a voiceprint DB. [Figure 4] 10 is a flowchart relating to processing of audio data. [Figure 5] FIG. 10 is a diagram showing a stored data DB. [Figure 6] 5 is a flowchart showing a specific process of step S76 in FIG. 4. [Figure 7] FIG. 10 is a diagram showing a keyword DB. [Figure 8] 5 is a flowchart showing a specific process of step S84 in FIG. 4. [Figure 9] 9 is a flowchart showing a specific process of step S122 in FIG. 8. [Figure 10] 10 is a flowchart showing specific processes of steps S142 and S148 in FIG. 9. [Figure 11] FIG. 10 is a diagram showing a specific conversion word DB. [Figure 12] 11 is a flowchart showing specific processes of steps S164, S166, S178, and S180 in FIG. [Figure 13] FIG. 10 is a diagram showing a place name DB. [Figure 14] FIG. 10 is a diagram illustrating a keyword storage DB. [Figure 15] 15(a) to 15(d) are diagrams showing examples of weight tables. [Figure 16] FIG. 10 is a diagram showing a keyword record DB. [Figure 17] FIG. 10 is a diagram illustrating a command DB. [Figure 18] FIG. 18(a) is a diagram showing an example of a displayed task list, and FIG. 18(b) is a diagram showing an example of a displayed recorded voice list. [Figure 19]10 is a flowchart showing a first process performed simultaneously in parallel in step S96. [Figure 20] 10 is a flowchart showing a second process performed simultaneously in parallel in step S96. [Figure 21] 10 is a flowchart showing a third process performed simultaneously in parallel in step S96. [Figure 22] FIG. 10 is a diagram illustrating a secure range DB. [Figure 23] FIG. 10 is a diagram showing an ambiguous word database. [Figure 24] 10 is a flowchart relating to the input and storage process of image data. [Figure 25] 10 is a flowchart relating to a display process of image data. [Figure 26] 10 is a flowchart showing a data erasure process on the portable terminal side. [Figure 27] 10 is a flowchart showing a process of deleting audio data on the server side; DETAILED DESCRIPTION OF THE INVENTION
[0017] An embodiment of the personal assistant system 100 will be described in detail below with reference to Figs. 1 to 27. Fig. 1 shows a block diagram of the personal assistant system 100. As shown in Fig. 1, the personal assistant system 100 includes a portable terminal 10 and a server 50.
[0018] The portable terminal 10 is a terminal that can be carried by a user, such as a mobile phone, a smartphone, a PHS (Personal Handy-phone System), or a PDA (Personal Digital Assistant). The size of the portable terminal 10 is such that it can fit in a breast pocket, for example. As shown in FIG. 1 , the portable terminal 10 has an input unit 12, a display unit 14, a playback unit 16, a warning unit 18, a biometric information input unit 20, a position detection unit 22, a time detection unit 24, a calendar unit 26, a flash memory 28, a communication unit 32, and a terminal-side control unit 30. The portable terminal 10 also has a portable first housing 11 that houses at least a portion of these units.
[0019] The input unit 12 has an audio input unit 42, an image input unit 43, and a text data input unit 44. The audio input unit 42 includes a microphone, and acquires the user's voice and voices emitted around the user, and outputs the voices to the audio encoding unit 45.
[0020] The audio encoding unit 45 encodes the audio input from the audio input unit 42 using an encoding method determined by the initial settings of the terminal-side control unit 30 , and outputs the encoded audio to the communication unit 32 .
[0021] The image input unit 43 includes an optical system, an imaging element, etc., and generates a moving image and outputs it to the image encoding unit 46. The image encoding unit 46 encodes the image input from the image input unit 43 using an encoding method determined by the initial setting of the terminal-side control unit 30, and outputs it to the communication unit 32. Note that if the input to the image input unit 43 is a moving image and there is also an input to the audio input unit 42 at the same time as the input of the moving image, the input data will be associated (performing a process similar to multiplexing). Note that the image input unit 43 is not limited to outputting moving image data, and may also output still image data to the image encoding unit 46.
[0022] The text data input unit 44 includes an input interface such as a keyboard or a touch panel, and acquires text data in response to user input operations. The input unit 12 also has a function of accepting user operation instructions from the touch panel or the like. The various data input through the input unit 12 may be stored in the flash memory 28 as needed. In this case, the various data stored in the flash memory 28 is deleted under the control of the server-side control unit 70, which will be described later.
[0023] The display unit 14 includes a display such as a liquid crystal display, an organic EL display, etc. Under the control of the terminal-side control unit 30 or the server-side control unit 70, the display unit 14 displays data such as image data and text data on the display, and displays menus for the user to operate.
[0024] The playback unit 16 includes a speaker and outputs voice and sound. For example, when the terminal-side control unit 30 or the server-side control unit 70 causes the display unit 14 to play back a video, the playback unit 16 can play back the audio of the video input by the audio input unit 42. Note that instead of playing back the audio, or in addition to playing back the audio, the server 50 may convert the audio into text data and display the converted text data as a caption on the display unit 14. In particular, in places where confidentiality and privacy cannot be ensured, it is desirable to play back text data instead of audio data.
[0025] The warning unit 18 warns the user when an error occurs in the mobile terminal 10, for example, by outputting a warning sound via the playback unit 16 or displaying a warning on the display unit 14.
[0026] The biometric information input unit 20 is a device that acquires at least one piece of biometric information, such as the user's muscle state (degree of tension and relaxation), blood pressure, heart rate, pulse, and body temperature, and inputs it to the terminal-side control unit 30. A wristwatch-type device, such as that described in Japanese Patent Application Laid-Open No. 2005-270543 (U.S. Patent No. 7,538,890), can be used to detect the biometric information. Blood pressure and pulse may be detected by an infrared pulse wave detection sensor, and heart rate may be detected by a vibration sensor. A higher heart rate than normal indicates a state of tension, and a lower heart rate indicates a state of relaxation. Since pupils dilate in a state of tension and constrict in a state of relaxation, a configuration may be applied in which pupils are detected to determine whether the user is in a state of tension or relaxation.
[0027] The position detection unit 22 detects the position (absolute position) of the user, and here, for example, a GPS (Global Positioning System) is adopted. Note that the position detection unit 22 may also adopt an absolute position measurement system using RFID (Radio Frequency Identification) or the like.
[0028] The time detection unit 24 has a timekeeping function that detects the current time. The calendar unit 26 stores dates and days of the week in association with each other. The flash memory 28 is a memory that temporarily stores data. The communication unit 32 has a wireless LAN unit for accessing an access point via Wi-Fi communication, a wired connection unit via an Ethernet (registered trademark) cable, or a USB connection unit for communicating with external devices such as computers. In this embodiment, the communication unit 32 can communicate with the communication unit 52 of the server 50.
[0029] The terminal-side control unit 30 comprehensively controls each component of the portable terminal 10 and executes processing on the side of the portable terminal 10. For example, the terminal-side control unit 30 acquires the time when voice data, image data, text data, etc. are input to the voice input unit 42, image input unit 43, or text data input unit 44 via the time detection unit 24, and acquires the position of the portable terminal 10 when the voice data, image data, or text data is input via the position detection unit 22. When transmitting the voice data, image data, or text data to the server 50, the terminal-side control unit 30 transmits information on the time and position as metadata together with the data.
[0030] The server 50 is installed, for example, in a company where a user of the mobile terminal 10 works. However, the location is not limited to this, and the server 50 may also be installed in a system management company. As shown in FIG. 1 , the server 50 includes a communication unit 52, a face recognition unit 53, a text data generation unit 54, a voiceprint analysis unit 55, a weighting unit 56, an extraction unit 58, a classification unit 60, a conversion unit 62, a flash memory 64, a hard disk 66, and a server-side control unit 70. The server 50 also includes a second housing 51 that houses at least a portion of these units. The second housing 51 and the first housing of the mobile terminal 10 described above are mechanically (physically) separated from each other.
[0031] The communication unit 52 is similar to the communication unit 32 on the mobile terminal 10 side, and in this embodiment, is capable of communicating with the communication unit 32 on the mobile terminal 10 side. Data (audio data, image data, and text data) received by the communication unit 52 is stored in the flash memory 64 via the server-side control unit 70. That is, the communication unit 52 functions as an input unit in the server 50 that inputs the audio data, text data, and image data to the server-side control unit 70.
[0032] The face recognition unit 53 detects the facial area of a subject (person) from an image transmitted from the mobile terminal 10 and stored in the flash memory 64, and compares the face in the detected facial area with a facial image registered on the hard disk 66 to identify the face included in the image transmitted from the mobile terminal 10. Various methods can be used for face recognition. In this embodiment, face recognition is performed using a face detection method based on edge detection or shape pattern detection, and / or a face detection method based on hue extraction or skin color extraction. The face recognition unit 53, in cooperation with the server-side control unit 70, can also perform noise reduction and edge enhancement on the image data transmitted from the mobile terminal 10 to facilitate face recognition. The face recognition unit 53 notifies the server-side control unit 70 of information (such as name and job title) of the subject (person) recognized from the face.
[0033] When a face is registered on the hard disk 66, the confidentiality level (high, medium, low) of the face may be set. For example, if a face is registered as having high confidentiality, the server-side control unit 70 may prohibit the transfer of the image of the face to the mobile terminal 10, or may reduce the resolution of the image or apply a mosaic when displaying it on the display unit 14. The server-side control unit 70 may also display information about the person notified by the face recognition unit 53 near the face image that has been processed by reducing the resolution or applying a mosaic. In this way, even if the face cannot be recognized from the image alone due to mosaic processing, the person's information can be displayed to identify the person in the mosaic-processed image. In this case, the server-side control unit 70 may display information about the person (e.g., name, job title, etc.) by converting it to initials.
[0034] Furthermore, when an image of a face that is not registered in the hard disk 66 is input, the server-side control unit 70 can prohibit the image of the unregistered face from being transferred to the mobile terminal 10, or can reduce the resolution of the image or apply a mosaic when displaying it on the display unit 14. In this case, the privacy of the unregistered person can be protected.
[0035] It should be noted that a character recognition unit may be provided instead of or in addition to the face recognition unit 53. In this case, the character recognition unit will notify the server-side control unit 70 of the results of recognizing characters present in the image (for example, characters written on a whiteboard or blackboard, or characters projected on a screen). If the server-side recognition unit 70 determines that the content of the characters is highly confidential, it may reduce the resolution of the character portion of the image or apply a mosaic to the character. The server-side control unit 70 may convert the recognition result by the character recognition unit according to a predetermined rule and display the converted content near the character portion of the image.
[0036] The text data generating unit 54 acquires the voice data stored in the flash memory 64 and converts the voice data to generate text data. The generated text data is stored in the flash memory 64 via the server-side control unit 70.
[0037] The voiceprint analysis unit 55 performs voiceprint analysis by pattern matching the volume (intensity), frequency, and length of the voice with registered voiceprint data to identify the person who spoke the voice. Note that voiceprint analysis does not necessarily use all of the volume (intensity), frequency, and length of the voice; it may identify the person who spoke the voice by using at least the frequency of the voice. Here, when the voices of multiple people are input to the voice input unit 42, the accuracy of the voiceprint analysis by the voiceprint analysis unit 55 may be reduced. In this case, if an image of the voice input is input to the image input unit 43, the voiceprint analysis unit 55 may perform voiceprint analysis taking into account the identification result by the face recognition unit 53. For example, the voiceprint analysis unit 55 may analyze the input voice by taking into account the mouth movements of a person who can be recognized by the face recognition unit 53.
[0038] The weighting unit 56 acquires the voice data stored in the flash memory 64, the text data generated from the voice data, or the text data input from the text data input unit 44, and weights each piece of text data. The weighting unit 56 stores the numerical value (task priority) obtained by weighting in the flash memory 64 together with the text data.
[0039] The weighting by the weighting unit 56 is performed based on, for example, the volume and frequency of the voice and the meaning of the text data. Specifically, the weighting unit 56 performs weighting based on the results of analysis by the voiceprint analysis unit 55 based on the volume and frequency of the voice (information about who made the voice), or performs weighting according to confidentiality based on the meaning of the text data. In this embodiment, confidentiality means the degree to which it is preferable that the data not be seen by others (unspecified third parties).
[0040] The weighting unit 56 is connected to a changing unit 72 and a setting unit 74. The changing unit 72 changes the weighting settings of the weighting unit 56, and the setting unit 74 changes the weighting settings of the weighting unit 56 based on a user instruction. The setting unit 74 may change the settings based on a user instruction input from an input unit (keyboard, etc.) of the server, or may change the settings in response to a user instruction input from the input unit 12 of the mobile terminal 10 via the communication units 52, 32.
[0041] The extraction unit 58 extracts predetermined words from the text data stored in the flash memory 64. That is, the extraction unit 58 extracts predetermined words from the information input to the input unit 12 of the mobile terminal 10. The predetermined words are words that should not be seen by others, that is, words that require relatively high confidentiality, and the words are predetermined in a keyword DB (see FIG. 7) stored on the hard disk 66.
[0042] The extraction unit 58 may cooperate with the face recognition unit 53 and the server-side control unit 70 to set a keyword for each user and extract confidentiality information.
[0043] The classification unit 60 classifies the words extracted by the extraction unit 58 into words with a high level of confidentiality (first words) and words with a slightly high (medium) level of confidentiality (second words). This classification is performed based on the keyword DB (see FIG. 7) stored on the hard disk 66.
[0044] The conversion unit 62 converts words with a "high" confidentiality level and words with a "medium" confidentiality level based on a predetermined rule, and stores the converted words in the flash memory 64.
[0045] The flash memory 64 temporarily stores data to be processed within the server 50. An erasure unit 76 is connected to the flash memory 64. The erasure unit 76 erases the voice data and text data stored in the flash memory 64 at a predetermined timing based on an instruction from the server-side control unit 70. The specific timing for erasing the data will be described later. Other volatile memories may be used instead of the flash memory 64.
[0046] Data such as databases used in various processes is stored on the hard disk 66. Note that instead of the hard disk 66, other non-volatile memory may be used.
[0047] The server-side control unit 70 comprehensively controls each unit in the server 50 and executes processing on the server 50 side. Note that the server 50 actually has a CPU (Central Processing Unit), ROM (Read Only Memory), RAM (Random Access Memory), etc., and the CPU executes programs stored (installed) in the ROM, etc., thereby realizing the functions of each unit, such as the face recognition unit 53, text data generation unit 54, weighting unit 56, extraction unit 58, classification unit 60, conversion unit 62, and voiceprint analysis unit 55, described above.
[0048] Furthermore, the server-side control unit 70 performs masking (reducing resolution, applying mosaic) on at least a portion of the image based on the facial recognition results of the subject (person) in the image performed by the face recognition unit 53. For example, if the person in the image is a person who has been registered in advance as an important person within or outside the company (a person who is required to maintain confidentiality), the server-side control unit 70 performs masking on the person's face. Alternatively, if the person in the image is a person who has been registered in advance as a telecommuter, the server-side control unit 70 performs masking on the person's background image (a portion of the image showing the inside of the telecommuter's home). Furthermore, if the face recognition unit 53 can detect the facial expression of the subject, such as whether the subject is smiling (a soft expression) or furrowing their brows (a stiff expression), the server-side control unit 70 can, for example, not perform masking on the image if the subject is smiling, but can perform masking on the subject and / or the subject's background if the subject is furrowing their brows. As described above, various data from the input unit 12 of the mobile terminal 10 includes time and location information as metadata. The server-side control unit 70 may use this metadata to determine whether to perform image masking. For example, even if the image is of the same person, if the time and location indicate that the image was acquired (photographed) inside the company or at the home of a telecommuter, the masking process may be performed. If the image is acquired outside the company, the masking process may not be performed. Also, even if the image is of the same person, if the time and location indicate that the image was viewed outside the company, the masking process may be performed. If the image is acquired inside the company, the masking process may not be performed. The server-side control unit 70 stores these masked images (converted image data) in the flash memory 64 together with the original image (original image data). The above-mentioned frown lines and smile detection may be detected by a facial expression detection unit (not shown). When detecting wrinkles between the eyebrows using a facial expression detection unit (not shown), an image with wrinkles between the eyebrows may be stored in flash memory 64 as a reference image and detected by pattern matching, or the wrinkles may be detected from the distribution of shadows in the area between the left and right eyes.Detection of wrinkles between the eyebrows is also disclosed in U.S. Patent Publication No. 2008-292148. When detecting a smile using a facial expression detection unit (not shown), evaluation points for the shapes of the eyebrows, pupils, and lips can be calculated by referencing smile data stored in flash memory 64. Smile detection is also disclosed in U.S. Patent Publication No. 2005-201594.
[0049] Next, the processing in the personal assistant system 100 of this embodiment will be described in detail with reference to FIGS.
[0050] First, the process (recording process) of storing the voice input from the voice input unit 42 in the flash memory 64 on the server 50 side will be described with reference to Figures 2(a) to 2(d). Of course, the recording process may be performed continuously, but in this embodiment, in order to save power and perform efficient recording, at least two of the processes shown in Figures 2(a) to 2(d) are performed simultaneously in parallel, or only one of the processes is performed.
[0051] If the portable terminal 10 is a pocket-type terminal, the image input process may be started when the user performs an image input operation using a switch (not shown). For example, an image input operation may be performed at the same time as the start of a meeting to capture (input) the situation of the meeting, or when a telecommuter has a meeting with someone in the office, an image input operation may be performed to capture (input) a picture of the person's face or documents. In this case, the terminal-side control unit 30 controls the operation of the image input unit 43 based on the image input operation.
[0052] Recently, a life log system that records one's own actions has also been proposed. When the mobile terminal 10 is a type that is hung from the user's neck or a type like a head-mounted display, or when the image input unit 43 is provided separately from the mobile terminal 10 (for example, when provided next to glasses), image input may be performed at the same time as sound recording processing.
[0053] (Recording timing example 1) 2(a) is a flowchart showing the process of recording only while a person is speaking. The voice input to the voice input unit 42 is assumed to be input to the server-side control unit 70 via the communication units 32 and 52.
[0054] In the process of FIG. 2(a), in step S10, the server-side control unit 70 determines whether or not voice has been input from the voice input unit 42. If the determination here is affirmative, in step S12, under the instruction of the server-side control unit 70, the voiceprint analysis unit 55 performs a voiceprint analysis of the input voice. This voiceprint analysis identifies the person who uttered the input voice by comparing (pattern matching) voice data contained in a voiceprint DB (see FIG. 3) stored on the hard disk 66 with the input voice data. Here, the voiceprint DB in FIG. 3 associates a person's name with the person's voiceprint data. When this system is used for business purposes, for example, the voiceprint data of all company employees can be registered in the voiceprint DB. When this system is used for private purposes, each user can individually register the voiceprint data of family members, relatives, friends, etc. in the voiceprint DB. This registration is performed through the voice input unit 42 of the mobile terminal 10.
[0055] Next, in step S14, the server-side control unit 70 determines whether or not a person was identified in step S12, i.e., whether or not the voice belongs to a person registered in the voiceprint DB. If the determination here is positive, the server-side control unit 70 starts recording (storing in flash memory 64) in step S16. Note that the recorded voice data is converted into text data in the text data generation unit 54, so the timing of this recording start can also be said to be the timing of generating the text data. On the other hand, if the determination in step S14 is negative, the process returns to step S10.
[0056] After the determination in step S14 is affirmative and the process proceeds to step S16, the server-side control unit 70 continues recording until the voice input ceases for a predetermined number of seconds in step S18. If there is no voice input for the predetermined number of seconds, that is, if the voice input is deemed to have ended, the determination in step S18 is affirmative. If the determination in step S18 is affirmative, the terminal-side control unit 30 ends recording in step S20 and returns to step S10.
[0057] Thereafter, by repeating the above process, recording is performed each time a person registered in the voiceprint DB speaks. Note that the person who decides the recording timing may be managed in a DB separate from the voiceprint DB. This makes it possible to limit the person who decides the recording timing to, for example, the organizer of the meeting.
[0058] 2(a) illustrates a case where recording is started when a person speaks based on their voiceprint, but the present invention is not limited to this. For example, recording may be started when a frequency related to a telephone call (for example, the frequency of a ringtone) is input from the voice input unit 42. This allows a telephone conversation to be recorded without missing a word.
[0059] (Recording timing example 2) Fig. 2(b) is a flowchart showing the process of executing recording at a pre-registered time. Note that the process of Fig. 2(b) differs from that of Fig. 2(a) in that the recording timing is switched by switching whether or not to transmit audio data from the communication unit 32 of the mobile terminal 10 to the server 50.
[0060] In FIG. 2(b), in step S24, the terminal-side control unit 30 detects the current time via the time detection unit 24. Next, in step S26, the terminal-side control unit 30 determines whether it is a predetermined recording start time. Here, the recording start time may be predetermined when the mobile terminal 10 is shipped, or may be input in advance by a user or the like via the input unit 12. This recording start time may be, for example, a time when there is a lot of conversation with people and the amount of information is thought to be large (e.g., one hour immediately after arriving at work), or a time when concentration is likely to be broken (e.g., 30 minutes before or after lunch break, or overtime (after 8 p.m., etc.) when fatigue reaches its peak).
[0061] If the determination in step S26 is positive, the process proceeds to step S28, where the communication unit 32 starts transmitting the voice data input to the voice input unit 42 to the server 50 under the instruction of the terminal-side control unit 30. In this case, the voice data is stored (recorded) in the flash memory 64 via the communication unit 52 and the server-side control unit 70.
[0062] Next, in step S30, the terminal-side control unit 30 detects the current time via the time detection unit 24. Then, in the next step S32, the terminal-side control unit 30 determines whether the predetermined recording end time has arrived. If the determination here is positive, the process proceeds to step S34, but if the determination is negative, the process returns to step S30. If the process proceeds to step S34, the communication unit 32, under the instruction of the terminal-side control unit 30, stops transmitting audio data to the server 50. This ends the recording. Thereafter, the process returns to step S24, and the above process is repeated. This allows recording to be performed each time the recording start time arrives.
[0063] (Recording timing example 3) Fig. 2(c) is a flowchart showing the process of recording at the end of a pre-registered conference. Note that, like Fig. 2(b), the process of Fig. 2(c) also switches the recording timing by switching whether or not to transmit audio data from the communication unit 32 to the server 50.
[0064] In Fig. 2(c), in step S36, the terminal-side control unit 30 detects the current time via the time detection unit 24. Next, in step S38, the terminal-side control unit 30 extracts a scheduled meeting from a task list (described later) stored in the flash memory 28, and determines whether the current time is a predetermined time (e.g., 10 minutes) before the end of the meeting. If the determination here is positive, in step S40, recording is started in the same manner as in step S28 of Fig. 2(b).
[0065] In the next step S42, the terminal-side control unit 30 detects the current time via the time detection unit 24. Then, in the next step S44, the terminal-side control unit 30 determines whether the end time of the conference used in the determination of step S38 has arrived. If the determination here is affirmative, the process proceeds to step S46, but if the determination is negative, the process returns to step S42. If the process proceeds to step S46, the communication unit 32 stops transmitting audio data to the server 50 under the instruction of the terminal-side control unit 30. Thereafter, the process returns to step S36 and the above process is repeated. This makes it possible to record a predetermined time towards the end of the conference. The reason for recording towards the end of the conference is that the closer the conference is to the end, the more likely it is that the conclusion of the conference will be stated or the schedule for the next conference will be announced.
[0066] The process in Figure 2(c) may be configured to record continuously throughout the duration of the conference. Also, if the chairperson or presenters of the conference are registered in the task list, it may be combined with the process in Figure 2(a) to record only the voices of the registered chairperson or presenters.
[0067] (Recording timing example 4) Fig. 2(d) is a flowchart showing the process of executing recording based on information (here, the state of the user's muscles (degree of tension and relaxation)) input from the biometric information input unit 20. Note that, like Fig. 2(b) and Fig. 2(c), the process of Fig. 2(d) also switches the recording timing by switching whether or not to transmit voice data from the communication unit 32 to the server 50 side.
[0068] In Fig. 2(d), in step S50, the terminal-side controller 30 acquires the state of the user's muscles via the biometric information input unit 20. Next, in step S52, the terminal-side controller 30 compares the state of the muscles with a predetermined threshold value to determine whether or not the muscles are in a predetermined relaxed state. If the determination here is positive, in step S54, recording is started in the same manner as in step S28 of Fig. 2(b).
[0069] In the next step S56, the terminal-side control unit 30 acquires the muscle state again. In the next step S58, the terminal-side control unit 30 compares the muscle state with a predetermined threshold to determine whether or not the muscle state is in a predetermined state of tension. If the determination here is positive, the process proceeds to step S60, but if the determination here is negative, the process returns to step S56. If the process proceeds to step S60, the communication unit 32 stops transmitting audio data to the server 50 under the instruction of the terminal-side control unit 30. Thereafter, the process returns to step S50 and the above process is repeated. Through the above process, it is possible to determine the user's level of tension from the muscle state and automatically record audio when the user is too relaxed and not listening to others (for example, when dozing off).
[0070] In Fig. 2(d), the voice is recorded only when the person is too relaxed, but in addition to or instead of this, the voice may be recorded when the person is moderately tense, because when the person is moderately tense, there is a high possibility that important talk is being held.
[0071] In addition, at least one of a sweat sensor and a pressure sensor may be provided in the handset (portable terminal housing) to detect whether the user is in a tense or relaxed state based on the amount of sweat in the hand holding the handset or the strength with which the handset is held.
[0072] The outputs of the sweat sensor and pressure sensor may be sent to the terminal control unit 30, and the terminal control unit 30 may start recording using the voice input unit 42 when it determines that the user is in a tense or relaxed state.
[0073] The sweat sensor can be provided with multiple electrodes to measure the impedance of the hand. Psychological sweating caused by emotions such as emotion, excitement, and tension produces less sweat and lasts for a shorter period of time, so a sweat sensor can be provided on the receiver corresponding to the palm side of the metacarpal region, which sweats more than the fingers.
[0074] The pressure sensor may be a capacitance type, a strain gauge, or an electrostrictive element, and it may be determined that the user is in a tense state when they grip the receiver with a pressure that is, for example, 10% or more greater than the pressure with which they would normally grip the receiver.
[0075] Furthermore, at least one of the perspiration sensor and the pressure sensor may be provided in the portable terminal 10, or may be provided in a mobile phone or the like.
[0076] 2(a) to 2(d), even if the timing to start recording arrives, recording may not start if, for example, the mobile terminal 10 is located in a recording prohibited location. An example of a recording prohibited location could be inside a company other than the company where the user works.
[0077] (Audio data processing) Next, the processing of audio data that is performed after the audio data is recorded will be described with reference to Fig. 4 to Fig. 23. Fig. 4 is a flowchart relating to the processing of audio data.
[0078] In step S70 of FIG. 4, the server-side control unit 70 determines whether voice data has been recorded in the flash memory 64. If the determination here is affirmative, the process proceeds to step S72, where the text data generation unit 54 converts the voice data into text under the instruction of the server-side control unit 70. In this case, the voice data is converted into text every time there is a predetermined interval of time between audio interruptions. The server-side control unit 70 also registers the text data (text data) obtained by converting the voice data into text, the time the voice data was input to the voice input unit 42, the location where the voice data was input, and the voice level of the voice data in the stored data DB (FIG. 5) in the flash memory 64. Note that the time and location information registered here is transmitted from the communication unit 32 along with the voice data, as described above. Next, in step S74, the voiceprint analysis unit 55, under the instruction of the server-side control unit 70, performs a voiceprint analysis to identify the person who made the voice and registers the identified person in the stored data DB. If the process of step S12 in FIG. 2(a) has been performed, the content of step S12 may be registered in the stored data DB without performing step S74.
[0079] The data structure of the stored data DB is shown in Figure 5. The stored data DB stores the above-mentioned time, location, text data, speaker, voice level, task flag, and task priority. The task flag and task priority items will be described later.
[0080] Returning to FIG. 4, in the next step S76, a task determination subroutine is executed. In the task determination subroutine, the process of FIG. 6 is executed, as an example. In the process of FIG. 6, in step S110, the server-side control unit 70 determines whether the text data includes a date and time. Note that the date and time here includes not only specific dates and times such as what year, what month, what day, and what time, but also dates and times such as tomorrow, the day after tomorrow, morning, and afternoon. If the determination here is positive, the data is determined to be a task in step S114, and the process proceeds to step S78 in FIG. 4; however, if the determination here is negative, the process proceeds to step S112.
[0081] In step S112, the server-side control unit 70 determines whether the text data contains a specific word. Here, the specific word is a word related to a task, such as "to do," "please," "do (or "should," "do")," "let's (or "will")," "will," "plan," etc. These specific words may be stored in advance as a table in the hard disk 66 before shipping the device, or may be added by the user as needed. If the determination in step S112 is affirmative, the data is determined to be a task in step S114, and the process proceeds to step S78 in FIG. 4. On the other hand, if the determination in step S112 is negative, the data is determined to be not a task in step S116, and the process proceeds to step S78 in FIG. 4.
[0082] 4, in step S78, the server-side control unit 70 determines whether or not the determination in step S78 is a task as a result of the processing in Fig. 6. Below, the processing when the determination in step S78 is positive and when it is negative will be described.
[0083] (If the determination in step S78 is affirmative (if it is a task)) If the determination in step S78 is positive, the process proceeds to step S80, where the server-side control unit 70 sets the task flag in the stored data DB (FIG. 5) to ON. Next, in step S82, the extraction unit 58, under the instruction of the server-side control unit 70, extracts keywords based on the keyword DB (FIG. 7) stored on the hard disk 66. As shown in FIG. 7, the keyword DB associates keywords with detailed information, attributes, and confidentiality levels for those keywords. Therefore, the extraction unit 58 focuses on the keyword item in the keyword DB and extracts keywords registered in the keyword DB from the text data.
[0084] For example, if the text data is "I'm planning to meet with Ichiro Aoyama of Daitokyo Co., Ltd. at 1:00 PM on November 20th to discuss the software specifications for Cool Blue Speaker 2." Let's say that was the case.
[0085] In this case, the extraction unit 58 extracts "Cool Blue Speaker 2," "Software," "Specifications," "Daitokyo Co., Ltd.", and "Aoyama Ichiro," which are registered in the keyword DB of FIG. 7, as keywords.
[0086] The keyword DB must be created in advance. The contents of the keyword DB can be added or changed as needed (for example, during maintenance). In addition to attributes such as personal names, company names, and technical terms, keywords can also be registered using attributes such as patent information, budget information, and business negotiation information, as shown in Figure 7.
[0087] Returning to Fig. 4, in the next step S84, an analysis subroutine for each keyword is executed. Fig. 8 is a flowchart showing specific processing of the analysis subroutine in step S84.
[0088] 8, in step S120, the classification unit 60 acquires the confidentiality levels of keywords from the keyword DB under the instruction of the server-side control unit 70. Specifically, the classification unit 60 acquires "medium" as the confidentiality level of "Cool Blue Speaker 2" from the keyword DB, "medium" as the confidentiality level of "software", "medium" as the confidentiality level of "specifications", "high" as the confidentiality level of "Daitokyo Co., Ltd.", and "high" as the level of "Aoyama Ichiro."
[0089] Next, in step S122, the conversion unit 62, under the direction of the server-side control unit 70, executes a subroutine to convert the keyword based on the confidentiality acquired in step S120 and store the converted keyword in the flash memory 64.
[0090] Fig. 9 is a flowchart showing the specific processing of the subroutine of step S122. As shown in Fig. 9, first, in step S138, the conversion unit 62 selects one keyword from the keywords extracted by the extraction unit 58. Here, as an example, it is assumed that "Daitokyo Co., Ltd." is selected.
[0091] Next, in step S140, the conversion unit 62 determines whether the confidentiality level of the selected keyword is "high." As described above, "Daitokyo Co., Ltd." has a confidentiality level of "high," so the determination here is affirmative and the process proceeds to step S142. In step S142, the conversion unit 62 executes a subroutine to convert the keyword according to the confidentiality level. Specifically, the process is executed in accordance with the flowchart shown in FIG. 10.
[0092] In step S160 of Fig. 10, the conversion unit 62 determines whether or not the selected keyword contains a specific conversion word. Here, the specific conversion word refers to, for example, a word that is frequently used in company names (such as a joint stock company, a limited liability company, (stock), or (owned)), a word that is frequently used in government organizations (such as an organization, a ministry, or an agency), or a word that is frequently used in educational institutions (such as a university or a high school), as defined in the specific conversion word DB of Fig. 11.
[0093] In the case of the selected keyword "Daitokyo Kabushiki Kaisha," since it contains the specific conversion word "Kabushiki Kaisha," the determination in step S160 is affirmative, and the process proceeds to step S162. In step S162, the conversion unit 62 converts the specific conversion word based on the specific conversion word DB. In this case, the "Kabushiki Kaisha" part of "Daitokyo Kabushiki Kaisha" is converted to "Sha." Next, in step S164, a conversion subroutine for words other than the specific conversion word is executed.
[0094] Fig. 12 is a flowchart showing the specific processing of the conversion subroutine of step S164. As shown in Fig. 12, in step S190, the conversion unit 62 determines whether the part to be converted (the part other than the specific conversion word) is a place name. Here, the part to be converted, "Dai-Tokyo," includes a place name but is not the place name itself, so the determination is negative and the process proceeds to step S194.
[0095] In step S194, the conversion unit 62 determines whether the portion to be converted is a name. In this case, since it is not a name, the determination is negative and the process proceeds to step S198. Then, in step S198, the conversion unit 62 converts the portion to be converted, "DaiTokyo," into an initial "D." When the processing of step S198 ends, the process proceeds to step S165 in FIG. 10.
[0096] In step S165, the conversion unit 62 combines the words converted in steps S162 and S164. Specifically, "D" and "company" are combined to form "D Company."
[0097] Next, in step S168, the conversion unit 62 determines whether or not information is attached to the keyword to be converted, "Daitokyo Co., Ltd." Here, "attached with information" means that information has been entered in the information column of the keyword DB in FIG. 7. In this case, "Daitokyo Co., Ltd." is attached with "Electrical machinery, Shinagawa Ward, Tokyo," so the determination in step S168 is affirmative, and the process proceeds to step S170.
[0098] In step S170, the conversion unit 62 selects one piece of information that has not yet been selected from the accompanying information. Next, in step S172, the conversion unit 62 determines whether the confidentiality level of the selected information (for example, "electrical equipment") is "high" or "medium." If the confidentiality level of "electrical equipment" is "low," the determination in step S172 is negative, and the process proceeds to step S182. In step S182, the conversion unit 62 determines whether all of the information has been selected. Here, "Shinagawa Ward, Tokyo" has not yet been selected, so the determination is negative, and the process returns to step S170.
[0099] Next, in step S170, the conversion unit 62 selects the unselected information "Shinagawa Ward, Tokyo," and in step S172, determines whether the confidentiality level of "Shinagawa Ward, Tokyo" is "high" or "medium." Here, as shown in the keyword DB in FIG. 7, place names are defined as being "low" or conforming to the confidentiality level of the associated keyword. Therefore, "Shinagawa Ward, Tokyo" has a confidentiality level of "high" conforming to the keyword "Daitokyo Co., Ltd." Therefore, the determination in step S172 is affirmative, and the process proceeds to step S174. In step S174, the conversion unit 62 determines whether "Shinagawa Ward, Tokyo" contains a specific conversion word. If the determination here is negative, the process proceeds to step S180, where a conversion subroutine for converting information is executed. This conversion subroutine in step S180 is basically the same process as step S164 described above (FIG. 12).
[0100] That is, in Fig. 12, the conversion unit 62 determines in step S190 whether "Shinagawa-ku, Tokyo" is a place name. If the determination here is positive, the conversion unit 62 executes a conversion process in step S192 based on the place name DB shown in Fig. 13. Specifically, the conversion unit 62 converts "Shinagawa-ku, Tokyo" to "Southern Kanto" by using a conversion method with a confidentiality level of "high." Note that in the place name DB of Fig. 13, when the confidentiality level is "high," the place name is expressed as a location within a relatively wide area, and when the confidentiality level is "medium," the place name is expressed as a location within a narrower area than when the confidentiality level is "high."
[0101] When the processing of step S192 is completed, the process then proceeds to step S182 in FIG. 10. At this stage of step S182, all information (electrical machinery, Shinagawa-ku, Tokyo) has already been selected, so the determination at step S182 is affirmative, and the process proceeds to step S184. At step S184, the conversion unit 62 associates the converted information with the converted keyword (step S165 or S166). In this case, it becomes "Company D (electrical machinery, southern Kanto)". Then, the process proceeds to step S144 in FIG. 9.
[0102] In step S144 of FIG. 9, the converted keyword is stored in area A of the keyword storage DB (see FIG. 14) stored in flash memory 64. As shown in FIG. 14, the keyword storage DB has storage areas O, B, and C in addition to area A. Raw keyword data (pre-conversion keyword) is stored in area O. When the storage process is completed, the process proceeds to step S154, where it is determined whether all of the keywords extracted by extraction unit 58 have been selected. If the determination here is negative, the process returns to step S138.
[0103] Next, a case will be described where the conversion unit 62 selects "cool blue speaker 2" as the keyword in step S138. In this case, the keyword is "cool blue speaker 2" and the confidentiality level is "medium," so the determination in step S140 is negative, while the determination in step S146 is positive, and the process proceeds to step S148.
[0104] In step S148, a subroutine for converting keywords according to confidentiality is executed. Specifically, similar to step S142, the process of FIG. 10 is executed. In the process of FIG. 10, in step S160, the conversion unit 62 determines whether or not a specific conversion word is included in "Cool Blue Speaker 2." Since the determination here is negative, the process proceeds to step S166, where the conversion subroutine is executed. In the conversion subroutine of step S166, the process of FIG. 12 is executed, similar to steps S164 and S180 described above. In FIG. 12, since "Cool Blue Speaker 2" is neither a place name nor a person's name, the determinations in steps S190 and S194 are negative, and the conversion unit 62 performs initial conversion in step S198. In this case, the English spelling "Cool Blue Speaker 2" written alongside "Cool Blue Speaker 2 (Japanese spelling)" in the keyword DB is converted to initials (capital letters are converted to initials) to convert it to "CBS2."
[0105] When the processing of FIG. 12 is completed as described above, the process proceeds to step S168 of FIG. 10. However, since no information is attached to "Cool Blue Speaker 2" in the keyword DB of FIG. 7, the determination in step S168 is negative, and the process proceeds to step S150 of FIG. 9. In step S150, the converted keyword is stored in area B of flash memory 64 shown in FIG. 14. That is, conversion unit 62 stores the keyword itself in area O, and also stores "CBS2" in area B corresponding to the keyword. When the storage process is completed, the process proceeds to step S154, where it is determined whether all of the keywords extracted by extraction unit 58 have been selected. If the determination here is negative, the process returns to step S138 again.
[0106] Next, a case will be described where the conversion unit 62 selects the keyword "Aoyama Ichiro" in step S138. In this case, the confidentiality level of "Aoyama Ichiro" is "high," so the determination in step S140 is affirmative, and the process proceeds to step S142.
[0107] In step S142, the processing of FIG. 10 is executed, as described above. In the processing of FIG. 10, the result of step S160 is negative, and the processing proceeds to step S166 (processing of FIG. 12). In step S190 of FIG. 12, the determination is negative, and the processing proceeds to step S194. In step S194, the conversion unit 62 determines whether "Aoyama Ichiro" is a name. If the determination here is positive, the processing proceeds to step S196. Note that in step S194, "Aoyama Ichiro" is determined to be a name because the attribute of "Aoyama Ichiro" in the keyword DB of FIG. 7 is the name of a business partner.
[0108] In step S196, the conversion unit 62 converts "Aoyama Ichiro" into initials. Note that in step S196, if the confidentiality level of the keyword is "high", both the last name and the first name are converted into initials. That is, "Aoyama Ichiro" is converted into "AI". On the other hand, if the confidentiality level of the keyword is "medium", for example, like "Ueda Saburo" registered in the keyword DB of FIG. 7, only the first name is converted into initials. That is, "Ueda Saburo" is converted into initials "Ueda S". Note that only the last name may be converted into initials, resulting in "U Saburo".
[0109] Upon completion of the processing of step S196, the process proceeds to step S168 in FIG. 10. Here, as shown in FIG. 7, the keyword "Aoyama Ichiro" is accompanied by the information "Daitokyo Co., Ltd. Camera AF Motor October 15, 2009 Patent Training Seminar (Tokyo)." Therefore, the determination in step S168 is affirmative, and the process proceeds to step S170. Then, in step S170, for example, the information "Daitokyo Co., Ltd." is selected. Since the confidentiality level of "Daitokyo Co., Ltd." is "High" as described above, the determination in step S172 is affirmative, and the process proceeds to step S174. Then, since "Daitokyo Co., Ltd." includes the specific conversion word "Kaisho" (Co., Ltd.), the determination in step S174 is affirmative, and the conversion of the specific conversion word (step S176) and the conversion of words other than the specific conversion word (step S178) are executed. Note that steps S176 and S178 are similar to steps S162 and S164 described above. If the determination in step S182 is negative, the process returns to step S170.
[0110] Thereafter, steps S170 to S182 are repeated until all information has been selected. After all information has been selected, in step S184, the converted information is associated with the converted keyword. Here, for example, it is associated with "AI (camera, AFM, October 15, 2009, T Meeting (Tokyo))". Then, when storage in area A is completed in step S144 of FIG. 9, the process proceeds to step S154, where it is determined whether all of the keywords extracted by extraction unit 58 have been selected. If the determination here is negative, the process returns to step S138 again.
[0111] In the above process, if the determination in step S146 in FIG. 9 is negative, that is, if the confidentiality level of the keyword is "low," then in step S152, the keyword is stored as is in area C (and area O). If information is attached to the keyword, that information is also stored in area C. For example, as shown in FIG. 14, if the keyword is "SVS Co., Ltd.", it is stored in area C as "SVS Co., Ltd. Machinery Germany Munich."
[0112] In the above process, for example, if the keyword selected in step S138 is "software," the software is converted to an initial "SW," and the information <sponge> shown in FIG. 7 is associated with SW without being converted. In this case, <xx>This notation means that the word is treated equally with the keyword. In other words, it means that either "software" or "sponge" can be used. Therefore, when the above process is performed on the keyword "software," "SW" and "sponge" will be stored equally in area B of flash memory 64. The difference between "SW" and "sponge" will be explained later.
[0113] The above process is also performed for other keywords (here, "specifications"), and if the determination in step S154 in FIG. 9 is affirmative, the process proceeds to step S124 in FIG.
[0114] In step S124, the server-side control unit 70 acquires the weighting for the speaker's attributes. In this case, the weighting (Tw) is acquired from the speaker's job title based on the attribute weighting table shown in Fig. 15(a). For example, if the speaker is Ueda Saburo in Fig. 7, the weighting (Tw) is acquired as "2" for manager (M).
[0115] Next, in step S126, the server-side control unit 70 acquires a weight related to the voice level. In this case, the server-side control unit 70 acquires the weight (Vw) based on the weight table related to the voice level shown in FIG. 15(b) and the voice level stored in the storage data DB (see FIG. 5). As shown in FIG. 5, when the voice level is 70 db, the weight (Vw) is 3. Note that the higher the voice level, the higher the weight (Vw) because the higher the voice level, the stronger the request and the higher the importance in many cases.
[0116] Next, in step S128 of Fig. 8, the server-side control unit 70 acquires a weight for the keyword. In this case, the server-side control unit 70 acquires a weight (Kw) based on the keyword weight table shown in Fig. 15(c) and the keywords included in the text data in the stored data DB. In Fig. 15(c), "important," "critical," "very important," and "very important" are registered, so if these keywords are included in the text data, a weight (Kw) of 2 or 3 is acquired. Also, in step S128, the server-side control unit 70 determines how many keywords with a confidentiality level of "high" and how many keywords with a confidentiality level of "medium" are included in the text data, and acquires a weight (Cw) for the confidentiality of the text data based on the determination result and the keyword confidentiality weight table shown in Fig. 15(d). For example, if the text data contains two keywords with a confidentiality level of "high" and one keyword with a confidentiality level of "medium," the server-side control unit 70 obtains Cw=8 (=3×2+2×1).
[0117] When the process of step S128 in Fig. 8 is completed, the process proceeds to step S86 in Fig. 4. In step S86, the server-side control unit 70 calculates the task priority (Tp) and registers it in the stored data DB (Fig. 5). Specifically, the server-side control unit 70 calculates the task priority (Tp) using the following equation (1). Tp=Uvw×Vw+Utw×Tw +Ufw×Fw+Ukw×Kw+Ucw×Cw …(1)
[0118] In addition, Uvw, Utw, Ufw, Ukw, and Ucw in the above equation (1) are weighting coefficients that take into account the importance of each weight (Vw, Tw, Fw, Kw, Cw), and these weighting coefficients can be set by a user or the like via the setting unit 74.
[0119] Next, the process proceeds to step S88 in FIG. 4, where the server-side control unit 70 registers the keywords contained in the text data in the keyword recording DB shown in FIG. 16. Note that the keyword recording DB in FIG. 16 is created, for example, on a weekly, monthly, or yearly basis. The keyword recording DB in FIG. 16 records related information such as keywords used simultaneously with keywords contained in the text data (referred to as registered keywords), the person who said the registered keywords, and the date, time, and place at which the registered keywords were said. In addition, the number of times the registered keywords are associated with related information is recorded as the degree of association. Furthermore, the number of times the registered keywords are uttered is recorded as the frequency of appearance. Note that the search frequency item in the keyword recording DB in FIG. 16 will be described later.
[0120] After the process of step S88 is completed, the process returns to step S70.
[0121] (If the determination in step S78 is negative (if it is not a task)) Next, a case where the determination in step S78 is negative will be described. If the determination in step S78 is negative, the process proceeds to step S90, where the server-side control unit 70 turns off the task flag. Next, in step S92, the server-side control unit 70 determines whether the speaker is the user. If the determination here is positive, the process proceeds to step S94, where it determines whether the words uttered by the user are commands. Here, for example, as shown in the command DB in FIG. 17, the words "task list" are assumed to be a command for displaying a task list, the words "voice recording text" are assumed to be a command for displaying a voice recording list, and the word "convert" is assumed to be a command for the conversion process. Note that this command DB is assumed to be stored in the flash memory 28 of the mobile terminal 10 or the hard disk 66 of the server 50. For example, this command DB defines that if the user's voice is "task list," a task list such as that shown in FIG. 18(a) is displayed. Note that the task list will be described in detail later. Furthermore, the command DB defines that if the user's voice is "voice recording list," a voice recording list such as that shown in FIG. 18(b) is displayed. The voice recording list will be described in detail later.
[0122] 4, if the determination in step S94 is negative, the process returns to step S70, but if the determination in step S94 is positive, the process proceeds to step S96, where the server-side control unit 70 executes a subroutine that executes processing in accordance with the command. Specifically, the processes in Figures 19, 20, and 21 are executed simultaneously in parallel.
[0123] First, the processing on the server 50 side will be described with reference to the flowchart in Fig. 19. On the server 50 side, in step S202, the server-side control unit 70 determines whether the command is a display request. In this case, as described above, the commands "task list" and "voice recording list" correspond to display requests.
[0124] Next, in step S204, the server-side control unit 70 extracts data necessary for displaying in response to the command from the flash memory 64. For example, if the command is "task list," the server-side control unit 70 extracts text data to be displayed in the task list (text data for which the task flag in FIG. 5 is on) from the flash memory 64. In this case, the text data for which the task flag is on includes not only text data converted from voice data, but also text data directly input from the text data input unit 44. The task flag of directly input text data is turned on or off by the same process as that described above with reference to FIG. 6.
[0125] Next, in step S206, the server-side control unit 70 acquires the current location of the user. In this case, the server-side control unit 70 acquires the location information detected by the location detection unit 22 of the mobile terminal 10 via the terminal-side control unit 30, the communication units 32, 52, etc.
[0126] Next, in step S208, the server-side control unit 70 determines whether the location is one where security can be ensured based on the acquired location information (current location). Here, an example of a location where security can be ensured is within a company. The company location is registered in the following manner.
[0127] For example, a user connects the mobile terminal 10 to a PC (Personal Computer) via a USB or the like and starts a dedicated application using map information on the PC. The user then registers the company's location by specifying the company's address in the application. The address is specified by drawing using a mouse or the like. The company's location is expressed as an area of a predetermined size. Therefore, the company's location can be expressed by two diagonal points (latitude and longitude) of a rectangular area, as shown in the securable range DB of FIG. 22. The securable range DB of FIG. 22 is stored in the hard disk 66 of the server-side control unit 70.
[0128] That is, in step S208, the server-side control unit 70 refers to the secure range DB of FIG. 22, and if the user is within the range, it is determined that the user is located in a place where security can be ensured.
[0129] If the determination in step S208 is affirmative, the process proceeds to step S210. In this step S210, the server-side control unit 70 acquires conversion words associated with the keywords included in the extracted data from areas O, A, B, and C, and proceeds to step S214. On the other hand, if the determination in step S208 is negative, the process proceeds to step S212. In this step S212, the server-side control unit 70 acquires conversion words associated with the keywords included in the extracted data from areas A and B, and proceeds to step S214.
[0130] In step S214, the server-side control unit 70 transmits the extracted data and the conversion word associated with the keyword to the mobile terminal 10 via the communication unit 52.
[0131] If the determination in step S202 is negative, that is, if the command is a command other than a display request, the server-side control unit 70 performs processing in accordance with the command in step S216.
[0132] Next, the processing in the mobile terminal 10 will be described with reference to Fig. 20. In step S220 in Fig. 20, the terminal-side control unit 30 determines whether data has been transmitted from the server. In this step, the determination is affirmative after step S214 in Fig. 19 is executed.
[0133] Next, in step S221, the terminal-side control unit 30 determines whether or not the converted words for areas A, B, and C have been transmitted. Here, the determination is positive if step S210 in FIG. 19 has been performed, and the determination is negative if step S212 has been performed.
[0134] If the determination in step S221 is affirmative, in step S222, the terminal-side control unit 30 converts the keywords included in the extracted data with the conversion words in areas A, B, and C. That is, for example, if the extracted data is "I'm planning to meet with Ichiro Aoyama of Daitokyo Co., Ltd. at 1:00 PM on November 20th to discuss the software specifications for Cool Blue Speaker 2." In this case, using the conversion words for areas A, B, and C, "I'm planning to meet with AI (camera, AFM, October 15, 2009, T-kai (Tokyo)) from D Company (electrical equipment, southern Kanto) at 1:00 PM on November 20th to discuss SWSP on CBS2." is converted to:
[0135] On the other hand, if the determination in step S221 is negative, in step S223, the terminal-side control unit 30 converts the extracted data with the conversion word in area B and deletes the word in area A. In this case, the extracted data is "I'm scheduled to meet with about SWSP on CBS2 on November 20th at 1pm." In this way, in this embodiment, the display mode of data is changed depending on whether security is ensured or not.
[0136] After step S222 or step S223 is performed as described above, the process proceeds to step S224, where the terminal-side control unit 30 executes a process of displaying the converted text data at a predetermined position on the display unit 14. While this display may be simply arranged so that the task times (dates and times) are closest to the current time (date and time), in this embodiment, the tasks are instead displayed in order of decreasing task priority. This reduces the likelihood of the user overlooking important tasks and allows the user to prioritize high-priority tasks even when multiple appointments are double-booked. In the event of double-booking, the terminal-side control unit 30 may issue a warning via the warning unit 18. If a task includes a person involved in a lower-priority appointment, the terminal-side control unit 30 may automatically request a change in the schedule of the task from that person via email. However, the display is not limited to the order of decreasing task priority as described above, and it is also possible to display the tasks in chronological order. Furthermore, the tasks may be displayed in chronological order, with the font, color, size, etc. of high-priority tasks being displayed more prominently. Furthermore, tasks may be arranged in descending order of task priority, and tasks with the same task priority may be displayed in chronological order.
[0137] As described above, the processing of FIGS. 19 and 20 results in the screen displays shown in FIGS. 18(a) and 18(b). The recorded voice list of FIG. 18(b) includes a task item. The user can switch the task flag on or off by touching the task item on the touch panel. In this case, the server-side control unit 70 changes the task flag in FIG. 5 when it recognizes the user's operation to switch the task flag. This allows the user to manually change the task flag even if the on / off status of the task flag differs from the user's understanding as a result of the processing of FIG. 6. Note that, once the user turns on a task flag, the server-side control unit 70 may automatically turn on the task flag for any text data similar to the text data of that task.
[0138] 20, the terminal-side control unit 30 transmits the current location acquired by the location detection unit 22 to the server-side control unit 70, and converts and displays the text data using the conversion word transmitted from the server-side control unit 70. Therefore, in this embodiment, it can be said that the terminal-side control unit 30 restricts the display on the display unit 14 depending on the current location acquired by the location detection unit 22.
[0139] Next, with reference to Fig. 21, a process performed in parallel with the process of Fig. 20 will be described. In Fig. 21, in step S232, the terminal-side control unit 30 determines whether or not the document conversion button has been pressed by the user. The document conversion button is the button displayed in the upper right corner in Figs. 18(a) and 18(b). The user presses the document conversion button by operating the touch panel, keyboard, or the like. If the determination in step S232 is affirmative, the process proceeds to step S234; if the determination is negative, the process proceeds to step S238.
[0140] In step S234, the terminal-side control unit 30 determines whether a convertible keyword is displayed. Here, a convertible keyword means a keyword in which multiple conversion words are associated with one keyword, such as "SW" and "sponge" shown in FIG. 14, as described above. Therefore, if such a keyword is included in the text data displayed on the display unit 14, the determination here is affirmative, and the process proceeds to step S236. On the other hand, if the determination in step S234 is negative, the process proceeds to step S238.
[0141] When the process proceeds to step S236, the terminal-side control unit 30 converts the keyword. "I'm planning to meet with AI (camera, AFM, October 15, 2009, T-kai (Tokyo)) from D Company (electrical equipment, southern Kanto) at 1:00 PM on November 20th to discuss SWSP on CBS2." In the sentence displayed as follows, "SW" can be converted to "sponge." Therefore, the terminal-side control unit 30 converts "I'm planning to meet with AI (camera, AFM, October 15, 2009, T-kai (Tokyo)) from D Company (electrical equipment, southern Kanto) at 1:00 PM on November 20th to discuss the Sponge SP for CBS2." and convert it to display.
[0142] Even if a user cannot think of software when they see "SW," they can press the document conversion button and see the word "sponge" to associate sponge with software. This association may not be possible when seeing the word sponge for the first time, but if this association method is widely known within the company, it will be easy to recall software.
[0143] Next, in step S238, the terminal-side control unit 30 determines whether or not the display before conversion button (see FIGS. 18(a) and 18(b)) has been pressed. The user presses the display before conversion button when he or she wants to see the text without the keywords being converted. If the determination here is negative, the process returns to step S232, but if the determination here is positive, the process proceeds to step S240. In step S240, the terminal-side control unit 30 acquires the user's current location, and in step S242, determines whether or not the current location is in a location where security can be ensured. If the determination here is negative, that is, if the user is in a location where security cannot be ensured, it is necessary to restrict the user from viewing the text before conversion. Therefore, in step S252, the user is notified that the text cannot be displayed, and the process returns to step S232. The notification method in step S252 can be a display on the display unit 14 or a warning via the warning unit 18.
[0144] If the determination in step S242 is affirmative, the process proceeds to step S244, where the terminal-side control unit 30 displays questions (questions that the user can easily answer) on the display unit 14. The questions are assumed to be stored in the hard disk 66 on the server 50 side, and the terminal-side control unit 30 reads the questions from the hard disk 66 and displays them on the display unit 14. The questions and example answers may be registered in advance by the user, for example.
[0145] Next, in step S246, the terminal-side control unit 30 determines whether the user has input an answer by voice to the input unit 12. If the determination here is positive, then in step S248 the terminal-side control unit 30 determines whether it is the user's voice and whether the answer is correct. Whether it is the user's voice is determined using the results of voice analysis in the voiceprint analysis unit 55 on the server 50 side described above. If the determination here is negative, then in step S252 the user is notified that it cannot be displayed. On the other hand, if the determination in step S248 is positive, then the process proceeds to step S250, where the keyword is converted to its pre-conversion state using the conversion word in area O and displayed. Specifically, the sentence as it was input by voice, that is, in the above example, "I'm planning to meet with Ichiro Aoyama of Daitokyo Co., Ltd. at 1:00 PM on November 20th to discuss the software specifications for Cool Blue Speaker 2." After that, the process proceeds to step S232, and the above process is repeated. Note that, although the above description has been given of the case where the user answers the question by voice, the present invention is not limited to this, and the answer may also be input from a keyboard or the like. In this case, the terminal-side control unit 30 may determine whether or not to display the state before conversion based on the result of biometric authentication such as fingerprint authentication in addition to the answer to the question.
[0146] When the process of step S96 in FIG. 4 is completed in this manner, the process returns to step S70.
[0147] On the other hand, if the determination in step S92 in FIG. 4 is negative, i.e., if the speaker is not the user, the process proceeds to step S100, where the terminal-side control unit 30 displays information about the speaker. Here, the terminal-side control unit 30 performs display based on information received from the server-side control unit 70. Specifically, if the speaker is Ichiro Aoyama, the terminal-side control unit 30 receives that information from the server-side control unit 70 and displays "Ichiro Aoyama." If information related to Ichiro Aoyama is received, that information may also be displayed. If a task related to Ichiro Aoyama is received from the server-side control unit 70, that task may also be displayed.
[0148] In this way, for example, when Mr. Ichiro Aoyama calls out to the user, such as "Good morning," his name, related information, tasks, etc. can be displayed on the display unit 14. This can help the user to remember a person's name or information, or tasks to be done related to that person.
[0149] Next, in step S102, the server-side control unit 70 determines whether a word registered in the fuzzy word DB shown in Fig. 23 has been uttered. If the determination here is negative, the process returns to step S70, but if the determination here is positive, the process proceeds to step S104.
[0150] In step S104, the server-side control unit 70 and the terminal-side control unit 30 execute processing corresponding to the uttered word based on the fuzzy word DB of Fig. 23. Specifically, when "that matter" or "that particular matter" is uttered, the server-side control unit 70 refers to the keyword recording DB, extracts keywords whose appearance frequency is higher than a predetermined threshold from among the keywords included in the related information of the speaker, and transmits the extracted keywords to the terminal-side control unit 30. The terminal-side control unit 30 then displays the received keywords on the display unit 14. For example, if the speaker is Manager Yamaguchi and the appearance frequency threshold is 10, the keyword "Project A" in the keyword recording DB of Fig. 16 will be displayed on the display unit 14. 23, when a user utters "about (place name)," e.g., "about Hokkaido," keywords for which the speaker is included in the related information and the location (latitude, longitude) where the voice data was input is within a predetermined range (e.g., within Hokkaido), or keywords for which the speaker is included in the related information and the related information contains the word "Hokkaido," are extracted and displayed on the display unit 14. Furthermore, when a user utters, e.g., "about the ____ month, ____ day (MM / DD)," keywords for which the speaker is included in the related information and the date and time when the voice data was input matches ____ month, ____ day (MM / DD), or keywords for which the speaker is included in the related information and the related information contains the word "____ month, ____ day (MM / DD)," are extracted and displayed on the display unit 14. Furthermore, there are cases where it is easy to estimate from the keyword record DB of FIG. 16 what a certain person will say at a certain time (date and time). In such cases, keywords related to the speaker and the current time may be displayed.
[0151] By performing the above process in step S104, even if the speaker asks an ambiguous question, it is possible to automatically determine what the speaker is asking from the question and display it to the user. Note that in step S104, every time a keyword is displayed, the server-side control unit 70 updates the search frequency in the keyword recording DB. This search frequency can be used, for example, to preferentially display keywords with a higher search frequency.
[0152] (Image data processing) Next, the processing of image data by the server-side control unit 70 when the image data is transmitted to the server-side control unit 70 will be described with reference to Figures 24 and 25. Figure 24 is a flowchart relating to the input and storage processing of image data.
[0153] 24, first, in step S302, the server-side control unit 70 determines whether image data has been input. If the determination here is affirmative, the process proceeds to step S304. In step S304, the server-side control unit 70 acquires metadata of the image data (position information detected by the position detection unit 22 or time information detected by the time detection unit 24 when the image input unit 43 of the portable terminal 10 acquires an image).
[0154] Next, in step S306, the server-side control unit 70 determines whether the image data is data of an image to be masked based on the location information or time information. In this case, if the location where the image was acquired is within the company or the home of a telecommuter, or if the time when the image was acquired is a time that can be estimated as the time when the image was taken within the company or the home of a telecommuter, the image data is determined to be data of an image to be masked (the determination in step S306 is YES).
[0155] If the determination in step S306 is affirmative, the process proceeds to step S308, where the server-side control unit 70 identifies a portion of the image to be masked. For example, the server-side control unit 70 acquires a face recognition result of a subject (person) in the image from the face recognition unit 53, and if the subject is an important person (person with high confidentiality) inside or outside the company based on the face recognition result, the server-side control unit 70 identifies the subject as a portion to be masked. Also, for example, if the subject is a telecommuter based on the face recognition result by the face recognition unit 53, the server-side control unit 70 identifies the background portion of the subject as a portion to be masked. Also, for example, if a character recognition unit (not shown) recognizes characters written on a whiteboard or blackboard or characters displayed on a screen, and the result shows that the characters contain highly confidential words, the server-side control unit 70 identifies the character portion as a portion to be masked.
[0156] Next, in step S310, the server-side control unit 70 masks the portion of the image data identified in step S308 to generate a masked image (converted image), and stores the converted image in the flash memory 64. In this case, the server-side control unit 70 stores the converted image in a converted image storage area provided in the flash memory 64.
[0157] Next, in step S312, the server-side control unit 70 stores the original image data in the flash memory 64. In this case, the original image data is stored in a storage area for original images provided in the flash memory 64. The converted image and the original image of the converted image are stored in the flash memory 64 in an associated state.
[0158] If the judgment in step S306 is negative, i.e., if the image input to the server-side control unit 70 is not an image to be masked, then in step S312, the server-side control unit 70 stores the input image in the flash memory 64.
[0159] In this way, the image data input and storage process in the server-side control unit 70 is completed.
[0160] In the above, whether or not to perform masking processing is determined based on the position and time at which the image was acquired, but this is not limiting. If the face recognition unit 53 can detect the facial expression of the subject, i.e., whether the subject is smiling (soft expression) or frowning (hard expression), the server-side control unit 70 may determine, for example, not to perform masking processing on the image if the subject is smiling, but to perform masking processing on the subject or the background of the subject if the subject is frowning.
[0161] Next, the display processing of image data by the server-side control unit 70 will be described with reference to Fig. 25. Fig. 25 is a flowchart relating to the display processing of image data.
[0162] In the process of FIG. 25, first, in step S330, the server-side control unit 70 waits until a display instruction is input from the user via the mobile terminal 10. Next, in step S332, the server-side control unit 70 identifies the image data identified in the display instruction (included in the display instruction). Next, in step S334, the server-side control unit 70 determines whether a converted image corresponding to the identified image data exists in the flash memory 64. If the determination here is negative, the process proceeds to step S342, where the server-side control unit 70 transmits the original image data (data input to the server-side control unit 70) to the terminal-side control unit 30 via the communication units 52 and 32. The terminal-side control unit 30 displays the received image data on the display unit 14.
[0163] On the other hand, if the determination in step S334 is affirmative, the process proceeds to step S336. In step S336, the server-side control unit 70 acquires location information or time information. In this case, the server-side control unit 70 acquires the location or time of the portable terminal 10 at which the user inputs a display instruction.
[0164] Next, in step S338, the server-side control unit 70 determines whether to display the converted image based on the acquired location or time information. For example, if the acquired location is a location where security is not maintained (e.g., outside the company) or if it can be estimated from the time that the mobile terminal 10 is in a location where security is not maintained, the server-side control unit 70 determines to display the converted image (step S338 is positive). In this case, the server-side control unit 70 transmits the converted image data to the terminal-side control unit 30 via the communication units 52 and 32 in step S340. The terminal-side control unit 30 displays the received converted image data on the display unit 14. On the other hand, for example, if the acquired location is a location where security is maintained (e.g., inside the company) or if it can be estimated from the time that the mobile terminal 10 is in a location where security is maintained, the server-side control unit 70 determines not to display the converted image, i.e., to display the original image (step S338 is negative). In this case, in step S342, the server-side control unit 70 transmits the original image data to the terminal-side control unit 30 via the communication units 52 and 32. The terminal side control unit 30 displays the received original image data on the display unit 14 .
[0165] When the processing of step S340 or step S342 is completed in this manner, all the processing of FIG. 25 ends.
[0166] In the above description, a converted image is generated in advance based on the position or time at which the image was acquired in the processing of Fig. 24, and whether or not to display the generated converted image is determined based on the position or time of the mobile terminal 10 in the processing of Fig. 25 has been described, but the present invention is not limited to this. For example, the processing of Fig. 24 may be omitted, and the converted image may be generated only when it is determined in the processing of Fig. 25 that the converted image should be displayed based on the position or time of the mobile terminal 10. Even in this case, the converted image may be generated based on the position and time at which the image was acquired and / or the position and time at which the image is displayed.
[0167] When a converted image is generated in the process of Fig. 24, the converted image may be always displayed (without converting the original image) regardless of the display position or time. Also, the user may be allowed to set whether to always display the converted image or to perform the process of Fig. 25.
[0168] In the above description, the server-side control unit 70 acquires information about the location or time at the time the display instruction is input and determines whether to display the converted image data or the original image data based on the information. However, this is not limited to this. For example, the server-side control unit 70 may monitor the location of the mobile terminal 10 that issued the display instruction, display the original image data while the mobile terminal 10 is in a secure location, and display the converted image data when the mobile terminal 10 moves to a location where security is not ensured. This allows for appropriate display that takes into account the movement of the mobile terminal 10.
[0169] If the image data is video data, the playback unit 16 can output audio associated with the image from the speaker when displaying the image (video). If the audio data of the video is converted into text data using the method described above, the playback unit 16 may display the text data as a caption when displaying the image. In this case, if the text data has been changed to a superordinate concept, the playback unit 16 may display the changed text data. When outputting audio from the speaker, the playback unit 16 may output audio data that is the text data converted to a superordinate concept. In this case, audio data does not exist in the text data converted to a superordinate concept. Therefore, in such cases, a speech synthesis technique is used that selects speech segments to be used for speech synthesis (e.g., "gu", "ru", "gu", "ru") from a large number of speech segments in pre-recorded speech waveform data (speech database) by referring to phonetic symbols. The speech synthesis technique is described, for example, in Japanese Patent No. 3,727,885. When the image data displayed on the display unit 14 is converted image data, sound may not be output.
[0170] Furthermore, if the masked portion of the converted image is a person's face, the server-side control unit 70 may display, near that portion, information about the person (such as name and position) recognized by the face recognition unit 53. In this case, the person's information may be initials, etc.
[0171] In the above, an example has been described in which the server-side control unit 70 displays a converted image when it is determined that security is not ensured based on the location and time, but this is not limiting. For example, when security is not ensured, neither the image data nor the converted image data may be displayed. Furthermore, depending on the level of security ensured, either the image data or the converted image data may be displayed, or neither image data may be displayed.
[0172] The determination of whether to generate and display converted image data is not limited to the above, and for example, the server-side control unit 70 may digitize elements such as the position and time at which the image was acquired, the position and time at which the image is displayed, the people appearing in the image, and the people's facial expressions, and determine whether to generate and display converted image data based on the total value of each element.Furthermore, the server-side control unit 70 may weight each element and determine whether to generate and display converted image data based on the total value of each weighted element.
[0173] Next, the process of erasing data acquired by the mobile terminal 10 and the server 50 will be described with reference to FIGS.
[0174] (Data deletion process (Part 1: Deleting converted data (text data))) FIG. 26 is a flowchart showing a process of erasing information acquired by the mobile terminal 10 from the server 50. As shown in FIG. 26, the terminal-side control unit 30 determines in step S260 whether a predetermined time (e.g., 2 to 3 hours) has passed since the data was acquired. If the determination here is affirmative, the process proceeds to step S262, where the terminal-side control unit 30 erases the text data (including the pre-conversion words and the post-conversion words) stored in the flash memory 28. On the other hand, even if the determination here is negative, the terminal-side control unit 30 determines in step S264 whether the user has moved from inside the company to outside the company. If the determination here is positive, the process proceeds to step S262, where the data is erased in the same manner as described above. Note that if the determination here is negative, the process returns to step S260. In this way, by erasing the data when a predetermined time has passed since the data was acquired or when security can no longer be ensured, it is possible to prevent important data from being leaked, etc. Although the above description has been given of the case where all text data is erased, this is not limiting, and in step S262, only the most important data may be erased. For example, only the data in area A and the data in area O may be erased.
[0175] In the process of FIG. 26, if the user (portable terminal 10) is located outside the company from the beginning, the converted data may be erased from the flash memory 28 immediately after being displayed on the display unit 14.
[0176] (Data deletion process (part 2: Deleting audio data)) The server-side control unit 70 executes the deletion process of Fig. 27 for each piece of voice data. In step S270 of Fig. 27, the server-side control unit 70 determines whether the text data generation unit 54 has (been able to) convert the voice data into text data. If the determination here is negative, the process proceeds to step S280, but if the determination here is positive, the process proceeds to step S272, where the server-side control unit 70 acquires the name of the person who spoke the voice data. Here, the server-side control unit 70 acquires the name of the person who spoke the voice data from the voiceprint analysis unit 55, and proceeds to step S274.
[0177] In step S274, the server-side control unit 70 determines whether the person who spoke is other than the user. If the determination here is positive, the server-side control unit 70 deletes the voice data converted into text data in step S276. On the other hand, if the determination in step S274 is negative, that is, if the voice data is the user's own voice data, the server-side control unit 70 proceeds to step S278, where the voice data is deleted after a predetermined time has elapsed, and all processing in FIG. 27 ends.
[0178] On the other hand, if the determination in step S270 is negative and the process proceeds to step S280, the server-side control unit 70 enables the audio data to be played back. Specifically, the server-side control unit 70 transmits the audio data to the flash memory 28 of the portable terminal 10. In step S280, the server-side control unit 70 warns the user via the warning unit 18 that the audio data could not be converted into text data. If, based on this warning, the user inputs an instruction to play back the audio data from the input unit 12 of the portable terminal 10, the audio data stored in the flash memory 28 is played back via the playback unit 16.
[0179] Next, in step S282, the server-side control unit 70 erases the audio data transmitted to the flash memory 28 (that is, the audio data reproduced by the reproduction unit 16), and all the processing in FIG. 27 ends.
[0180] By performing the voice data deletion process as described above, the amount of voice data stored in the server 50 can be reduced, and therefore the storage capacity of the flash memory 64 of the server 50 can be reduced. Also, by deleting voice data other than that of the user immediately after it has been converted into text data, privacy can be respected. In this embodiment, if a voice synthesis unit is provided in the server 50, even if the voice data is deleted, the voice can be reproduced from the text data stored in the flash memory 64 using the function of the voice synthesis unit.
[0181] (Data Deletion Process (Part 3: Deletion of Image Data)) The image data (including the original image data and the converted image data) is erased from the flash memory 64 by the server-side control unit 70, for example, after a predetermined time has elapsed since the data was acquired. However, this is not limiting, and the data may be erased using the same logic as for text data, for example.
[0182] (Data Deletion Process (Part 4: Task Deletion)) The server-side control unit 70 deletes tasks according to the following rules. (1) If the task is about an external meeting In this case, the task is deleted when the current location detected by the location detection unit 22 matches the conference location specified in the task and the current time detected by the time detection unit 24 has passed the conference start time specified in the task. Note that if the current time has passed the conference start time but the current location does not match the conference location, the server-side control unit 70 causes the warning unit 18 to issue a warning to the user via the terminal-side control unit 30. This can prevent the user from forgetting to perform the task. Alternatively, the warning can be issued a predetermined time before the task is to be performed (e.g., 30 minutes before). This can prevent the user from forgetting to perform the task. (2) If the task is about an internal meeting In this case, a position detection unit that can detect entry into a conference room, such as RFID, is employed as the position detection unit 22, and the task is deleted when the current position detected by the position detection unit 22 matches the conference room specified in the task and the current time detected by the time detection unit 24 is past the conference start time specified in the task. In this case, a warning can also be used in conjunction with this, as in (1) above. (3) The task is about shopping and the location of the shopping is specified. In this case, the task is deleted when the current location detected by the location detection unit 22 matches the location specified in the task and a voice such as "Thank you" is input from the voice input unit 42 or purchase information is input wirelessly from the POS register terminal to the input unit 12. Note that, in addition to input from the POS register terminal, if the portable terminal has an electronic money function, for example, the task may be deleted when payment is completed using that function. Furthermore, if the image input unit 43 is a life log camera that is hung from the neck or attached to the ear, the task may be deleted based on the image capture results of the life log camera. (4) Other cases where a time is specified in the task In this case, if the current time detected by the time detection unit 24 passes the execution time specified in the task, the task is deleted.
[0183] As described above in detail, according to this embodiment, the communication unit 52 inputs an image, location information, and time information, and the server-side control unit 70, upon receiving an image output instruction (display instruction) from the user, determines whether to impose a restriction on the display of the image based on the location or time. This makes it possible to restrict the display of the image according to the location and time (masking processing: mosaic processing, resolution adjustment, etc.). Furthermore, in this embodiment, for example, if the location information and time information input by the communication unit 52 are information on the location and time at which the image was captured, the server-side control unit 70 can determine to impose a restriction on the display of the image if it is determined that at least a portion of the image is likely to be confidential based on the location and time at which the image was captured. Furthermore, for example, if the location information and time information input by the communication unit 52 are information on the location and time at which the image will be viewed, the server-side control unit 70 can estimate the location at which the image will be viewed based on the information and determine to impose a restriction on the display of the image according to the location. In this way, the server-side control unit 70 determines whether or not to restrict the display of images based on location information and time information, thereby preventing information leakage from images, protecting privacy, and improving usability.
[0184] Furthermore, in this embodiment, when the server-side control unit 70 determines to impose a restriction on image display, it imposes a restriction on at least a part of the image, so that, for example, in the case of an image taken during a meeting, it is possible to restrict the display of a person's face in the image, or to restrict the display of text in the image (text written on a whiteboard or blackboard, text projected on a screen, etc.). This makes it possible to impose appropriate display restrictions (such as hiding confidential parts to the extent that the content of the image can be revealed) even when displaying an image in which the parts that should be restricted from display differ depending on location information or time information.
[0185] Furthermore, in this embodiment, when the server-side control unit 70 determines from the location information that a location has changed from a secure location to an unsecured location, it imposes restrictions on at least a portion of the image. Therefore, when a user moves from a secure location to an unsecured location, the user can view the image without having to change the way he or she views the image accordingly (for example, changing from openly viewing images to viewing images while being mindful of the surroundings).
[0186] Furthermore, in this embodiment, the server-side control unit 70 identifies the location in the image where display restrictions are to be applied based on at least one of the location information and the time information, thereby enabling appropriate display restrictions (such as hiding confidential parts to the extent that the contents of the image can be seen) according to the location information and the time information.
[0187] In this embodiment, the face recognition unit 53 inputs information about the subject in the image to the server-side control unit 70, and the server-side control unit 70 identifies the portion of the image where a display restriction is to be applied based on the input information. This allows a display restriction to be applied so as to hide a person of confidentiality among the subjects. It is also possible to apply a display restriction so as to hide text of confidentiality among the subjects.
[0188] Furthermore, in this embodiment, the server-side control unit 70 displays the information about the subject input to the server-side control unit 70 by the face recognition unit 53 together with the image. Therefore, even if a display restriction is imposed on the image to hide the face, the information about the subject (e.g., the person's name) is displayed, allowing the user to easily identify the person to whom the display restriction is imposed.
[0189] In this embodiment, the server-side control unit 70 converts information about the subject based on a predetermined rule and displays it together with the image. This allows the converted information (initials, etc.) of a person's name to be displayed together with the image, making it possible to keep the person's name confidential.
[0190] Furthermore, in this embodiment, the server-side control unit 70 can determine whether to impose restrictions on image output based on information about the facial expression of the person in the image. That is, if the facial expression of the person in the image is stiff, it is highly likely that the image has confidentiality, so it can determine to impose a display restriction on the image, and if the facial expression of the person in the image is soft, it is low likely that the image has confidentiality, so it can determine not to impose a display restriction on the image. This enables appropriate image display restrictions.
[0191] Furthermore, in this embodiment, when at least one of voice data and text data is input from the communication unit 52, the server-side control unit 70 imposes restrictions on the output of at least one of the voice data and text data based on at least one of the location information and the time information. As a result, in the voice data and text data, similar to images, restrictions are imposed on the output of confidential parts (words), thereby ensuring security and privacy protection.
[0192] In this case, in this embodiment, if the server-side control unit 70 determines that it is necessary to restrict the output of voice data, it outputs text data converted and generated from the voice data, so that even if it is necessary to restrict the output of voice data, the content of the voice data can be confirmed in text data.
[0193] Furthermore, in this embodiment, when the server-side control unit 70 determines that it is necessary to restrict the output of voice data, it converts specific nouns in at least one of the voice data and text data into other nouns and outputs them. Therefore, by converting specific nouns into nouns that can be understood by specific people, it is possible to ensure security and privacy protection while allowing specific people to recognize the contents of the voice data and text data.
[0194] Furthermore, in this embodiment, the device has a first housing 11 having a display unit 14 and a second housing 51 having a server-side control unit 70. Since the first housing 11 and the second housing 51 are separate, the first housing 11 side is the mobile terminal and the second housing 51 side is the server, which makes it possible to make the mobile terminal smaller and lighter than when the mobile terminal is provided with functions such as the server-side control unit 70.
[0195] In this embodiment, the first housing 11 is provided with a position detection unit 22 that detects position information of the first housing 11, and the communication unit 52 inputs the detection result by the position detection unit 22. This makes it possible to restrict image display according to the position of the first housing 11.
[0196] Furthermore, this embodiment includes a communication unit 52 to which information is input, an extraction unit 58 that extracts predetermined keywords from the data input to the communication unit 52, a classification unit 60 that classifies the keywords extracted by the extraction unit 58 into keywords with a "high" confidentiality level and keywords with a "medium" confidentiality level, and a conversion unit 62 that converts keywords with a "high" confidentiality level using a predetermined conversion method and converts keywords with a "medium" confidentiality level using a conversion method different from that used for keywords with a "high" confidentiality level. In this way, by classifying keywords according to confidentiality level and performing different conversions according to each level, it is possible to display data taking the confidentiality level into consideration, thereby improving usability.
[0197] Furthermore, even if the voice data does not contain confidential keywords, if the image data has confidentiality or privacy issues, the server-side control unit 70 may prohibit or mask the display of the image data on the display unit 14 of the mobile terminal 10, and may also prohibit or restrict the playback of text data and voice data.
[0198] Furthermore, in this embodiment, the communication unit 52 that communicates with the portable terminal 10 transmits the result of the conversion by the conversion unit 62 to the portable terminal 10, so that the portable terminal 10 can display data that takes into consideration the confidentiality level without processing the data.
[0199] Furthermore, in this embodiment, a text data generation unit 54 is provided that generates text data from audio data, and the extraction unit 58 extracts keywords from the text data generated by the text data generation unit 54, making it possible to easily extract keywords.
[0200] In addition, in this embodiment, since keywords are converted to initials, each keyword can be easily converted without creating a special conversion table for each keyword. Furthermore, if the keyword is a name, if the confidentiality level is "high," both the last name and first name are converted to initials, and if the confidentiality level is "medium," either the last name or first name is converted to initials, making it possible to display information according to the confidentiality level. Furthermore, if the keyword is a place name, if the confidentiality level is "high," the information is converted to information about a specified area (location information within a wide range), and if the confidentiality level is "medium," the information is converted to information about an area smaller than the specified area (location information within a narrower range). This also makes it possible to display information according to the confidentiality level.
[0201] In addition, this embodiment includes a position detection unit 22 that detects position information, an input unit 12 for inputting, a display unit 14 that displays information related to the input, and a terminal-side control unit 30 that restricts display on the display unit 14 according to the position detected by the position detection unit 22. In this way, by restricting display according to the position, it is possible to perform display that takes security into consideration, and ultimately to improve usability.
[0202] Furthermore, in this embodiment, when it is determined that security cannot be maintained based on the output of the position detection unit 22, the terminal-side control unit 30 restricts the display on the display unit 14, thereby making it possible to restrict display while taking security into appropriate consideration. Furthermore, in this embodiment, when it is determined that security can be maintained based on the output of the position detection unit 22, the terminal-side control unit 30 releases at least a part of the restriction on display on the display unit 14, making it possible to restrict display while taking security into appropriate consideration.
[0203] Furthermore, since the personal assistant system 100 of this embodiment includes the portable terminal 10 that imposes display restrictions in consideration of security as described above, and the server 50 that imposes display restrictions on at least a portion of the data input from the portable terminal 10, data with display restrictions can be displayed on the display unit 14 of the portable terminal 10 without imposing display restrictions on at least a portion of the data on the portable terminal 10. This reduces the processing load on the portable terminal 10, and as a result, the portable terminal 10 can be simplified and made smaller and lighter.
[0204] Furthermore, this embodiment includes a display unit 14 that displays text data, a voice input unit 42 that inputs voice, and a terminal-side control unit 30 that displays information related to the voice on the display unit according to the results of voice analysis. Therefore, as shown in step S100 of FIG. 4, when a person utters a voice such as "Good morning," information about that person (such as their name, other registered information, or tasks to be performed for that person) can be displayed on the display unit 14. This allows the user to recall the person by looking at the display unit 14, even if they have forgotten who uttered the voice. Thus, this embodiment provides a user-friendly personal assistant system 100 and a portable terminal 10. In this case, appropriate display is possible by utilizing the analysis results of the voiceprint analysis unit 55 that analyzes the voiceprint of the voice.
[0205] Furthermore, in this embodiment, the terminal-side control unit 30 and the server-side control unit 70 display information related to predetermined words (e.g., "that matter" or "the matter in Hokkaido") included in the voice data on the display unit 14 in accordance with the analysis results of the voiceprint analysis unit 55. Therefore, even when an ambiguous question such as "that matter" or "the matter in Hokkaido" is asked, the user can recall the matter by checking the display unit 14. From this point of view, this embodiment can also be said to provide an easy-to-use personal assistant system 100 and portable terminal 10. Furthermore, in this embodiment, information related to the predetermined word (e.g., "the matter in Hokkaido") is selected according to the frequency of input into the input unit together with the predetermined word (e.g., "Hokkaido") and displayed on the display unit 14 (step S104 in FIG. 4), thereby enabling appropriate information display.
[0206] Furthermore, in step S104 of FIG. 4, information corresponding to the position at which the voice data was input is displayed on the display unit 14, so that appropriate information display is also possible in this respect.
[0207] Furthermore, in step S104 of FIG. 4, information corresponding to the time when the voice data was input (such as information input within a predetermined time from the time when the voice data was input) is displayed on the display unit 14, so that appropriate information can be displayed in this respect as well.
[0208] This embodiment also includes a voice input unit 42 that inputs voice, a text data generation unit 54 that generates text data based on the voice data input to the voice input unit 42, a voiceprint analysis unit 55 that analyzes voiceprint data of the voice data input to the voice input unit 42, and an erasure unit 76 that erases the voice data after the text data has been generated by the text data generation unit 54 in accordance with the analysis results of the voiceprint analysis unit 55. This makes it possible to reduce the storage capacity required for the flash memory 64 by erasing the voice data after the text data has been generated. Furthermore, this embodiment erases the voice data in accordance with the analysis results of the voiceprint analysis unit 55, thereby making it possible to achieve good usability that takes privacy into consideration by erasing the voice data of a specific person.
[0209] In addition, this embodiment includes a communication unit 52 to which information is input, an extraction unit 58 that extracts predetermined keywords from the data input to the communication unit 52, a classification unit 60 that classifies the keywords extracted by the extraction unit 58 into keywords with a confidentiality level of "high" and keywords with a confidentiality level of "medium", and a conversion unit 62 that converts keywords with a confidentiality level of "high" using a predetermined conversion method and converts keywords with a confidentiality level of "medium" using a conversion method different from that used for keywords with a confidentiality level of "high". In this way, by classifying keywords according to confidentiality level and performing different conversions according to each level, it becomes possible to display data taking the confidentiality level into consideration.
[0210] Furthermore, in this embodiment, the voiceprint analysis unit 55 analyzes whether the voiceprint data of the voice data is that of a registered user, and the erasure unit 76 erases voices other than those of the user, thereby effectively reducing the storage capacity of the flash memory 64 and further enhancing consideration for privacy.
[0211] In this embodiment, the erasure unit 76 determines the time period from analysis until erasure for the user's voice to that for the voice of a person other than the user (steps S276 and S278). This allows the user's voice to be erased after a predetermined time, further reducing the storage capacity.
[0212] In this embodiment, when the text data generation unit 54 cannot generate text data from the voice data, the warning unit 18 issues a warning, allowing the user to recognize that text data could not be generated from the voice data. Also, when the text data generation unit 54 cannot generate text data from the voice data (when step S270 is negative), the playback unit 16 plays back the voice data in response to a user instruction, allowing the user to check the content that could not be converted into text data by playing back the voice data.
[0213] Furthermore, according to this embodiment, the system includes a display unit 14 for displaying, a voice input unit 42 for inputting voice, a weighting unit 56 for weighting the input voice based on at least one of the volume, frequency, and meaning, and control units 70, 30 for changing the display mode of tasks on the display unit based on the voice input by the voice input unit 42 and the weighting by the weighting unit 56. As a result, the display mode of tasks on the display unit 14 is changed based on the weighting performed by the weighting unit 56 depending on the input method of the voice data, the content of the voice data, etc., so that a display mode according to the weight (importance) of the voice data can be realized. This makes it possible to improve usability.
[0214] Furthermore, according to this embodiment, the weighting unit 56 uses at least the frequency of the voice data to identify the person who made the voice and assigns weighting according to that person (job title in this embodiment), thereby enabling appropriate weighting of the importance of the voice data.
[0215] Furthermore, according to this embodiment, the weighting unit 56 assigns weighting in accordance with confidentiality based on the meaning of the voice, so that appropriate weighting can be performed in relation to the importance of the voice data.
[0216] In addition, in this embodiment, if the voice input from the voice input unit 42 includes date information, tasks can be displayed based on the date information, so that the device can also function as a regular schedule. In addition, in this embodiment, tasks are displayed taking into account information about the time detected by the time detection unit 24 or the date information in the calendar unit 26, so that tasks to be done can be displayed in order of proximity to the current time or furthest from the current time.
[0217] Furthermore, in this embodiment, the text data generating unit 54 that converts the voice input from the voice input unit 42 into text data is provided, so the weighting unit 56 can weight the text data. This allows weighting to be performed more easily than when handling voice data.
[0218] Furthermore, in this embodiment, the display order, color, display size, display font, etc. are changed based on the weighting results, so the weighting results can be expressed in various ways.
[0219] Furthermore, in this embodiment, the display mode on the display unit is changed according to the output of the position detection unit 22 that detects the position. In other words, when it is determined that a task has been executed based on the current position, the task is not displayed (deleted), which makes it possible to reduce storage capacity.
[0220] Furthermore, in this embodiment, whether or not the voice data is a task is determined based on whether or not it contains a standard word, and this determination result is used to determine whether or not to display it on the display unit 14, so it is possible to automatically determine whether or not the voice data is a task, and also to automatically determine whether or not to display it on the display unit.
[0221] In addition, in this embodiment, a setting unit 74 is provided in the server 50 to enable the user to set the weighting, so that the user can set the weighting according to his or her own preferences.
[0222] Furthermore, according to this embodiment, the system includes a voice input unit 42 that inputs voice, a text data generation unit 54 that converts the input voice into text data, and a server-side control unit 70 that starts conversion by the text data generation unit 54, i.e., starts recording and starts conversion into text data, when the voice input unit 42 inputs a specific frequency. Therefore, when a person speaks and a voice of a specific frequency is input, recording and conversion into text data are started based on the input voice (see FIG. 2(a)), so that recording and conversion into text data can be started automatically. This simplifies user operation and improves usability.
[0223] In addition, in this embodiment, since conversion to text data can be started when the voice input unit 42 inputs a frequency related to a telephone call, it is possible to record the voice of a telephone call and convert it to text data from the moment the telephone rings, for example. This makes it possible to record and convert the entire telephone conversation into text data without missing a word.
[0224] In addition, in this embodiment, since recording and conversion to text data can be started at an appropriate timing based on the task, for example, when the date and time of the meeting arrives, it is possible to simplify the user's operation and improve usability. Also, since recording and conversion to text data can be started according to the end time of the meeting (see FIG. 2(c)), it is possible to automatically start recording and converting voice data to text data during the time period when the most important things in the meeting are likely to be discussed.
[0225] Furthermore, in this embodiment, recording and conversion to text data can be started at an appropriate time based on the user's biometric information (see FIG. 2(d)), which also simplifies user operations and improves usability.
[0226] Furthermore, in this embodiment, recording and conversion to text data can be started when the current time reaches a predetermined time (see Figure 2(b)), which also simplifies user operations and improves usability.
[0227] Furthermore, in this embodiment, conversion by the text data generator 54 can be prohibited depending on the detection result of the position detector 22, so that recording can be automatically prohibited in cases where recording is problematic, such as an external conference, for example. This further improves usability.
[0228] In the above embodiment, the confidentiality level is determined for each word, but this is not limited to this. For example, the classification unit 60 may classify words used in business as words with a high confidentiality level and words used privately as words with a low confidentiality level.
[0229] In the above embodiment, a case has been described in which a keyword is converted and displayed when the current location detected by the location detection unit 22 of the mobile terminal 10 is a location where security is not ensured, i.e., a case in which the display on the display unit 14 is restricted. However, this is not limiting. For example, if the time detected by the time detection unit 24 is a predetermined time (e.g., during working hours), the display on the display unit 14 may be restricted. This also makes it possible to provide a display that takes security into consideration, as in the above embodiment. In case of performing such control, the current time may be acquired instead of acquiring the user's current location in step S206 of FIG. 19, and it may be determined whether the current time is a time when security can be ensured, instead of determining whether the current location is a location where security can be ensured in step S208.
[0230] In the above embodiment, the determination of whether the audio data is a task is made based on the presence or absence of date and time information and the type of ending of the audio data, but this is not limited to this, and the task determination may also be made based on, for example, the intonation of the audio data.
[0231] In the above embodiment, a case has been described in which words with a "high" confidentiality level and words with a "medium" confidentiality level are converted into their superordinate initials. However, this is not limiting. For example, a converted word for each word may be defined in the keyword DB. In this case, for example, a converted word for the keyword "camera" may be defined as a superordinate concept of camera, such as "precision equipment," or a subordinate concept, such as "photography equipment." In this case, if "camera" has a "high" confidentiality level, it may be converted to "precision equipment," and if "camera" has a "medium" confidentiality level, it may be converted to "photography equipment." In this way, by converting into superordinate and intermediate concept words according to the confidentiality level, it is possible to perform display taking into account the security level. Furthermore, when monetary information such as a budget is registered in the keyword DB, a superordinate concept of the monetary information expressed in digits may be defined.
[0232] In the above embodiment, the voice is in Japanese, but it may be in a foreign language such as English. In a foreign language (e.g., English), it may be determined whether or not it is a task based on the presence or absence of a predetermined word or a predetermined syntax.
[0233] In the above embodiment, a case where a flash memory 28 is installed in order to make the portable terminal 10 smaller and lighter has been described, but in addition to or instead of this, a storage device such as a hard disk may also be installed in the portable terminal 10.
[0234] In the above embodiment, the case where the mobile terminal 10 is connected to an external PC and the setting is performed on the external PC when setting the company location, etc., has been described. However, this is not limited to this. For example, the company location may be registered in advance on the hard disk 66 of the server 50, and the company location may be downloaded from the hard disk 66. Also, for example, an application for setting the company location, etc. may be installed on the mobile terminal 10, so that the company location, etc. can be set on the mobile terminal 10.
[0235] In the above embodiment, the task priority is calculated based on the above formula (1), but this is not limiting and other formulas may be used to calculate the task priority. For example, the weights may simply be added or multiplied. Furthermore, instead of using the above formula (1) to calculate the task priority, it is also possible to select one of the weights and determine the task priority in descending order of the selected weight. In this case, the user may be able to set which weight is used to determine the task priority.
[0236] In the above embodiment, the initial keyword (for example, "SW" for software) and the image-based keyword (for example, "sponge" for software) are displayed first, but the initial keyword may be displayed first. Alternatively, the initial keyword and the image-based keyword may be displayed simultaneously.
[0237] In the above embodiment, when the voice of a person other than the user is input to the input unit 12, the name and information of the speaker are displayed. However, the present invention is not limited to this. For example, an image related to the speaker, such as a photograph of the speaker's face, may be displayed. In this case, it is necessary to store the image in the hard disk 66 of the server 50, for example, and to register the image in the information section of the keyword DB.
[0238] In the above embodiment, the degree of intimacy with the user may be used as the weight. In this case, for example, a person who receives a relatively large amount of voice input or a person who has a mobile terminal and has many opportunities to approach the user may be determined to be a person with a high degree of intimacy.
[0239] The configuration described in the above embodiment is merely an example. That is, at least a part of the configuration of the server 50 described in the above embodiment may be provided on the mobile terminal 10 side, or at least a part of the configuration of the mobile terminal 10 described in the above embodiment may be provided on the server 50 side. Specifically, for example, the voiceprint analysis unit 55 and the text data generation unit 54 of the server 50 may be provided on the mobile terminal 10.
[0240] In the above embodiment, the present invention has been described mainly in terms of business use, but it may also be used for private purposes, or may of course be used for both private and business purposes.
[0241] The above-described embodiment is a preferred example of the present invention, but the present invention is not limited to this and can be modified in various ways without departing from the spirit of the present invention. [Explanation of symbols]
[0242] 10...portable terminal, 12...input unit, 14...display unit, 16...playback unit, 18...warning unit, 20...biometric information input unit, 22...position detection unit, 24...time detection unit, 26...calendar unit, 28...flash memory, 30...terminal side control unit, 32...communication unit, 50...server, 52...communication unit, 53...face recognition unit, 54...text data generation unit, 55...voiceprint analysis unit, 56...weighting unit, 58...extraction unit, 60...classification unit, 62...conversion unit, 64...flash memory, 66...hard disk, 70...server side control unit, 72...change unit, 74...setting unit.< / xx>
Claims
[Claim 1] An information processing device that processes an image to be displayed on a display device, a first input unit for inputting first image data; a second input unit for inputting position information of the display device; An information processing device comprising: a first processing unit that outputs either the first image data or second image data that is the first image data with restrictions applied, based on position information of the display device input by the second input unit.
Citation Information
Patent Citations
Face collating device and face collating method
JP2004062560A