Mobile information terminal and method

The portable information terminal achieves a hands-free video call by using a wide-angle camera and image processing to detect and correct the user's face image, addressing the usability challenges of existing devices with a single flat housing.

JP7693746B2Active Publication Date: 2025-06-17MAXELL LTD
View PDF 16 Cites 0 Cited by

Patent Information

Application Number
JP2023085363
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2025-06-17
Estimated Expiration
2038-06-07

AI Technical Summary

Technical Problem

Existing mobile information terminals with a video call function face challenges in achieving a hands-free video call with good usability, especially when the device has a single substantially flat plate-shaped housing without a deformable structure.

Method used

A portable information terminal equipped with a wide-angle camera on the front surface, allowing the device to be placed flat on a surface, and using image processing to detect and trim the user's face from a wide-angle image, correct distortion, and transmit a suitable image to the other party.

Benefits of technology

Enables a hands-free video call with improved usability, allowing users to conduct video calls with both hands free without the need for special fixing devices or deformable structures, and provides a visually comfortable transmission image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693746000001
    Figure 0007693746000001
  • Figure 0007693746000002
    Figure 0007693746000002
  • Figure 0007693746000003
    Figure 0007693746000003
Patent Text Reader

Abstract

To provide a technology that can realize a hands-free videophone with more suitable usability.SOLUTION: A mobile information terminal 1 has a videophone function and is equipped with a first camera (in-camera C1) including a wide-angle lens at a predetermined position (point PC1) on a front surface s1 having a display screen DP in a flat plate-shaped housing. When a first user (user A) makes a videophone call with another second user by using the videophone function, the housing is placed flat on a first surface (horizontal plane s0) of an object and a state in which the face of the first user is included in a range of an angle of view AV1 of the in-camera C1 is a first state. The mobile information terminal 1 detects a first area including the face of the first user from a wide-angle image taken by the in-camera C1, trims a first image corresponding to the first area, and creates a transmission image to be transmitted to another terminal on the basis of the first image, and transmits the transmission image to the other terminal.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology of mobile information terminals such as smartphones, and particularly relates to the video phone function.

Background Art

[0002] In recent years, in mobile information terminals such as smartphones and tablet terminals, miniaturization such as high-density mounting on a single substantially flat housing has progressed, and various functions and advanced functions have been incorporated. These functions include a television receiving function, a digital camera function, a video phone function, etc. In particular, the video phone function in mobile information terminals has become easily available, for example, through applications and services such as Skype (registered trademark).

[0003] In addition, as a digital camera function, a mobile information terminal may have a camera (also called an in-camera) provided on the front surface of the main surface of the housing on the side where the display screen is located, or a camera (also called an out-camera) provided on the back surface opposite to the front surface.

[0004] When a user makes a video phone (sometimes referred to as a call) using a mobile information terminal, the user uses the in-camera on the front surface, which is the side that can capture the user himself / herself. The mobile information terminal displays the image of the other party received from the terminal of the video phone partner on the display screen, and transmits the image of the user himself / herself captured by the in-camera to the terminal of the other party. Therefore, usually, the user needs to hold the front surface of the mobile information terminal held in one hand, for example, in a position facing the user's face and line of sight directly, for example, in a position close to a state where the housing stands vertically. In this case, the user's hands are not free.

[0005] When a user makes a video call using a mobile information terminal, there are cases where the user wants to make the call with both hands free (sometimes referred to as hands-free). In such a case, for example, if the user places the housing of the mobile information terminal on a horizontal surface such as a desk and positions their face in a direction perpendicular to the display screen and the in-camera, a hands-free video call is possible.

[0006] As prior art examples related to the above mobile information terminal and video call function, Japanese Patent Application Laid-Open No. 2005-175777 (Patent Document 1) and Japanese Patent Application Laid-Open No. 2007-17596 (Patent Document 2) can be cited. Patent Document 1 describes, as a mobile phone, that it improves the visibility of the user when placed on a desk and that the shape of the main body can be changed to support hands-free video calls. Patent Document 2 describes, as a mobile terminal device, that based on an image captured by the camera of the main body, it acquires information about the user's face, grasps the relative positional relationship between the direction of the face and the direction of the main body, and determines the direction of the information to be displayed on the display screen.

Prior Art Documents

Patent Documents

[0007]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0008] When a user makes a video call using a mobile information terminal, in a normal state where both hands are not free (sometimes referred to as non-hands-free), there may be a lack of convenience. For example, the user cannot talk while operating a PC with both hands, and it is also difficult to talk while showing some object such as a document to the other party on the call.

[0009] In addition, when a user makes a video call using a portable information terminal and attempts to achieve a hands-free state where both hands are free, for example, the user has to assume an awkward posture to face the face towards a housing placed flat on a horizontal surface such as a desk, which is not user-friendly. Alternatively, it can be achieved by arranging the housing obliquely on a horizontal surface such as a desk using a fixing device, but this requires a fixing device and lacks convenience. Alternatively, the same state can be achieved by using a portable information terminal having a structure such as a foldable type or a separable type whose posture can be deformed, but it cannot be applied to a portable information terminal having a single substantially flat plate-shaped housing.

[0010] An object of the present invention is to provide a technology that can realize a hands-free video call with better usability for a portable information terminal having a video call function, on the premise of a configuration having a substantially flat plate-shaped housing and not having a special deformable structure.

Means for Solving the Problems

[0011] A typical embodiment of the present invention is a portable information terminal, characterized by having the following configuration. A portable information terminal according to an embodiment is a portable information terminal having a video call function, and includes a first camera including a wide-angle lens at a predetermined position on the front surface having a display screen in a flat plate-shaped housing. When a first user makes a video call with a second user using the video call function, the housing is flatly arranged on a first surface of an object, and the face of the first user is included within the range of a first viewing angle of the first camera, which is defined as a first state. In the first state, a first region including the face of the first user is detected from a wide-angle image captured by the first camera, a first image corresponding to the first region is trimmed, a transmission image for transmission to a counterpart terminal, which is the portable information terminal of the second user, is created based on the first image, and the transmission image is transmitted to the counterpart terminal.

Effects of the Invention

[0012] According to a typical embodiment of the present invention, a hands-free video phone can be realized with more suitable usability.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Embodiments for Carrying Out the Invention

[0014] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In all the drawings for explaining the embodiments, the same parts are generally denoted by the same reference numerals, and repeated explanations thereof are omitted. As directions for explanation, the X direction, Y direction, and Z direction are used. The X direction and Y direction are two orthogonal directions constituting a horizontal plane, and the Z direction is the vertical direction. The X direction is particularly the left - right direction as viewed from the user, and the Y direction is particularly the front - rear direction as viewed from the user.

[0015] [Problems, etc.] Supplementary explanation will be given about problems, etc. FIG. 26 shows, as a comparative example, an example of the usage state when attempting to realize a hands - free state using a conventional portable information terminal with a video - phone function. (A) shows the first example, and (B) shows the second example. In FIG. 26(A), user A has placed the housing of the portable information terminal 260 flat on a horizontal plane s0 such as a desk. A display screen and an in - camera 261 are provided on the front surface of the housing. The lens of the in - camera 261 is arranged at the position of point p2. The in - camera 261 is a normal camera with a normal angle of view. The representative points of user A's face and eyes are indicated by point p1. The line of sight from point p1 is in the vertically downward direction. User A is in a posture with the neck bent so that his / her face and eyes (point p1) face the display screen and the in - camera 261 on the front surface of the housing. In such a state, user A can make a video - phone call with both hands free. However, since the posture is strained, it places a burden on the body and the usability is not good.

[0016] In FIG. 26(B), user A has fixed and arranged the housing of the portable information terminal 260 on a horizontal plane s0 such as a desk using a fixing device 262 so that the front surface is in an inclined state. The direction of the line of sight from user A's face and eyes (point p1) is obliquely downward. The direction of the optical axis from the in - camera 261 (point p2) is obliquely upward (for example, an elevation angle of about 45 degrees). In such a state, user A can make a video - phone call with both hands free. However, such a state cannot be realized without using a fixing device 262 such as a stand, and it lacks versatility and convenience, such as user A having to carry the fixing device 262.

[0017] Also, if a portable information terminal with a deformable posture structure is used instead of the fixing device 262, a similar state can be realized. For example, in the technology of Patent Document 1, by changing the shape of a foldable portable information terminal, the camera and the display screen can be arranged to face the user's face. However, in the case of the portable information terminal 261 having a single substantially flat plate-shaped housing, since it does not have a structure that moves for deformation, such a technology cannot be applied.

[0018] (Embodiment 1) The portable information terminal according to Embodiment 1 of the present invention will be described with reference to FIGS. 1 to 19. In the portable information terminal according to Embodiment 1, as shown in FIG. 2 described later, when the user makes a video call using the portable information terminal (which may be simply referred to as a terminal), it is not necessary to hold the housing by hand, and it is in a hands-free state. Further, when the user wants to make a video call in a hands-free state, it is not necessary to consider and take the trouble to arrange the portable information terminal housing (front face) so that it faces the user's face. The portable information terminal according to Embodiment 1 has a configuration such as a camera in order to realize a suitable video call in a hands-free state, and defines the positional relationship between the user's face and the terminal and the arrangement state of the terminal as follows.

[0019] In the portable information terminal according to Embodiment 1, as an in-camera (first camera) provided on the front face of the housing, it has a wide-angle camera having a wide-angle field of view. This in-camera has an optical axis perpendicular to the front face of the housing, and has a predetermined wide field of view (for example, about 180 degrees (at least an angular range from 30 degrees to 150 degrees)) in a cross section formed by the front face and the optical axis (360 degrees as the horizontal field of view).

[0020] When using a video phone in a hands-free state, the user places a generally flat-shaped housing flat on a substantially horizontal plane such as a desk, with the front in-camera facing upward. And the user's face is in a positional relationship in the direction of an upward diagonal elevation angle as viewed from the position of the in-camera of the terminal. From the user's eyes, the display screen and the in-camera of the housing are in a positional relationship in the direction of looking downward diagonally. In this state (the first state), the direction of face shooting from the in-camera and the direction of viewing the display screen from the user's eyes are generally the same or sufficiently close. For example, these directions are at an elevation angle of about 45 degrees with respect to the horizontal plane (for example, an angle within the angle range from 30 degrees to 60 degrees).

[0021] In such a positional relationship and arrangement state, in the image (wide-angle image) of the in-camera, the user's face is reflected in a partial area (for example, an area within the angle range from 0 degrees to 90 degrees). The portable information terminal can capture the area including the user's face using the wide-angle image. The portable information terminal detects and trims (cuts out) the area including the user's face from the wide-angle image.

[0022] However, in the wide-angle image through the wide-angle lens, distortion (deformation) peculiar to the wide-angle lens occurs in the entire wide-angle image including the user's face. When the wide-angle image is transmitted to the other party's portable information terminal, the other party's user may have difficulty recognizing the user's face and the like from the wide-angle image. Therefore, the portable information terminal of Embodiment 1 performs distortion correction processing on the area including the face of the wide-angle image to obtain a flattened image in which the distortion is eliminated or reduced. Thereby, a visually more comfortable and suitable transmission image is obtained.

[0023] The portable information terminal 1 creates a transmission image for transmitting to the terminal of the other party in the video phone from the image including the face of User A after distortion correction, and transmits it to the terminal of the other party. As described above, in the portable information terminal of Embodiment 1, the user can use the video phone in a hands-free state, and the usability is good.

[0024] [(1) Communication system and video phone system] FIG. 1 shows the overall configuration of a communication system and a videophone system including the mobile information terminal 1 of Embodiment 1. In the communication system and the videophone system of FIG. 1, the mobile information terminal 1 of the first user (User A) and the mobile information terminal 2 of the second user (User B) are connected via the mobile communication network 101 and the Internet 102. Videophone communication is performed between the mobile information terminal 1 of the first user and the mobile information terminal 2 of the second user. The base station 103 is a radio base station constituting the mobile communication network 101. The mobile information terminals 1 and 2 are connected to the mobile communication network 101 through the base station 103. The access point 104 is an access point device for wireless communication in a wireless LAN or the like. The mobile information terminals 1 and 2 are connected to the Internet 102 including a wireless LAN or the like through the access point 104.

[0025] The mobile information terminals 1 and 2 are devices such as smartphones, and both are equipped with a videophone function. The mobile information terminal 1 of the first user is the mobile information terminal of Embodiment 1 and has a specific function related to videophone. User A is the first user who is one of the callers in a videophone call, and User B is the other caller in the videophone call and the second user who is the other party as seen from User A. Hereinafter, the mobile information terminal 1 of the first user will be mainly described. User B uses, for example, a conventional mobile information terminal 2 with a videophone function. The mobile information terminal 2 of the second user may also have a specific videophone function similar to that of the mobile information terminal 1 of the first user. Note that, during videophone communication, a server or the like that provides services related to the videophone function may be interposed on the mobile communication network 101 or the Internet 102.

[0026] [(2) Outline of Videophone Use] FIG. 2 is a schematic diagram showing an outline, situation, and method of using the mobile information terminal 1 by User A during a videophone call between User A and User B in FIG. 1. FIG. 2 shows the positional relationship between the face of User A and the mobile information terminal 1 and the arrangement state of the terminal when User A makes a hands-free videophone call. The outline of videophone use is as follows.

[0027] (1) As shown in FIG. 2, when user A makes a hands-free video call, the user places the flat housing of the portable information terminal 1 flat on the horizontal plane s0 (X-Y plane, the first plane) of an arbitrary object such as a desk, with the in-camera C1 on the front surface s1 facing vertically upward. The back surface s2 of the housing is in contact with and hidden by the horizontal plane s0. On the front surface s1 of the vertically long housing of the portable information terminal 1, mainly a rectangular display screen DP is provided, and a camera, operation buttons, a microphone, a speaker, etc. are provided in the frame area on the outer periphery of the display screen DP. In this portable information terminal 1, the in-camera C1 (particularly the wide-angle lens part) is arranged at the position PC1 on the upper side of the frame area. When placing it, user A places the housing such that the in-camera C1 (position PC1) on the front surface s1 is at a position in the far-back direction Y1 in the Y direction as seen from user A.

[0028] User A places his / her face and eyes at a diagonally upward position with respect to the portable information terminal 1. In other words, as seen from the eyes (point P1) of user A, the display screen DP (point PD) of the portable information terminal 1 is arranged at a diagonally downward position. Let the representative point associated with the head, face, or eyes of user A be point P1. An image of the other party (user B) etc. is displayed on the display screen DP (FIG. 7). User A looks at the face image etc. of the other party within the display screen DP. The direction J1 indicates the line-of-sight direction from the eyes (point P1) of user A to the display screen DP (point PD) of the terminal. The angle θ1 indicates the elevation angle corresponding to the direction J1 (with the horizontal plane s0 as the reference of 0 degrees). The angle θ1 is an angle within the range from 30 degrees to 60 degrees, for example, about 45 degrees.

[0029] In this state, naturally, the face and eyes of user A can be photographed by the in-camera C1 of the terminal in the diagonally upward elevation angle direction. The optical axis of the in-camera C1 is in the direction perpendicular to the front surface s1 (vertically upward), which is shown as the direction J2. The angular field of view AV1 of the in-camera C1 has a wide angular range centered on the optical axis, with a horizontal angular field of view of 360 degrees and an angular field of view in the Y-Z cross-section of about 180 degrees, particularly having an angular range from the first angle ANG1 to the second angle ANG2. For example, the first angle ANG1 is 30 degrees or less, and the second angle ANG2 is 150 degrees or more.

[0030] In this state, the in-camera C1 can capture the face of user A with a wide-angle field of view AV1. That is, in this state, within the field of view AV1 of the in-camera C1, there is included a field of view AV2 (face-capturing field of view) corresponding to the range for capturing the face of user A, particularly within the angular range from the first angle ANG1 to 90 degrees. Correspondingly, the face of user A appears in a partial region within the wide-angle image of the in-camera C1. The direction from the in-camera C1 (point PC1) to the face of user A (point P1) is indicated by direction J3. The elevation angle corresponding to direction J3 is indicated by angle θ3. Angle θ3 is, for example, an angle slightly smaller than 45 degrees.

[0031] The direction J1 of the line of sight of user A and the face-capturing direction J3 from the in-camera C1 are in a state of being sufficiently close, and the angular difference AD1 between angle θ1 and angle θ3 is sufficiently small. Therefore, the in-camera C1 can capture the face of user A in the direction J3 and the field of view AV2 where the line of sight of user A can be confirmed. It is more preferable that these direction differences and angular differences are as small as possible because the state of the line of sight in the image becomes more natural. In the portable information terminal 1 of the first embodiment, since the wide-angle image of the in-camera C1 is used, even in the positional relationship as shown in FIG. 2, the face of user A can be captured like the field of view AV2.

[0032] (2) In the state of FIG. 2, user A can make a video call while looking at the image of the other party (FIG. 7) displayed on the display screen DP of the portable information terminal 1. The portable information terminal 1 outputs the voice received from the portable information terminal 2 of the other party from the speaker. The portable information terminal 1 transmits the voice of user A input by the microphone to the portable information terminal 2 of the other party.

[0033] The portable information terminal 1 detects and trims the region including the face of user A from the wide-angle image by the in-camera C1. The portable information terminal 1 creates a transmission image for transmission to the portable information terminal 2 of the other party using the trimmed image. However, in the wide-angle image captured by the in-camera C1, there is distortion depending on the wide-angle lens, including the face of user A.

[0034] Therefore, the mobile information terminal 1 performs distortion correction processing on the trimmed image so that the distortion is eliminated or reduced. The mobile information terminal 1 creates a transmission image for transmission to the other party's mobile information terminal 2 and a monitor image for confirming the state of the user A's own face or the like corresponding to the transmission image from the image after distortion correction. The mobile information terminal 1 displays the monitor image corresponding to the transmission image within the display screen DP (Fig. 7). In the monitor image (and the transmission image), the direction of the user A's line of sight is generally facing forward. The user A can view and confirm the image of the other party (user B) and the monitor image of the user A himself / herself within the display screen DP of the mobile information terminal 1. The user A can also reject the transmission of the transmission image corresponding to the monitor image if necessary.

[0035] The mobile information terminal 1 transmits data for video phone including the voice of the user A input by the microphone and the above transmission image to the other party's mobile information terminal 2. The other party's mobile information terminal 2 outputs an image and voice related to the user A based on the received data, and the user B can make a video call with the user A.

[0036] [(3) Mobile Information Terminal] Fig. 3 shows the configuration of the mobile information terminal 1 of Embodiment 1. The mobile information terminal 1 includes a controller 10, a camera unit 11, a ROM 14, a RAM 15, an external memory 16, a microphone 17, a speaker 18, a bus 19, a display unit (touch panel) 21, a LAN communication unit 22, a mobile network communication unit 23, sensors 30 such as an acceleration sensor 31 and a gyro sensor 32, and the like.

[0037] The controller 10 controls the entire mobile information terminal 1 and gives instructions to each part. The controller 10 realizes a video phone function 100 based on a video phone application. The controller 10 of the mobile information terminal 1 controls the video phone function 100 using each part and each function. The controller 10 is composed of a microprocessor unit (MPU) or the like and controls the entire mobile information terminal 1 according to the program in the ROM 14. Each part such as the controller 10 performs data transmission and reception with each part within the mobile information terminal 1 via the bus 19 (including the system bus).

[0038] The camera unit 11 includes an in-camera C1, a normal camera (out-camera) C2, a photographing processing unit 12, and a memory 13. As shown in FIG. 2 described above, the in-camera C1 is provided on the front surface s1 of the housing, and the normal camera C2 is provided on the back surface s2 of the housing. The in-camera C1 is composed of known elements such as a wide-angle lens, a camera sensor (imaging device), and a drive circuit. The camera sensor is composed of, for example, a CCD or a CMOS sensor. The normal camera C2 is composed of known elements such as a normal lens, a camera sensor, and a drive circuit. The normal camera C2 has a normal angle of view, and the normal angle of view is narrower than the angle of view AV1 of the wide-angle lens of the in-camera C1. The optical axis of the normal camera C2 is in the opposite direction to the optical axis of the in-camera C1. In the videophone function in Embodiment 1, the normal camera C2 is not used.

[0039] The photographing processing unit 12 is a part that performs photographing processing and image processing using a camera based on the control of the controller 10. In Embodiment 1, the photographing processing unit 12 is implemented as a circuit such as an LSI as a part separate from the controller 10. However, the photographing processing unit 12 may be integrally implemented by the program processing of the controller 10 or the like, either partially or entirely. Functions such as the face detection function 201 may be implemented entirely or partially by software program processing, or may be implemented by a hardware circuit or the like for speed improvement. The memory 13 is a memory that stores image data and the like related to photographing processing.

[0040] As known functions and processing units, the photographing processing unit 12 includes an autofocus function, a zoom function, a codec unit, an image quality improvement processing unit, an angle / rotation correction unit, and the like. The autofocus function is a function that automatically adjusts the focus of the camera to the object to be photographed. The zoom function is a function that enlarges or reduces the object in the image. The codec unit is a processing unit that compresses and decompresses the photographed image and video. The image quality improvement processing unit is a processing unit that improves the image quality of the photographed image, such as noise removal. The angle / rotation correction unit is a processing unit that performs angle correction and rotation correction on the photographed image.

[0041] The microphone 17 collects the ambient sound around the terminal, including the voice of User A, to obtain voice data. The speaker 18 outputs the voice including the video call voice from the mobile information terminal 2 of the call partner (User B).

[0042] The display unit 21 includes the display screen DP in FIG. 2, and is particularly a touch panel such as a liquid crystal touch panel, and allows touch input operations by the user. The shooting image and various other information are displayed on the display screen DP.

[0043] The LAN communication unit 22 performs communication processing corresponding to communication on the Internet 102, including wireless LAN communication with the access point 104 in FIG. 1. The mobile network communication unit 23 performs communication processing corresponding to communication on the mobile network 101, including wireless communication with the base station 103 in FIG. 1.

[0044] The sensors 30 include known sensor devices such as an acceleration sensor 31, a gyro sensor 32, a GPS receiver (not shown), a proximity sensor, an illuminance sensor, and a temperature sensor. The controller 10 detects the orientation and movement of the mobile information terminal 1 using the detection information of the sensors 30. The controller 10 can also grasp states such as whether the mobile information terminal 1 is held by User A or whether it is flatly placed on the horizontal plane s0 as shown in FIG. 2 using the sensors 30.

[0045] The shooting processing unit 12 has a face detection function 201, a trimming function 202, and a distortion correction function 203 as functions realized by a program, a circuit, or the like. The outline of the shooting processing and functions of the shooting processing unit 12 is as follows. In the camera mode where the in-camera C1 is used for shooting, the shooting processing unit 12 inputs a wide-angle image that is an image captured by the in-camera C1. In the first embodiment, the in-camera C1 can capture a moving image (a plurality of image frames in time series), and the shooting processing unit 11 processes the moving image. However, it is not limited to this, and a still image at a predetermined timing may be treated as the image of the in-camera C1.

[0046] The imaging processing unit 12 detects, by means of the face detection function 201, a region including the face of user A from the wide-angle image. The imaging processing unit 12 trims, by means of the trimming function 202, a region including the face of user A from the wide-angle image. The imaging processing unit 12 performs a distortion correction process on the trimmed image by means of the distortion correction function 203. The imaging processing unit 12 stores the image after distortion correction in the memory 13. The imaging processing unit 12 (or the controller 10) creates a transmission image for transmission to the other party's portable information terminal 2 and a monitor image for self-confirmation from the corrected image.

[0047] The controller 10 (video phone function 100) of the portable information terminal 1 creates video phone data by combining the transmission image of user A with the voice data of user A input from the microphone 17. The controller 10 transmits the data to the other party's portable information terminal 2 using the LAN communication unit 22 or the mobile network communication unit 23 or the like. The other party's portable information terminal 2 receives the data, displays the transmission image on the display screen, and outputs the voice.

[0048] The controller 10 (video phone function 100) receives video phone data (including the other party's image and voice) from the other party's portable information terminal 2 through the LAN communication unit 22 or the mobile network communication unit 23. The controller 10 displays the other party's image among the received data within the display screen DP and also displays the monitor image of user A. The controller 10 outputs the other party's voice from the speaker 18.

[0049] [(4) Software Configuration] FIG. 4 shows the software configuration of the portable information terminal 1. In the ROM 14, basic operation programs 14a such as an OS and middleware, and other application programs and the like are stored. As the ROM 14, a rewritable ROM such as an EEPROM or a flash ROM is used, for example. Through communication or the like, the programs in the ROM 14 can be updated as appropriate, and version upgrades and function expansions are possible. The ROM 14 or the like may be integrated with the controller 10.

[0050] The RAM 15 is used as a work area during the execution of the basic operation program 14a, the videophone application program 16b, and the like. The RAM 15 also has a temporary storage area 15c for temporarily holding data and information as needed during the execution of various programs. The controller 10 (MPU) expands the basic operation program 14a in the ROM 14 into the RAM 15 and executes processing according to the program. As a result, a basic operation execution unit 15a is configured in the RAM 15. Similarly, a videophone processing execution unit 15b is configured in the RAM 15 in accordance with the processing of the videophone application program 16b. Data for processing related to the videophone function 100 is stored in the temporary storage area 15c, and information such as the position and posture state of the portable information terminal 1 is also stored.

[0051] Programs such as the shooting program 16a and the videophone application program 16b are stored in the external memory 16, and the external memory 16 also has a data storage area 16c for accumulating images captured by the shooting processing unit 12 and data and information related to various processes. The external memory 16 is composed of a non-volatile memory device that retains data even in a power-off state, and for example, a flash ROM, an SSD, or the like is used. In the data storage area 16c, for example, setting values for the functions and operations of the portable information terminal 1 are also stored. The various programs may be stored in the ROM 14 or other non-volatile memory devices. The portable information terminal 1 may acquire programs and information from an external server device or the like.

[0052] The shooting program 16a realizes shooting control processing for the shooting processing unit 12 of the camera unit 11. This shooting control processing includes general camera shooting control processing not limited to videophones and shooting control processing for videophones. The shooting program 16a is expanded into the RAM 15 or within the shooting processing unit 12 to form an execution unit.

[0053] The controller 10 (MPU) instructs the shooting processing unit 12 regarding the camera mode, start and end of video shooting by the camera, detailed shooting settings (for example, focus, exposure), and the like. The camera mode indicates which of the plurality of cameras is used for shooting.

[0054] The video phone processing execution unit 15b based on the video phone application program 16b performs processing corresponding to the video phone function 100. When realizing the video phone function 100, the controller 10 (MPU) performs control processing for each function of the imaging processing unit 12 and control processing for each related unit.

[0055] [(5) Camera unit, imaging processing unit] FIG. 5 shows the detailed configurations of the camera unit 11, the imaging processing unit 12, and the memory 13. The face detection function 201 of the imaging processing unit 12 includes the personal recognition function 201B. As processing, the distortion correction function 203 performs orthographic conversion processing 203A, trapezoid correction processing 203B, aberration correction processing 203C, etc. Data such as the registered image D10, the corrected image D11, the transmitted image D12, and the monitor image D13 are stored in the memory 13. The memory 13 temporarily holds the captured image data and is also used as a work area related to the processing of each function. The memory 13 may be inside the imaging processing unit 12. The outline of the processing of the imaging processing unit 12 is as follows.

[0056] (1) First, the wide-angle image (data D1) captured through the in-camera C1 is input to the face detection function 201 of the imaging processing unit 12. The face detection function 201 performs processing to detect the area including the face of user A from the data D1 of the wide-angle image based on image processing. The face detection function 201 extracts, for example, a feature point group from within the wide-angle image, detects eyes, ears, nose, mouth, etc., and also detects the face and head contours based on the difference in pixel color and luminance. Thereby, as shown in FIG. 9 described later, the face area B1 etc. can be detected. From the face detection function 201, the data D2 of the wide-angle image and the detection result are output.

[0057] Also, in the personal recognition function 201B, it is recognized whether the face image is the face of a specific user A. The imaging processing unit 12 applies the subsequent processing only when, for example, the face of a specific user A is detected.

[0058] (2) Next, the trimming function 202 performs a process of trimming a trimming area corresponding to the detected face-containing area based on the data D2 of the wide-angle image to obtain a trimmed image. Data D3 such as the trimmed image is output from the trimming function 202. As a trimming method, for example, an area having a predetermined shape and size is trimmed based on the center point (point P1) in the detected face area. Note that the trimming area may be only the face area, may be the head area, or may be an area including the head and its surroundings. The type and size of the trimming area can be changed using the user setting function.

[0059] (3) Next, the distortion correction function 203 performs a process of correcting the distortion due to the wide-angle lens in the trimmed image to a plane having a non-distorted orthographic image based on the data D3. The distortion correction function 203 first performs an orthographic conversion process 203A (FIG. 14) on the trimmed image. As a result, an image (flattened image) that is a plane having a non-distorted orthographic image is obtained, and data D4 including the flattened image is output.

[0060] Next, the distortion correction function 203 performs a trapezoid correction process 203B (FIG. 15) on the data D4 of the flattened image. As a result, an image in which the trapezoidal image content is made into a rectangular image content is obtained, and its data D5 is output. By the trapezoidal conversion, the image content is made to have a more suitable appearance.

[0061] Next, in the aberration correction process 203C, the distortion correction function 203 performs a known process of correcting various aberrations caused by the characteristics of the lens system other than the wide-angle distortion on the data D5. As a result, a corrected image D11 is obtained. For example, when the lens system is fixed, lens system correction parameters D14 are stored in, for example, the memory 13 at the time of product shipment in advance. The lens system correction parameters D14 are setting information and initial values for aberration correction, and may be settable. The aberration correction process 203C refers to the lens system correction parameters D14.

[0062] Note that, normally, in the state of the image after the positive image conversion process 203A, the state of the face of user A is visually not uncomfortable enough (at least in a state where it can be used for a video phone). Therefore, the trapezoidal conversion process 203B and the aberration correction process 203C may be omitted. Also, the processes performed by the distortion correction function 203 do not necessarily have to be performed in this order, and the processes may be performed in any order. Also, depending on the conditions, control may be performed so as not to perform a specific process.

[0063] (4) The corrected image D11 by the distortion correction function 203 is stored in the memory 13. In the corrected image D11, the distortion caused by the wide-angle image and the aberration of the lens are eliminated or reduced, and the image is at a level where the user can recognize the face etc. with little discomfort. The shooting processing unit 12 or the controller 10 (video phone function 100) creates a transmission image D12 for transmitting to the other party (user B) and a monitor image D13 for self-confirmation using the corrected image D11. The controller 10 creates data for the video phone using the transmission image D12 and the voice input by the microphone 17.

[0064] The transmission image D12 is an image to which cuts, enlargements / reductions, etc. are appropriately applied so as to match, for example, the image size (display screen size, etc.) requested by the other party's portable information terminal 2. The monitor image D13 is an image to which cuts, enlargements / reductions, etc. are appropriately applied so as to match the size of the monitor image display area (area R2 in FIG. 7) within the display screen DP.

[0065] (5) The controller 10 displays the monitor image D13 in an area within the display screen DP. When the result of confirmation by user A for the monitor image D13 is affirmative (transmission permission), the controller 10 transmits the data including the transmission image D12 to the other party's portable information terminal 2 via the communication unit.

[0066] (6) Further, when the photographing processing unit 12 cannot detect the face of user A by the face detection function 201, or when receiving an instruction to reject transmission after user A has confirmed the monitor image D13, etc., the transmission image D12 is created using the registered image D10. The registered image D10 includes the face image of user A.

[0067] The photographing processing unit 12 similarly repeats the above processing for each photographed image at a predetermined time interval. At that time, when the face of user A cannot be captured from the image at a certain point in time, etc., an alternative transmission image D12 may be created using the image detected last in the past or the face image of the registered image D10.

[0068] [(6) Processing flow] FIG. 6 shows the processing flow of the video phone function 100 in the mobile information terminal 1. The flow of FIG. 6 has steps S1 to S13. Hereinafter, it will be described in the order of the steps.

[0069] (S1) First, in S1, when user A makes a video call (when making a call from oneself to the other party or when receiving a call from the other party to oneself), the controller 10 (video phone function 100) of the mobile information terminal 1 shifts the control state of the own device to the video phone mode. Specifically, for example, when user A wants to make a video call with the other party (user B), user A dials the phone number of the other party. Along with this, the video phone application program 16b (video phone processing execution unit 15b) is activated. The controller 10 controls the photographing of the in-camera C1, the voice input of the microphone 17, the voice output of the speaker 18, the display of the display unit 21, various communications, etc. simultaneously and in parallel in the video phone mode.

[0070] Also, the mobile information terminal 1 shifts to the video phone mode in response to the selection operation of a voice call (non-video call) or a video call by user A. For example, the mobile information terminal 1 displays selection buttons for a voice call or a video call on the display screen DP, and shifts to the corresponding mode according to the selection operation. Also, since the in-camera C1 is used in the video phone mode, the mobile information terminal 1 sets the camera mode of the camera unit 11 to the mode using the in-camera C1.

[0071] Furthermore, in Embodiment 1, as details of the videophone mode, two types are provided: a normal mode (non-hands-free mode) and a hands-free mode. The normal mode (non-hands-free mode) is the first mode corresponding to the state as shown in FIG. 19. The hands-free mode is the second mode corresponding to the state as shown in FIG. 2. The portable information terminal 1 selects from these modes according to a predetermined instruction operation by User A or an automatic determination of the terminal state using sensors 30. For example, the portable information terminal 1 may display selection buttons for the normal mode and the hands-free mode on the display screen DP and shift to the corresponding mode according to the selection operation. Alternatively, the portable information terminal 1 may grasp the state of whether User A is holding the housing or placing it flat on the horizontal plane s0 from the detection information of the acceleration sensor 31 or the like, and automatically determine the mode according to the state.

[0072] In this example, User A makes a videophone call in the hands-free state (corresponding hands-free mode) as shown in FIG. 2. User A sets the housing of the portable information terminal 1 to the state as shown in FIG. 2, and the portable information terminal 1 selects the hands-free mode. Note that in other embodiments, it may not be necessary to distinguish between the above two types of modes.

[0073] (S2) The portable information terminal 1, together with grasping the terminal state, sets the camera unit 11 to the mode using the in-camera C1 (in-camera mode) in the videophone mode and starts shooting. The shooting processing unit 12 inputs the moving image from the in-camera C1.

[0074] (S3) The shooting processing unit 12 of the portable information terminal 1 detects a region including the face of User A (for example, region B1 in FIG. 9) from the wide-angle image of the in-camera C1 by the face detection function 201.

[0075] (S4) The shooting processing unit 12 trims a predetermined region including the face as a trimming region (for example, trimming region TRM1 in FIG. 10) for the region detected in S3 by the trimming function 202 to obtain a trimmed image (for example, image GT1 in FIG. 12).

[0076] (S5) The image processing unit 12 performs distortion correction processing (normal image conversion processing 203A) on the trimmed image obtained in S4 by the distortion correction function 203. The distortion correction function 203 also performs the keystone correction processing 203B and aberration correction processing 203C described above. As a result, a corrected image D11 (for example, image GP1 in FIG. 12) is obtained.

[0077] (S6) The image capturing and processing unit 12 (or the controller 10) creates a transmission image D12 and a monitor image D13 (for example, images GP11 and GP12 in FIG. 12) using the image D11 corrected in S5.

[0078] (S7) The controller 10 displays the image received from the mobile information terminal 2 of the other party (user B) in an area (area R1 in FIG. 7) on the display screen DP. The controller 10 also displays the monitor image D13 of user A in an area (area R2 in FIG. 7) on the display screen DP.

[0079] (S8) The controller 10 asks the user A whether the face condition in the monitor image D13 is acceptable for the corresponding transmission image D12 (sometimes referred to as transmission confirmation). For example, transmission confirmation information (e.g., "May I send the image?") or operation buttons (e.g., a transmission permission button, a transmission refusal button) may be displayed in the display screen DP. The user A looks at the monitor image D13, etc., and judges whether the image content is acceptable for transmission. For example, the user A presses the transmission permission button or the transmission refusal button in the display screen DP. If the transmission is acceptable (Y) in S8, proceed to S10, and if the transmission is refusal (N), proceed to S9. If the user A looks at the monitor image D13 and feels something is wrong with the condition of the eyes, for example, he or she can select transmission refusal.

[0080] (S9) The controller 10 creates an alternative transmission image D12 using the registered image D10. At this time, the controller 10 may also confirm with the user A on the display screen DP whether the registered image D10 may be used as the alternative transmission image D12. For example, the face image of the registered image D10 is displayed in the area R2 for monitor image display in FIG. 7, and confirmation information and operation buttons are displayed.

[0081] (S10) The controller 10 transmits data in the form of a videophone including the transmission image D12, etc. to the other party's mobile information terminal 2 via the communication unit.

[0082] (S11) The controller 10 processes a videophone call (including voice input and output, image display, etc.) between the user A and the user B. Regarding the voice data of the callers, it may be transmitted and received constantly separately from the image as in the case of a normal phone.

[0083] (S12) The controller 10 checks whether the videophone has ended. For example, when the user A ends the videophone, the user A presses the end button. Alternatively, the mobile information terminal 1 receives information indicating the end from the other party's mobile information terminal 2. In the case of ending the videophone (Y), the process proceeds to S13, and in the case of continuation (N), it returns to S2. From S2, the processing for each time point is similarly repeated in a loop. Note that the face of the user A is automatically tracked by the loop. In each process in the loop, the processing is optimized so as not to repeat the same process as much as possible. For example, in the face detection process of S3, for the face area once detected at a certain time point, the face is automatically tracked by motion detection, etc. at a later time point.

[0084] (S13) The controller 10 performs an end process for the videophone mode and ends the startup (execution) of the videophone application. The end process includes resetting settings related to the videophone (for example, the number of retry times), deleting image data, etc.

[0085] Supplements and variations of the above processing flow are as follows. In the above processing flow, basically, the same processing is performed in a loop for each image of the video of the in-camera C1. Regarding the transmission confirmation using the monitor image D13 in S8, for example, it is performed once at the start of a video call, that is, using an image in the first period of the video. If transmission is permitted in that transmission confirmation, during the subsequent video call, the transmission image D12 created at each time point is automatically transmitted. Not limited to this, the transmission confirmation may be performed at regular intervals during the video call, may be performed in response to a predetermined user operation, or may not be performed at all. In the user setting function of the video call application, settings regarding the presence or absence, timing, etc. of the above transmission confirmation are possible. When the setting is not to perform the transmission confirmation, the creation of the monitor image D13 in S6, the display of the monitor image D13 in S7, the transmission confirmation in S8, etc. can be omitted, and the mobile information terminal 1 automatically transmits the transmission image D12 with transmission permitted.

[0086] Also, when transmission is permitted by transmission confirmation at the beginning of a video call, thereafter, until the end of the video call or until the next transmission confirmation, the same one as the transmission image D12 created first may be continuously used.

[0087] In the face detection process in S3, for example, when the face area cannot be detected from a wide-angle image at a certain time point, the face detection process may be retried using an image at another time point according to the preset number of retries. Also, when the face area cannot be detected or tracking cannot be performed, the face image detected last in the past or the registered image D10 may be used as an alternative. Further, the mobile information terminal 1 may display a message indicating that the face cannot be detected, etc. within the display screen DP to the user A, and confirm whether to use the registered image D10 as the transmission image D12 instead and respond accordingly.

[0088] Also, when user A performs an instruction operation to reject transmission during the transmission confirmation of S8, the mobile information terminal 1 may immediately use the registered image of S9, but it is not limited to this. For example, it may return to steps such as S5 or S3 and retry the process up to a predetermined number of retry times. The settings such as the number of retry times may be the default settings of the video phone application or may be changeable by user settings. According to the user's needs and operations, using the user setting function of the video phone application, it is possible to set the availability and operation details of various functions including transmission confirmation and creation of transmission images using registered images. Depending on the user settings, it is possible to use real-time camera images from the beginning to the end of the video phone, or only use the registered image D10.

[0089] [(7) Mobile Information Terminal - Display Screen] FIG. 7 shows the configuration of a display screen DP and the like in the X-Y plane when the front surface s1 of the mobile information terminal 1 is viewed in plan view during a video phone call. Among the main surfaces (front surface s1, back surface s2) of the flat housing of the mobile information terminal 1, an in-camera C1 including a wide-angle lens unit is provided on the front surface s1 side having the display screen DP. In the vertically long rectangular area of the front surface s1, the in-camera C1 is provided at, for example, the central position (point PC1) of the upper side portion in the frame area outside the area of the main display screen DP.

[0090] Based on the control of the video phone application and the image data received from the mobile information terminal 2 of user B, an image (opponent image) g1 including the face of the call opponent (user B) is displayed in the area R1 within the display screen DP of the mobile information terminal 1.

[0091] Along with the display of the image g1 in the area R1 within the display screen DP, in some predetermined areas R2, a monitor image g2 including the face of user A, created based on the image of the in-camera C1, is displayed. This monitor image g2 is provided so that user A can check the state of their own image to be sent to the other party during a video call. The function of displaying this monitor image g2 is not essential, but when it is displayed, it can enhance usability. User A can check the state of their face etc. by looking at this monitor image g2 and, if necessary, can also reject the transmission of the transmission image D12 corresponding to this monitor image g2.

[0092] In the display example of Fig. 7(A), within the display screen DP, the image g1 of the other party is displayed in the main area R1 corresponding almost entirely. In the area R1, for example, at the position of the upper right corner close to the in-camera C1, an overlapping area R2 is provided, and the monitor image g2 is displayed in that area R2.

[0093] In another display example of Fig. 7(B), within the display screen DA, the main area R1 for displaying the image g1 of the other party is arranged at a position closer to the upper side near the in-camera C1. Below the area R1, a separated area R2 is provided, and the monitor image g2 is displayed in that area R2. It is not limited to these, and various display methods are possible and can be changed by the user setting function.

[0094] Also, in the display example of Fig. 7, the monitor image g2 in the area R2 is smaller in size than the image g1 of the other party in the area R1. It is not limited to this, and the size of the monitor image g2 in the area R2 can also be set and changed. Also, it is possible to perform, for example, an enlarged display of only the monitor image g2 in response to a touch operation on the display screen DP.

[0095] [(8) Image Examples, Processing Examples] Figs. 8 to 12 show image examples and processing examples regarding face detection, trimming, and distortion correction based on the wide-angle image of the in-camera C1 of the portable information terminal 1.

[0096] (1) First, FIG. 8 shows, for comparative explanation purposes, an example of an image (normal image) obtained by photographing the face of user A from the front with a normal camera C2. This normal image shows the case when it has a square size. Let point P1 be the center point or representative point on the face or head of user A. The midpoint between both eyes or the like may be used as point P1. Schematically, the face region A1, the head region A2, and the region A3 are shown. The face region A1 is a region taken to include the face (including eyes, nose, mouth, ears, skin, etc.). The head region A2 is a region that is wider than the face region A1 and is taken to include the head (including hair, etc.). The region A3 is wider than the face region A1 and the head region A2, and is a region taken to include to a certain extent the peripheral region outside the face or head. The region A3 may be, for example, a region up to a predetermined distance from point P1, or may be a region taken according to the ratio of the face region A1 or the like within the region A3. The shape of each region is not limited to a rectangle, and may be an ellipse or the like. Note that in the case of an image photographed in the positional relationship of looking up diagonally upward from the camera as shown in FIG. 2, the actual image content becomes a somewhat trapezoidal image content as shown in FIG. 15(B).

[0097] (2) FIG. 9 schematically shows an overview of the wide-angle image G1 captured by the in-camera C1 and face region detection. This wide-angle image G1 has a circular region. Point PG1 indicates the center point of the wide-angle image G1 and corresponds to the direction J2 of the optical axis. The position coordinates within the wide-angle image G1 are indicated by (x, y). The region (face region) B1 shown by the dashed line frame indicates a rectangular region that generally corresponds to the face region A1. Similarly, the region (head region) B2 indicates a rectangular region corresponding to the head region A2, and the region B3 indicates a rectangular region corresponding to the region A3. The region B4 further shows the case where it is larger than the region B3 and takes a rectangular region that is sufficiently large for processing. Although each region is shown as a rectangle, it is not limited to this, and it may have a shape adapted to the coordinate system of the wide-angle image.

[0098] In the wide-angle image G1 of FIG. 9, based on the positional relationship as shown in FIG. 2, in a partial region within the wide-angle image G1, particularly in the region near the lower position (point P1) from the central point PG1, the face of user A or the like is captured. Thus, distortion depending on the wide-angle lens occurs in the whole including the face within the wide-angle image G1 captured through the wide-angle lens of the in-camera C1. In the wide-angle image G1, the distortion may be greater at the outer peripheral position than at the center (point PG1).

[0099] The face detection function 201 detects a region including a face from within the wide-angle image G1. For example, the face detection function 201 detects a face region B1 or a head region B2.

[0100] (3) FIG. 10 shows, for the wide-angle image G1 of FIG. 9, a state in the (x, y) plane of the original image of FIGS. 13 and 14 described later superimposed thereon. The trimming function 202 sets a trimming region corresponding to the face region B1, the region B3, etc. from the wide-angle image G1 to obtain a trimmed image. In the example of FIG. 10, a shield-shaped region corresponding to the region B3 in the (x, y) plane coordinate system of the original image is set as the trimming region TRM1 (broken line frame) from the wide-angle image G1. The shield shape of this trimming region TRM1 is, for example, a shape in which the width in the x direction decreases as it goes from the center to the outer periphery in the y direction.

[0101] (4) Further, FIG. 11 shows, for comparison and explanation, an example of a trimmed image when trimming from the wide-angle image G1 of FIG. 9 as a rectangular (right-angled quadrilateral) trimming region. Region 111 shows a rectangular region including the head and its surroundings. Region 112 shows an example of a rectangular region when taking a larger size than region 111 for processing. In region 111, generally, the horizontal width H1 of the face region or the like and the overall horizontal width H2 are shown. Regarding the determination of the size of region 111, the overall width H2 is set so as to have a predetermined ratio (H1 / H2) with respect to the horizontal width H1 of the face region or the like. For example, H1 / H2 = 1 / 2, or 2 / 3, etc. This ratio can be set by the user. Alternatively, regarding the sizes of regions 111 and 112, they may be determined by taking predetermined distances K1, K2, etc. in the horizontal and vertical directions from the central point P1 of the face region.

[0102] (5) FIG. 12 shows an example of a trimming image and distortion correction. (A) of FIG. 12 shows a trimming image GT1 corresponding to the trimming area TRM1 of FIG. 10. Further, the trimming area TRM2 and the trimming image GT2 show an example when they are taken larger than the trimming area TRM1 and the trimming image GT1 for processing. On the (x, y) plane of the original image, the trimming area TRM1 has a size (for example, a size including the upper body) that includes the periphery of the face of user A.

[0103] (B) of FIG. 12 shows a flattened image GP1 which is an image of the result of distortion correction for the trimming image GT1 of (A). In the distortion correction function 203, an orthographic conversion process 203A (FIG. 14) is performed on the trimming image GT1. As a result, a flattened image GP1 with almost no distortion is obtained. The flattened image GP1 has a rectangular (right-angled quadrilateral) plane PL1. Further, the plane PL2 indicates the plane of the flattened image obtained in the same manner for the trimming image GT2. In particular, the images GP11 and GP12 show examples when extracting a partial area of the flattened image GP1. The image GP11 corresponds to the face area, and the image GP12 corresponds to an area including the head area and its periphery.

[0104] The imaging processing unit 12 acquires the flattened image GP1 as a corrected image D11. Further, the imaging processing unit 12 may create a monitor image D13 by extracting some images such as GP11 from this flattened image GP1 and appropriately processing them.

[0105] The imaging processing unit 12 performs the orthographic conversion process 203A using the distortion correction function 203 so as to be in a state without distortion from the state of the wide-angle image G1 (original image). Here, when performing the calculation of the orthographic conversion process 203A on the entire circular area of the original image, there is a concern that the calculation amount will increase. Therefore, in the first embodiment, as described above, as an example of processing, the imaging processing unit 12 performs calculations such as the orthographic conversion process 203A limitedly on an image obtained by trimming a partial area of the wide-angle image G1 (original image).

[0106] [(9) Distortion correction] Using FIGS. 13 and 14, a distortion correction method for correcting a wide-angle image with distortion captured by the in-camera C1 into a flattened image without distortion will be described.

[0107] FIG. 13 shows a model and a coordinate system for orthographic transformation when using the wide-angle lens of the in-camera C1. A hemispherical surface 500 corresponding to the wide-angle lens, a planar imaging surface 501 of the camera sensor, a flattened image 502, etc. are shown. The imaging surface 501 is also shown as the (x, y) plane of the original image. The flattened image 502 corresponds to the captured image of the object to be photographed and is shown as the (u, v) plane. Although the radius and the like of the wide-angle lens vary depending on the angle of view and the like, it has a shape close to a spherical surface and is shown as the hemispherical surface 500. The imaging surface 501 is arranged at a position perpendicular to the Z-axis of the imaging range of the hemispherical surface 500. The origin O, radius R, and three-dimensional coordinates (X, Y, Z) of the hemispherical surface 500 and the imaging surface 501 are shown. The plane of the bottom surface of the spherical coordinate system (hemispherical surface 500) corresponds to the imaging surface 501 on the camera sensor and is shown as the (x, y) plane of the original image.

[0108] For the position and the angle of view AV2 where the face of user A, which is the object to be photographed, is photographed within the angle of view AV1 of the wide-angle image captured through the wide-angle lens of the in-camera C1, the mobile information terminal 1 can determine the distance, angle, range, etc. from the center position based on the model of FIG. 13. Therefore, the imaging processing unit 12 can perform distortion correction processing (orthographic transformation of FIG. 14) based on the model of FIG. 13 to convert a face image with distortion into a flattened image with the distortion eliminated or reduced.

[0109] As shown in FIG. 2 described above, the angle of view AV1 by the wide-angle lens of the in-camera C1 is as large as about 180 degrees. An optical image that captures and transmits an object (for example, a face) with a shape close to a spherical surface (hemispherical surface 500) corresponding to the wide-angle lens is received by a camera sensor having a plane shown by the imaging surface 501. Then, as shown on the left side of FIG. 14, distortion occurs in the image (wide-angle image) on the camera sensor. The relationship (indicated by the angle β) between the front angle (Z direction) of the camera sensor and the direction of the captured image (n direction) becomes tighter toward the outer periphery compared to the center.

[0110] On the left side of FIG. 14, the distortion in the (x, y) plane of the original image is conceptually shown. In the (x, y) plane, the unit region of the coordinate system is not a right-angled rectangle (e.g., distortion amounts δu, δv). On the right side of FIG. 14, the (u, v) plane is shown as the flattened image 502. This (u, v) plane is the image that we want to extract as an image without distortion. In the (u, v) plane, the unit region of the coordinate system is a right-angled rectangle (e.g., Δu, Δv).

[0111] The specifications regarding the focal length and lens shape of the wide-angle lens are known in advance. Therefore, it is possible to easily perform coordinate conversion from the spherical image (original image) to the planar image (referred to as the flattened image). As this coordinate conversion, an orthographic transformation as shown in FIG. 14 can be applied. This orthographic transformation converts an image with distortion into an image without the distortion as seen by the human eye, and is used, for example, in distortion correction of a fish-eye lens or the like. Using this orthographic transformation, the pixel at each point position in the (x, y) plane of the original image on the left side of FIG. 14 is converted into the pixel at each point position in the (u, v) plane of the flattened image on the right side. The details of the conversion are as follows.

[0112] As shown in FIG. 13, with respect to the shooting direction n of the target object, let the azimuth angle be α, the zenith angle be β, and the rotation angle of the plane to be extracted (flattened image 502) be φ. Let the magnification of the lens be m. Then, from the known orthographic transformation equations, the following equations 1 and 2 are obtained. Also, A, B, C, and D in those equations are given by equations 3 to 6. Equation 1: x = R(uA - vB + mRsinβsinα) / √(u 2 + v 2 + m 2 R 2 ) Equation 2: y = R(uC - vD + mRsinβcosα) / √(u 2 + v 2 + m 2 R 2 ) Equation 3: A = cosφcosα - sinφsinαcosβ Equation 4: B = sinφcosα + cosφsinαcosβ Equation 5: C = cosφsinα + sinφcosαcosβ Equation 6: D = sinφsinα - cosφcosαcosβ

[0113] By performing calculations according to the above equation, each pixel of the original image can be converted into each pixel of the flattened image to eliminate distortion. The distortion amounts δu and δv at each pixel position in the (x, y) plane of the original image are made into a distortion-free state (Δu, Δv) like a square unit area in the (u, v) plane of the flattened image.

[0114] [(10) Trapezoidal correction] FIG. 15 shows an image example and the like of the trapezoidal correction process 203B. As shown in FIG. 15(A), when the portable information terminal 1 is placed flat on the horizontal plane s0, the face of user A is photographed in the obliquely upward direction J3 as viewed from the in-camera C1. Therefore, in the angle of view AV2 of the face photographing image, the portion (such as the chin) on the lower side in the Z direction of the face and the head is closer than the portion on the upper side in the Z direction. The distance DST1 is an example of the distance between the position PC1 of the in-camera C1 and the upper portion of the head, and the distance DST2 is an example of the distance between the position PC1 of the in-camera C1 and the lower portion of the head. DST1 > DST2.

[0115] Therefore, in the photographed image (wide-angle image), as shown in FIG. 15(B), the face area of user A and the like become a trapezoidal image content (referred to as a trapezoidal image). In the trapezoidal image 151, the upper side in the Z direction is smaller than the lower side. The trapezoidal image 151 schematically shows the shape of the wide-angle image in a distortion-free state after the orthographic transformation. In the trapezoidal image 151, for example, the top of the head side is relatively small and the chin side is relatively large.

[0116] When the transmission image D12 is configured using this trapezoidal image, the other party (user B) may feel a slight sense of discomfort when viewing the transmission image D12. Therefore, in order to create a more suitable transmission image D12, the trapezoidal correction process 203B is used. The portable information terminal 1 can determine the state such as the range of the elevation angle (angle θ1) when viewing the representative position point P1 of the face of user A from the position PC1 of the in-camera C1 using the image and sensors of the in-camera C1. The photographing processing unit 12 performs the trapezoidal correction process 203B based on the determined information.

[0117] Figure 15(C) shows the image 152 obtained as a result of trapezoid correction from (B). This image 152 is a rectangular image, and the upper side and the lower side have the same length. By creating the transmission image D12 using this image 152, for the other party (user B), the image will be such that user A's face is seen from the front.

[0118] [(11) Monitor function, image correction function] As shown in Figure 7, the mobile information terminal 1 of Embodiment 1 also has a monitor function for the image of user A himself / herself. The mobile information terminal 1 displays, within the display screen DP, the monitor image D13 of user A corresponding to the transmission image D12, and checks with user A whether the monitor image D12 can also be used as the transmission image in the state of the face. With this monitor function, if user A looks at the monitor image D13 and feels discomfort or dislike for his / her own face image after distortion correction, user A can reject the transmission of the transmission image D12 corresponding to that monitor image D13. When the mobile information terminal 1 receives the instruction operation of rejecting the transmission, it does not transmit the transmission image D12 corresponding to that monitor image D13.

[0119] Also, in that case, the mobile information terminal 1 may create a new transmission image D12 based on the registered image D10, replace it with the original transmission image D12, and transmit it. The mobile information terminal 1 may use the face image of the registered image D10 as it is, or process it to create the transmission image D12. The registered image D10 may be a still image or a moving image taken by user A of his / her own face with an arbitrary camera, or may be any other arbitrary image (icon image, animation image, etc.) other than the face.

[0120] As another function, the mobile information terminal 1 may be provided with an image correction function. When receiving an instruction to reject transmission for a corrected image D11 (monitor image D13, transmission image D12) once created, the mobile information terminal 1 uses this function to correct the face image. At this time, the mobile information terminal 1 performs a correction process on the corrected image D11 based on the face image of the registered image D10 to create a corrected face image. For example, the mobile information terminal 1 processes the state of both eyes in the face image so that the line of sight is directed forward to create a corrected face image. Note that the corrected face image may also be displayed on the display screen DP and sent to user A for confirmation.

[0121] As a specific example, in the face image of user A taken at a certain timing, the eyes may not be captured well. For example, the line of sight may deviate greatly from the front (direction J3). In that case, the mobile information terminal 1 corrects the eye part of the face image by synthesizing or replacing it with the eye part of the registered image D10. As a result, the eyes in the corrected face image face forward.

[0122] Also, when user A gives an instruction to reject transmission after checking the monitor image D13, the mobile information terminal 1 may reset the processing results (corrected image D11, transmission image D12) at that time and retry creating the transmission image D12 based on input images at different timings. The number of retries can also be set. If, as a result of retrying up to a predetermined number of times, a sufficient image cannot finally be obtained (when user A does not give a transmission instruction), the mobile information terminal 1 may use the face image of the registered image D10 as the transmission image D12.

[0123] Also, the mobile information terminal 1 may switch whether the image to be the transmission image D12 is an image created in real time or the face image of the registered image D12 according to a touch operation of user A on the area within the display screen DP.

[0124] [(12) Registration function and registered image] By using the registration function and the registration image D10, the accuracy of face detection and the like can be improved, and various additional functions can also be utilized. The data of the registration image D1 is stored in the memory 13, the external memory 16, or the like. The method of registering the face image of user A as the registration image D10 is as follows. For example, user A operates the registration function among the user setting functions of the videophone application, uses the normal camera C2 (or the in-camera C3 in Embodiment 2 described later), captures his / her own face from the front, and registers the face image without distortion as the registration image D10. Note that the registration image D10 may be created using the in-camera C1 and the distortion correction function 203, or the registration may be performed by reading data from another camera or an external device.

[0125] The registration image D10 may include not only the face image captured from the front of the face of user A but also a plurality of images captured of the face from various other directions. In this case, even when user A changes the orientation of his / her face or moves during a videophone call, the face detection function 201 of the portable information terminal 1 can detect the state of the face at that time using the registration image D10. The image correction function of the portable information terminal 1 can perform correction processing according to the state of the face.

[0126] Also, when user A checks the monitor image D13 on the display screen DP and rejects the transmission, and instead uses the registration image D10, the transmitted image D12 can also be formed using the face image selected by user A among the plurality of face images of the registration image D10.

[0127] Also, the registration image D10 may include not only the face image of a single user A but also a plurality of face images of a plurality of other users who may make a videophone call using the portable information terminal 1.

[0128] [(13) Personal recognition function] In Embodiment 1, the face detection function 201 of the imaging processing unit 12 also includes a function (personal recognition function 201B) for recognizing the face of a specific individual user. The mobile information terminal 1 may not only detect an unspecified face area from a wide-angle image, but also recognize the face of a specific individual user. In that case, the mobile information terminal 1 may detect only the face of a specific user and create a transmission image D12.

[0129] The face detection function 201 detects an arbitrary face area from a wide-angle image, for example. Then, in the personal recognition function 201B, the face area is compared with a face image for personal recognition of user A registered in the registered image D10 in advance. In the personal recognition function 201B, based on the result of the comparison, it is determined whether the face area in the wide-angle image corresponds to the face of a specific user A based on the similarity. The personal recognition function 201B outputs personal recognition result information.

[0130] The mobile information terminal 1 applies the control of the hands-free television phone function only when it is the face of a specific user A, and creates a transmission image D12, for example. When the faces of a plurality of users are reflected in the wide-angle image, the transmission image D12 can be created targeting only the face of the specific user A. For example, the face of a passerby who is only in the background of user A in the wide-angle image does not need to be targeted. As a modification, the personal recognition function 201B may not be provided. Further, the imaging processing unit 12 may recognize the face of a specific individual with respect to the image before distortion correction, or may recognize the face of a specific individual with respect to the image after distortion correction.

[0131] [(14) Face tracking function] In Embodiment 1, the imaging processing unit 12 (particularly the face detection function 201) also includes a function (face tracking function) for automatically tracking the movement of the user's face area based on the video of the in-camera C1 (a plurality of images at a predetermined rate). By using the wide-angle in-camera C1, the mobile information terminal 1 can track the face by face detection within the wide-angle image even if the user himself / herself moves slightly. The mobile information terminal 1 can also set the tracked transmission image D12 so that the user's face is always at the center of the image.

[0132] During a videophone call, user A may not always be stationary at the same position and may be moving. The shooting processing unit 12 detects the face region of user A for each wide-angle image at predetermined time points from the moving image. For example, after the face region is once detected at a certain time point by the face detection function 201, at subsequent time points, it searches near the detected face region to determine the movement of the face region. Thereby, even when the user is moving, the face region of user A can be continuously tracked on the time axis while suppressing the amount of image processing.

[0133] Also, during a videophone call in the state of FIG. 2, user A (especially the face) may move from the initial position. For example, user A may temporarily move away from the initial position and then return to the initial position. Even in that case, the shooting processing unit 12 tracks the moving face region as much as possible by the face tracking function. When the face of user A is not shown in the wide-angle image, that is, when tracking is impossible, at subsequent times, the shooting processing unit 12 responds as follows, for example. The shooting processing unit 12 responds using the last detected image in the past and the created transmission image D12. Alternatively, the shooting processing unit 12 temporarily switches to the face image of the registered image D10 for response. When the face of user A appears in the wide-angle image again, the face region is detected and tracked in the same manner thereafter. Also, the shooting processing unit 11 can similarly respond by the face tracking function even when the face of user A is temporarily hidden by an arbitrary object.

[0134] [(15) Other usage states, arrangement states, and guidance functions] FIGS. 16 to 18 show examples of other usage states and arrangement states of the portable information terminal 1 according to the first embodiment. Regarding the arrangement state of the terminal when user A makes a hands-free videophone call, it is not limited to the state of FIG. 2, and the following states are also possible.

[0135] FIG. 16 shows a first example of a state. In FIG. 16, there is an arbitrary object 160 such as a stand having an inclined surface s5 inclined at a certain angle 161 with respect to the horizontal plane s0. User A places the housing of the portable information terminal 1 flat along the inclined surface s5 of the object 160 such as the stand. The object 160 and the angle 161 are not particularly limited as long as the housing is in a stationary state. In this state, User A views the display screen DP of the housing in front. The in-camera C1 is arranged in the direction J2 of the optical axis according to the angle 161. The in-camera C1 captures the face (point P1) of User A with the angle of view AV2 in the direction J3. In this way, even if the housing is arranged at a certain inclination, a hands-free video call can be realized as in FIG. 2.

[0136] FIG. 17 shows a second example where the position PC1 of the in-camera C1 on the front surface s1 of the housing is in the position in the front side direction Y2 in the Y direction as viewed from User A. In this case, the face-capturing direction J3 (angle θ3) of the in-camera C1 is different from the state in FIG. 2. For example, the elevation angle is larger. In this case, in the wide-angle image, the face region of User A is shown inverted. The imaging processing unit 12 can recognize the inverted state from the wide-angle image. The imaging processing unit 12 appropriately performs image inversion processing and displays a monitor image D13 with the top and bottom in the appropriate direction on the display screen DP.

[0137] FIG. 18 shows a third example where the housing is arranged along the X direction (the left-right direction as viewed from User A) in the longitudinal direction of the housing. In this case, the position PC1 of the in-camera C1 on the front surface s1 is in the position on one side (for example, the direction X1) in the X direction as viewed from User A. The imaging processing unit 12 appropriately performs image rotation processing and displays a monitor image D13 in the appropriate direction on the display screen DP.

[0138] In Embodiment 1, although the arrangement state shown in FIG. 2 is particularly recommended for the user, the video phone function can be realized in substantially the same manner in each of the above arrangement states. Further, in Embodiment 1, regardless of whether the arrangement state of the portable information terminal 1 on the horizontal plane s0 and the positional relationship with the user A are in any of the states shown in FIGS. 2, 17, 18, etc., the in-camera C1 can respond in substantially the same manner. Therefore, the user A can change the state of the terminal and his own position to some extent freely during a video call, which is highly convenient.

[0139] Further, the portable information terminal 1 may be provided with a function (guidance function) for recommending and guiding the user regarding the arrangement state of the housing. The portable information terminal 1 grasps the arrangement state of the housing by using the camera image and the sensors 30. The portable information terminal 1 may display guidance information (e.g., "It is recommended to place the camera on the back side.") on the display screen DP or output it as voice, for example, so as to recommend the arrangement state shown in FIG. 2.

[0140] Further, when the arrangement state of the housing is not appropriate, the portable information terminal 1 may output guidance information to that effect. For example, when the angle θ3 regarding the direction J3 for photographing the face (point P1) of the user A from the position PC1 of the in-camera C1 is not within a predetermined angle range (when the elevation angle is too small or too large) in a certain arrangement state, the portable information terminal 1 may output guidance information indicating that the arrangement position is not appropriate.

[0141] Further, for example, in the positional relationship between the user A and the portable information terminal 1, a case where the face cannot be recognized particularly from the in-camera C1 is also assumed. In that case, the portable information terminal 1 may output guidance information to the user A in order to make the positional relationship (terminal position and user position) appropriate. For example, it may output information indicating to change the position or arrangement state of the portable information terminal 1, or information indicating to change the position of the face of the user A. At this time, when the portable information terminal 1 grasps the positional relationship with the user A, it may output instruction information on which direction and position to change to.

[0142] [(16) Comparative Example - Hands-Free Off State] FIG. 19 shows a normal video phone mode (normal mode) in the mobile information terminal of the comparative example, and a hands-free off mode as another video phone mode in Embodiment 1. In this state and mode, User A holds the housing of the mobile information terminal in hand, and both hands are not free. In the example of FIG. 19, the housing is in a state of standing vertically upward. The in-camera CX on the front surface of the housing is a normal camera having a normal lens with a normal angle of view (so-called narrow angle). This normal angle of view is narrower than the angle of view AV1 of the in-camera C1. The direction JX2 of the optical axis of the in-camera CX is shown, and in this example, it faces the front side direction Y2 in the Y direction. The angle of view AVX1 of the in-camera CX is, for example, in the angular range from 45 degrees to 135 degrees. The angle of view AVX2 corresponding to the face shooting range within the angle of view AVX1 is shown.

[0143] Also, the direction JX1 of the line of sight of User A, the face shooting direction JX3 of the in-camera CX, and the angle difference AD2 formed by them are shown. The larger such an angle difference is, the more downward the direction of the line of sight of User A in the image becomes. When viewed from the other party (User B), User A's eyes are not facing the front.

[0144] On the other hand, in the case of the hands-free mode in FIG. 2 of Embodiment 1, the angle difference AD1 can be made smaller than the angle difference AD2 of the comparative example. Therefore, in the case of the hands-free mode in Embodiment 1, the direction of the line of sight of the user in the image is closer to the front than in the case of the normal mode. As a result, when viewed from the other party (User B), User A's line of sight is more facing the front, so a more natural and less uncomfortable video phone is possible.

[0145] Note that the conventional in-camera CX of a mobile information terminal has a normal lens, and the shooting range is limited. Therefore, when this in-camera CX is used for the video phone function and user A makes a video call in a non-hands-free state while holding the housing by hand, the following considerations and efforts are required. In order for user A to appropriately capture their own face and convey it to the other party, it is necessary to continuously maintain the positional relationship between the face and the housing while adjusting the orientation of the housing by hand. On the other hand, in the mobile information terminal 1 of Embodiment 1, a video call can be made in a hands-free state using the in-camera C1, and the above considerations and efforts are unnecessary.

[0146] Note that in Embodiment 1, when the normal mode is used, the face of user A will appear at a position near the center within the wide-angle image of the in-camera C1. In the case of this normal mode, the mobile information terminal 1 may omit distortion correction processing and the like, and different effects can be obtained as an operation different from the hands-free mode.

[0147] [(17) Effects, etc.] As described above, according to the mobile information terminal 1 having the video phone function of Embodiment 1, a hands-free video call can be realized with a more suitable usability. The user can make a video call with both hands free, and the convenience is also high. Note that the in-camera C1 on the front surface s1 of the mobile information terminal 1 is not limited to being dedicated to video calls, but is a general one that can also be used for other purposes (such as selfies). In Embodiment 1, a hands-free video call is realized by making good use of the in-camera C1. The mobile information terminal 1 of Embodiment 1 does not need to deform the housing during a hands-free video call, nor does it need to use another fixing device or the like, and has good usability and high versatility.

[0148] In addition, in the mobile information terminal 1 of Embodiment 1, the optical axis of the camera (in-camera C1) is in the general plane perpendicular direction, which is different from the prior art examples in which the direction of the optical axis of the camera is oblique (for example, 45 degrees), or the prior art examples in which the direction of the optical axis of the camera can be mechanically driven and changed, and the implementation is also easy.

[0149] (Embodiment 2) Using FIGS. 20 and 21, the portable information terminal according to Embodiment 2 of the present invention will be described. The basic configuration of Embodiment 2 is the same as that of Embodiment 1. Hereinafter, the components different from those of Embodiment 1 in Embodiment 2 will be described. The portable information terminal 1 of Embodiment 2 is provided with a plurality of in-cameras on the front surface s1 of the housing and uses them properly.

[0150] FIG. 20 shows the configuration of the camera unit 11 in the portable information terminal 1 according to Embodiment 2. In addition to the above-described in-camera C1 and normal camera C2 (which may be a wide-angle camera in particular), this camera unit 11 includes an in-camera C3 having a normal angle of view. The above-described in-camera C1 corresponds to the first in-camera, and the in-camera C3 corresponds to the second in-camera.

[0151] The imaging processing unit 12 includes an out-camera processing unit 12B that processes a normal image of the normal camera C2, a first in-camera processing unit 12A that processes a wide-angle image of the in-camera C1, a second in-camera processing unit 12C that processes a normal image of the in-camera C3, and a mode control unit 12D. The imaging processing unit 12 switches, with the mode control unit 12D, the mode such as which of these plurality of cameras to use. The portable information terminal 1 switches the camera mode according to the grasped positional relationship between the terminal and the face of User A and the arrangement state of the terminal.

[0152] FIG. 21 shows an example of the usage state in the portable information terminal 1 according to Embodiment 2, the angle of view of the in-camera C3, etc. In FIG. 21(A), the housing of the portable information terminal 1 is flatly arranged on the horizontal plane s0. For example, on the front surface s1 of the housing, an in-camera C3 (especially a normal lens unit) is provided at a position PC3 near the position PC1 of the in-camera C1. The direction J4 of the optical axis of the in-camera C3 is vertically upward, similar to the in-camera C1. The angle of view AV4 of the in-camera C3 (the angular range from the first angle ANG3 to the second angle ANG4) is shown. This angle of view AV4 is narrower than the angle of view AV1 of the in-camera C1. For example, the first angle ANG3 is about 60 degrees, and the second angle ANG3 is about 135 degrees.

[0153] State 211 is the same as the state in FIG. 2, indicating a case where the face of user A can be photographed within the viewing angle AV2 within the viewing angle AV1 of the in-camera C1. State 212 indicates a positional relationship where the face of user A can be photographed by the viewing angle AV4 of the in-camera C3. For example, there is a representative point P4 of the face of user A at the tip of the direction J4 of the optical axis of the in-camera C3.

[0154] In the case of a positional relationship where the face can be photographed by the in-camera C3 as in state 212, the portable information terminal 1 switches to a camera mode using the in-camera C3 among the plurality of cameras. Also, in the case of a positional relationship where the face can be photographed by a viewing angle other than the viewing angle AV4 within the viewing angle AV1 of the in-camera C1 as in state 211, the portable information terminal 1 switches to a camera mode using the in-camera C1.

[0155] (B) of FIG. 21 shows an example of a non-hands-free state in Embodiment 2, which is a state close to the non-hands-free state of FIG. 19. For example, when the position of the face of user A shifts from state 211 to state 213 in (B) of FIG. 21, the portable information terminal 1 switches the camera mode from the in-camera C1 to the in-camera C3 and controls it to perform a video call in the non-hands-free mode (normal mode). Similarly, when the position of the face of user A shifts from state 213 to state 211, the portable information terminal 1 switches the camera mode from the in-camera C3 to the in-camera C1 and controls it to perform a hands-free video call. In response to the mode switch, the wide-angle image by the in-camera C1 and the normal image by the in-camera C3 are switched as the input image.

[0156] The portable information terminal 1 may automatically select and switch between the modes of the above two types of cameras based on state detection, or may do so based on an instruction operation or setting by the user. In the mode using the normal image of the in-camera C3, since distortion correction processing and the like are unnecessary, the processing can be made more efficient.

[0157] When the in-camera C1 or in-camera C3 of the mobile information terminal 1 is operating, the face detection function 201 is used to detect the face area of the user A in the image, and based on the position, direction, and angle of view of the face area, it may be possible to select which camera to use and switch the mode. For example, the mobile information terminal 1 may select whether to use the in-camera C3 or the in-camera C1 depending on whether the face is within a predetermined angle of view corresponding to the angle of view AV4 of the in-camera C3 within the angle of view AV1 of the in-camera C1.

[0158] As described above, according to the mobile information terminal 1 of the second embodiment, in addition to the same effects as those of the first embodiment, by using the in-camera C3 with a normal angle of view in combination, the processing can be made more efficient. Regarding the positions of the in-camera C1 and the in-camera C3 on the front surface s1 of the housing of the mobile information terminal 1, it is not limited to the above-described configuration and is possible, for example, as the position of the portion that enters the rectangle of the display screen DP on the front surface s1.

[0159] (Other embodiments) The following are also possible as other embodiments (modification examples) related to the first and second embodiments.

[0160] [Modification example (1) - Shooting process] In the shooting processing unit 12 in the first embodiment, as shown in FIG. 5 and the like, after detecting and trimming the face area from the wide-angle image, distortion correction processing is performed on the trimmed image area. The shooting processing method is not limited to this and is possible.

[0161] FIG. 22 shows an example image of the shooting process on the mobile information terminal 1 of the modified example. First, the shooting processing unit 12 of the mobile information terminal 1 performs distortion correction processing on the entire wide-angle image G1, and then detects and trims a face area or the like from the distortion-corrected image. The mobile information terminal 1 performs distortion correction processing on a 360-degree area (range 221) with a horizontal angle of view or a 180-degree area (the lower semi-circular range 222 of the x-axis) in the wide-angle image G1 to obtain a corresponding flattened image GP3. In the example of FIG. 22, as the distortion-corrected image which is the flattened image GP3, a panoramic image in the case of a 180-degree range 222 with a horizontal angle of view is schematically shown. The mobile information terminal 1 detects an area (for example, area 224) including the face of user A from the flattened image GP3. Then, the mobile information terminal 1 takes a trimming area 225 for the area 224 and trims it to obtain a trimmed image 226, and creates a transmission image D12 or the like from the trimmed image 226.

[0162] In this modified example, compared with the aforementioned Embodiment 1, the area of the image region to be subjected to the distortion correction processing is larger. In the shooting process of the aforementioned Embodiment 1, since the area of the image region to be subjected to the distortion correction processing is smaller, it is advantageous in terms of processing efficiency and the like. In the case of a terminal with high computing performance, the modified example may be adopted. In the modified example, since the image for face detection is the flattened image, it is advantageous in terms of the ease of face detection image processing.

[0163] Also, in another modified example, when detecting the face area from the wide-angle image with distortion, the shooting processing unit 12 performs comparison and collation using the face image of the registered image D10 of user A. The face image of the registered image D10 at that time may be a face image with distortion photographed in advance by the in-camera C1.

[0164] In another modification example, among the viewing angles of the wide-angle image, when performing processing such as face detection, the mobile information terminal 1 refers to an image area in a partial range, for example, the range 222 of the lower semi-circle from the x-axis in FIG. 22 (the range of the elevation angle from 0 degrees to 90 degrees in FIG. 2), and narrows down the processing target image area to that range, and may ignore the upper half of the image area. Furthermore, it may be narrowed down to a narrower range in the horizontal viewing angle as in the example of the area 223. Also, for example, when the mobile information terminal 1 grasps the state as shown in FIG. 2, it may narrow down the processing target image range as described above. In the first state of FIG. 2, among the wide-angle images, the face of user A is captured within the range 222 of the lower semi-circle, and there is almost no possibility of being captured within the range of the upper semi-circle. Therefore, the above processing is effective.

[0165] As another method of shooting processing, the mobile information terminal 1 may first perform a simple first distortion correction process on the wide-angle image by the distortion correction function 203, then perform face detection and trimming, and finally perform a more accurate second distortion correction process on the trimmed image.

[0166] As a modification example, whether the image after distortion correction is allowed as the transmission image D12 may not be determined by the confirmation or operation of user A, but may be automatically determined by the mobile information terminal 1. For example, the mobile information terminal 1 compares the face area of the image after distortion correction with the face area of the registered image D10, evaluates the degree of face reproduction, and calculates an evaluation value. When the evaluation value is equal to or greater than the set threshold value, the mobile information terminal 1 determines that transmission is permitted.

[0167] [Modification Example (2) - Object Recognition Function] In the mobile information terminal 1 of the modification example, the imaging processing unit 12 (particularly the face detection function 201) may have an object recognition function. This object recognition function is a function that recognizes a predetermined object other than a face from a wide-angle image based on image processing and detects the object area. Since user A can freely move their hands during a hands-free video call as shown in FIG. 2, an object held in the hand can be imaged by the in-camera C1. As a result, not only the face of user A but also any object held in the hand can be imaged around it in the transmission image D12 and shown to the other party (user B).

[0168] The predetermined object is an object defined in advance in information processing. The imaging processing unit 12 has a detection algorithm corresponding to the object. Examples of the predetermined object include the user's materials, photos, notes, notebook PC screens, or the user's articles, animals, etc. The predetermined object is defined as an area having a predetermined shape such as a rectangle or a circle, or a predetermined color, etc. In the imaging processing unit 12 (object recognition function), based on the detected face area, for example, a predetermined distance range around the face area may be searched to detect the area of the predetermined object.

[0169] FIG. 23(A) shows an image example when using the object recognition function. This image shows a state of a flattened image without distortion. User A is making a call while showing the object 230 to the other party (user B) during a hands-free video call. The object 230 is, for example, an A4-sized material, etc., and is approximately rectangular in the image after distortion correction. The mobile information terminal 1 uses the object recognition function to detect not only the face area 231 but also the area 232 of the specific object 230 from the image. The mobile information terminal 1 also performs distortion correction processing, etc. on the area 232 of the object 230. The mobile information terminal 1 may search for a specific object, for example, in a range from the point P1 of the face area 231 to a predetermined distance 233 around it. The mobile information terminal 1 may create the transmission image D12, for example, by taking a rectangular area 234 that includes the face area 231 and the area 232 of the object 230. Also, not limited to the image centered on the face area 231 (point P1), it may be an image of the minimum rectangle that includes the face and the object, like the area 235.

[0170] Alternatively, the mobile information terminal 1 may separately divide the face area 231 and the area 232 of the object 230, create respective transmission images D12, display them as monitor images D13 (images 236, 237), and perform transmission confirmation. Further, the mobile information terminal 1 may create, as the monitor image D13, an image that focuses on the detected object 230 (an image enlarged with the object 230 as the center). Also, if an area for securing a large distance around the face is taken using the aforementioned area A3 or area B3, even if the object recognition process is omitted, the object can be automatically captured within that area.

[0171] In this object recognition function, since a wide-angle image is used, even if the distance between the face and the object is somewhat large, both of their images can be obtained. For example, FIG. 23(B) shows a flattened panoramic image GP4 based on a wide-angle image. In this panoramic image GP4, at a position corresponding to the direction Y2, direction J3 in FIG. 2, and the lower side of the y-axis in FIG. 10 (assuming 0 degrees in the horizontal angle of view), there is an area r1 where the face of user A is captured. And from that position, at a position somewhat separated in the horizontal direction, for example, at a position corresponding to 90 degrees in the horizontal angle of view and on the right side of the x-axis, there is an area r2 where a predetermined object is captured. Both of these can be captured within one wide-angle image and used as the transmission image D12.

[0172] As another modification, the photographing processing unit 12 (object recognition function) may detect the hand of user A from the wide-angle image, take an area including the face and hand of user A, and create the transmission image D12. Also, when both hands are captured in the wide-angle image, it can be determined that the user is not holding the housing. Therefore, the mobile information terminal 1 may be configured to switch to the hands-free mode when both hands of user A are detected from the wide-angle image.

[0173] [Modification Example (3) - Face Images of Multiple Users] As a modification example, for instance, with one portable information terminal 1 placed on an airplane, it is also possible to have a usage method where multiple users use the same portable information terminal 1 to conduct a video call with a counterpart (user B) as one of the callers on the transmitting side. In that case, the portable information terminal 1 in the modification example performs processes such as face detection and distortion correction simultaneously and in parallel for the multiple faces of the multiple users included in the wide-angle image. Also, at that time, the portable information terminal 1 may create multiple transmission images D12 separately for each face of each user shown in the wide-angle image, or may create one transmission image D12 including multiple faces.

[0174] Figure 24(A) shows an example image including the faces of multiple users in this modification example. It shows a case where, as one of the callers in a video call, in addition to the main user A, there is another user C. The portable information terminal 1 detects the face region RU1 of user A and the face region RU2 of user C from within the wide-angle image of the in-camera C1. The portable information terminal 1 takes, for example, a region 241 (e.g., a horizontally long rectangle) that encompasses those two face regions and creates it as the transmission image D12. Alternatively, the portable information terminal 1 takes those two face regions as respective trimmed images 242 and 243, creates them as respective transmission images D12, and may display them in parallel. Even when there are three or more users, it can be basically realized in the same way, but it is restricted to the faces of a predetermined number of people (e.g., four people) so as not to have too many people. The portable information terminal 1 may display monitor images D13 of the faces of multiple users within the display screen DP.

[0175] Also, the portable information terminal 1 may grasp which user's face among the multiple faces in the wide-angle image is currently speaking by performing image processing on the wide-angle image, particularly detecting the state of the mouth, and create a transmission image D12, etc., for the face of the currently speaking user. Furthermore, the portable information terminal 1 may grasp which user's face among the multiple faces in the wide-angle image is currently speaking by linking the image processing with the voice processing of the microphone 17. The portable information terminal 1 may display the monitor images D13 of multiple users in parallel within the display screen DP or may display them by switching on the time axis.

[0176] Also, when dealing with the faces of the above-mentioned multiple users, the above processing may be performed only on multiple users whose face images are registered in the mobile information terminal 1 in advance as the registered image D10. The mobile information terminal 1 does not deal with the faces of people who are not registered in advance (such as passers-by). Also, when the processing of the faces of some users is not in time, etc., the registered image D10 may be used as a substitute, or an image of another icon, scenery, etc. may be used as a substitute.

[0177] In the image example of (A) above, the faces of multiple people (User A, User C) are shown in the area within the angle of view in a certain direction as seen from the in-camera C1 (for example, the area L1 below the y-axis in (B)). Not limited to this, by using the wide-angle view angle of the in-camera C1, even when each person is at a position with a different horizontal view angle around the position of the mobile information terminal 1 on the horizontal plane s0, it is possible to handle. For example, even if there is a face in any of the areas L1 to L4 above and below the y-axis and to the left and right of the x-axis in the area near the outer periphery in the wide-angle image G1 of (B) in FIG. 24, it is possible to handle. That is, within one wide-angle image, the multiple faces of the multiple people can be captured and used as the transmission image D12.

[0178] [Modification Example (4) - Opponent Image Correction Function] FIG. 25 shows the opponent image correction function provided in the mobile information terminal 1 of the modification example. When the mobile information terminal 1 of the modification example displays the image of the opponent in the display screen DP (area R1) as shown in FIG. 7, it uses this opponent image correction function to display an image corrected for trapezoid distortion.

[0179] FIG. 25(A) shows an example of a normal opponent image received from the mobile information terminal 2 of the opponent (User B). In the area R1 in the display screen DP of the mobile information terminal 1, a rectangular image g1 of the opponent is displayed.

[0180] (B) of FIG. 25 schematically shows how the image g1 of (A) looks when viewed obliquely downward from the eyes (point P1) of user A in the state as shown in FIG. 2. In the state of (B), the image g1 appears as a trapezoidal shape with the upper side smaller than the lower side. That is, when viewed from user A, the head side of user B appears relatively slightly smaller.

[0181] The portable information terminal 1 grasps the positional relationship between user A and the terminal and the arrangement state of the terminal based on the analysis of the wide-angle image of the in-camera C1 and the detection information of the sensors 30. For example, the portable information terminal 1 infers the state such as the position of user A's face and the distance from the terminal from the position of user A's eyes, the direction of the line of sight, and the size of the face in the image. The portable information terminal 1 sets the ratio (the ratio of the upper side to the lower side) etc. at the time of inverse trapezoidal correction according to the grasped state. The portable information terminal 1 performs inverse trapezoidal correction processing on the rectangular image received from the other party's portable information terminal 2 in accordance with the ratio etc. to obtain an inverse trapezoidal-shaped image. The above ratio may be a preset value.

[0182] (C) of FIG. 25 shows the image g1b after inverse trapezoidal correction of the image g1 of (A). This image g1b is an inverse trapezoidal shape, which is a trapezoid with a large upper side and a small lower side. The portable information terminal 1 displays the inverse trapezoidal-shaped image g1b in the area R1 within the display screen DP. User A views the image g1b of the other party in the area R1 within the display screen DP obliquely downward from the eyes (point P1) in the state of FIG. 2. Then, in this state, the other party's image appears to be close to a rectangular shape as in (A) when viewed from user A. As a result, user A can more easily visually recognize the other party's image and it is more convenient to use.

[0183] [Modification Example (5) - 3D Image Processing Function] As a modification, the portable information terminal 1 may use a function for processing three-dimensional images (three-dimensional image processing function), not limited to two-dimensional images. For example, the camera unit 11 (for example, a normal camera C2) may be provided with a known infrared camera function and a three-dimensional sensor module. Using this, the imaging processing unit 12 processes the wide-angle image of the in-camera C1 as a three-dimensional image. For example, the portable information terminal 1 irradiates, for example, tens of thousands or more infrared dots onto the user's face using this infrared camera function and three-dimensional sensor module. The portable information terminal 1 captures the infrared dots with an infrared camera, reads the subtle unevenness on the face surface from the image, and creates a face three-dimensional map (corresponding three-dimensional image). The portable information terminal 1 may perform distortion correction processing or the like on the three-dimensional image. Further, the portable information terminal 1 may perform three-dimensional face correction processing by collating the three-dimensional image with the three-dimensional face image information among the registered images D10. In that case, clearer and finer correction can be achieved.

[0184] Also, when performing such advanced three-dimensional correction, the portable information terminal 1 may add analysis using machine learning such as deep learning, rather than simply collating images. For example, the portable information terminal 1 may incorporate an AI engine having a deep learning function (software and hardware that perform deep learning using a convolutional neural network). The portable information terminal 1 learns about the user's face from the camera image using the AI engine to improve the performance of face detection and face correction. As a result, for example, specifically, it is possible to detect, recognize, and correct the face of user A considering differences and influences such as changes due to a person's hairstyle and makeup, the presence or absence of glasses and sunglasses, and the growth of a beard.

[0185] Also, in the personal recognition function 201B, the portable information terminal 1 may compare and collate the three-dimensional face image of the registered image D10 with the three-dimensional face image captured by the in-camera C1 and distortion-corrected. Thereby, personal recognition of whether the user is A himself / herself can be realized with higher accuracy.

[0186] [Modification Example (6) - Directional Microphone, Directional Speaker] As a modification, the microphone 17 of the portable information terminal 1 in FIG. 3 may be a directional microphone. The directional microphone includes a voice processing function such as a noise cancellation function. The controller 10 preferentially collects the voice from the direction where the face of user A is located using the microphone 17. The controller 10 cancels the noise from the input voice by the noise cancellation function to obtain a clear voice of user A. The controller 10 transmits the voice data to the other party's portable information terminal 2 together with the transmission image D12. Regarding the direction where the face of user A is located as viewed from the portable information terminal 1, it can be grasped by grasping the state of the portable information terminal 1 and face detection in the image.

[0187] Regarding the microphone 17, a known MEMS microphone or the like may be applied, and the directivity and noise cancellation functions may be realized by a known beamforming technique. For example, when realizing the noise cancellation function, basically a plurality of microphones are required. However, when the portable information terminal 1 is small, it may not be possible to mount a plurality of microphones. In that case, in the portable information terminal 1, by mounting a MEMS microphone, a specific sound source can be separated and emphasized from a plurality of sound sources by the beamforming technique. Thereby, it is possible to obtain only by emphasizing the voice of user A.

[0188] In addition, the portable information terminal 1 can recognize the position and direction of user A with a certain degree of accuracy by using the in-camera C1. Therefore, the portable information terminal 1 roughly specifies the position and direction of user A using the in-camera C1. The portable information terminal 1 may preferentially emphasize and acquire the voice from that direction using the microphone 17 and the beamforming technique for the specified position and direction.

[0189] Also, as a modification, the portable information terminal 1 may roughly estimate the position and direction of the face of user A based on the analysis of the voice of the microphone 17. The portable information terminal 1 may perform processing such as face detection on the wide-angle image according to the position and direction of the face.

[0190] Similarly, a directional speaker may be used as the speaker 18. The directivity, volume, etc. of the audio output of the speaker 18 may be controlled according to the position of the face of user A with respect to the terminal.

[0191] As described above, the present invention has been specifically described based on the embodiments. However, the present invention is not limited to the above embodiments, and various modifications can be made without departing from the gist thereof.

Explanation of Reference Numerals

[0192] 1... Portable information terminal, 2... Portable information terminal, s0... Horizontal plane, s1... Front surface, s2... Rear surface, DP... Display screen, C1... In-camera, J1, J2, J3... Directions, AV1, AV2... Viewing angles, θ1, θ2... Angles, P1, PC1, PD... Points, ANG1, ANG2... Angles, AD1... Angle difference.

Claims

1. A portable information terminal having a videophone function for transmitting a transmission image via a communication unit, a display, a camera disposed on the same housing surface as the housing surface on which the display is disposed, a control unit that repeatedly executes a process of generating an image signal from a signal output by the camera and detecting a face from the image signal, and the control unit sets a region including the detected face as a trimming region, extracts a signal of a region corresponding to the trimming region from the image signal to generate a trimming image, generates the transmission image of a predetermined size from the trimming image, further, the control unit when the position of the detected face is not fixed within the image signal during a period of repeatedly executing the face detection, sets a first trimming region at a first time point and a second trimming region at a second time point to be different from each other, the control unit switches the angle of view of the image signal based on the result of the process of detecting the face, or switches the angle of view of the image signal based on whether the face fits within the image signal, further includes, on the same housing surface, a camera having a different angle of view from the camera, the switching of the angle of view is performed by switching between the camera and the camera having a different angle of view, A portable information terminal, characterized by the above.

2. the control unit when a face is detected in a first region by the process of detecting the face, executes the process of detecting the face next from the vicinity of the first region, The portable information terminal according to claim 1.

3. the control unit when a face cannot be detected in the process of detecting the face, outputs a notification to the display. The mobile information terminal according to claim 1.

4. The control unit performs distortion correction processing so that distortion caused by the lens of the camera is eliminated or reduced with respect to the trimming image, and generates the transmission image based on the corrected trimming image. The mobile information terminal according to claim 1.

5. The control unit creates a monitor image of the image content corresponding to the transmission image, displays the monitor image in a predetermined area of the display, and determines whether to permit or reject the transmission of the transmission image based on a predetermined operation on the operation unit. The mobile information terminal according to claim 1.

6. comprises a storage unit that stores an arbitrary image in advance as a registered image, and the control unit when the predetermined operation is an operation for rejecting the transmission of the transmission image, transmits the registered image as an alternative image to the transmission image. The mobile information terminal according to claim 5.

7. comprises a storage unit that stores an image including the face of the user of the mobile information terminal in advance as a registered image, and the control unit based on the comparison result between the trimming image and the registered image, when the user using the mobile information terminal is the user of the mobile information terminal, generates the transmission image, and when the user using the mobile information terminal is not the user of the mobile information terminal, does not generate the transmission image. The mobile information terminal according to claim 1.

8. The control unit detects an object held by the user using the mobile information terminal, Extract a signal of a region corresponding to the region including the object from the image signal to generate a second trimmed image, Create the transmission image using the second trimmed image and the trimmed image, The mobile information terminal according to claim 1.

9. The control unit, Detect a plurality of faces from the image signal, Set a region including the plurality of faces as the trimming region, The mobile information terminal according to claim 1.

10. The control unit, Detect a plurality of faces from the image signal, Set a plurality of trimming regions each including a face of each of the plurality of faces, Generate the transmission image using a plurality of trimmed images generated based on the plurality of trimming regions, The mobile information terminal according to claim 1.

11. A method executed by a mobile information terminal having a videophone function of transmitting a transmission image via a communication unit, The mobile information terminal, A display, A camera disposed on the same housing surface as the housing surface on which the display is disposed, A control unit that repeatedly executes a process of generating an image signal from a signal output by the camera and detecting a face from the image signal, The control unit, Set a region including the detected face as a trimming region, Extract a signal of a region corresponding to the trimming region from the image signal to generate a trimmed image, A step of generating the transmission image of a predetermined size from the trimmed image, Further, the control unit, During the period of repeatedly executing the face detection, when the position of the detected face is not fixed within the image signal, setting the first trimming region at the first time point and the second trimming region at the second time point to be different from each other; The control unit switching the angle of view of the image signal based on the result of the process of detecting the face, or switching the angle of view of the image signal based on whether the face is included in the image signal; The mobile information terminal further includes a camera having a different angle of view from the camera on the same housing surface; The switching of the angle of view is performed by switching between the camera and the camera having a different angle of view; A method characterized by the above.

12. The control unit, When a face is detected in the first region by the process of detecting the face, performing the process of detecting the face to be executed next from the vicinity of the first region; The method according to claim 11.

13. The control unit, When a face cannot be detected in the process of detecting the face, outputting a notification to the display; The method according to claim 11.

14. The control unit, Performing distortion correction processing so that the distortion caused by the lens of the camera is eliminated or reduced for the trimmed image, Generating the transmission image based on the trimmed image after correction; The method according to claim 11.

15. The control unit, Creating a monitor image of the image content corresponding to the transmission image, Displaying the monitor image in a predetermined region of the display, and determining whether to permit or reject the transmission of the transmission image based on a predetermined operation on the operation unit; The method according to claim 11.

16. The mobile information terminal includes a storage unit that stores an arbitrary image in advance as a registered image, and when the control unit, in a case where the predetermined operation is an operation that rejects transmission of the transmission image, transmits the registered image as an alternative image to the transmission image, The method according to claim 15.

17. The mobile information terminal includes a storage unit that stores an image including the face of the user of the mobile information terminal in advance as a registered image, and the control unit, based on a comparison result between the trimmed image and the registered image, when the user using the mobile information terminal is the user of the mobile information terminal, generates the transmission image, when the user using the mobile information terminal is not the user of the mobile information terminal, does not generate the transmission image, The method according to claim 11.

18. The control unit, detects an object held by the user using the mobile information terminal, extracts a signal of a region corresponding to the region including the object from the image signal to generate a second trimmed image, and creates the transmission image using the second trimmed image and the trimmed image, The method according to claim 11.

19. The control unit, detects a plurality of faces from the image signal, and sets a region including the plurality of faces as the trimming region, The method according to claim 11.

20. The control unit, detects a plurality of faces from the image signal, sets a plurality of trimming regions each including a face of each of the plurality of faces, and generates the transmission image using a plurality of trimmed images generated based on the plurality of trimming regions. The method according to claim 11.

Citation Information

Patent Citations

  • Utterer detection system and video conference system using same

    JP2004118314A

  • Device for adjusting angle of field

    JP2004282535A

  • Portable telephone set

    JP2005175777A

  • Communication terminal unit, videophone control method and its program

    JP2006067436A

  • Portable terminal device

    JP2007017596A