Image processing system, program, and image processing method

The image processing device generates avatars that synchronize with facial movements, offering privacy and expression enhancement by emphasizing specific facial features, addressing the limitations of existing technologies in controlling facial expressions.

JP2025140839AActive Publication Date: 2025-09-29SOFTBANK CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024040440
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-14
Publication Date
2025-09-29
Estimated Expiration
2044-03-14

AI Technical Summary

Technical Problem

Existing image processing technologies do not effectively allow users to control or enhance facial expressions through avatars in a way that hides their own facial movements or emphasizes specific features, which can be useful for privacy or expression enhancement.

Method used

An image processing device that analyzes facial movements, applies weighted feature amounts to generate avatars that move in sync with the user's face, allowing modes to emphasize eyes, mouth, or overall expressions, and superimposes these avatars on the user's face.

Benefits of technology

Enables users to conceal their facial expressions while conveying intentions strongly or provide an interesting experience by emphasizing specific facial features, benefiting those with facial disabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025140839000001_ABST
    Figure 2025140839000001_ABST
Patent Text Reader

Abstract

To provide an image processing system, a program, and an image processing method for displaying an avatar that follows movement of a user.SOLUTION: An image processing system 100 includes: a first calculation unit that analyzes image data of a face of a user and calculates a feature amount including a movement direction and a movement speed of each of a plurality of feature points of the face of the user; a second calculation unit that calculates a weighted feature amount by applying weight corresponding to a part of the face of the user to each feature amount of at least some of the feature points of the face of the user; an avatar generation unit that generates, on the basis of the weighted feature amount of each of the plurality of feature points, an avatar that moves in accordance with the movement of the face of the user; and a display control unit that controls display of the avatar generated by the avatar generation unit so as to be superimposed on a part of the face of the user in image data.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device, a program, and an image processing method. [Background technology]

[0002] Patent Document 1 describes a technique for displaying an avatar that follows the movements of a user. [Prior art document] [Patent documents] [Patent Document 1] JP 2023-094549 A Summary of the Invention [Means for solving the problem]

[0003] According to one embodiment of the present invention, there is provided an image processing device. The image processing device may include a first calculation unit that analyzes imaging data capturing an image of a user's face and calculates feature amounts including a movement direction and a movement speed of each of a plurality of feature points on the user's face. The image processing device may include a second calculation unit that calculates weighted feature amounts by applying weights corresponding to parts of the face to the feature amounts of each of at least some of the feature points on the user's face. The image processing device may include an avatar generation unit that generates an avatar that moves in accordance with the movement of the user's face, the avatar generation unit generating the avatar based on the weighted feature amounts of each of the plurality of feature points. The image processing device may include a display control unit that controls the avatar generated by the avatar generation unit to be displayed superimposed on the user's face in the imaging data.

[0004] In the image processing device, the second calculation unit may calculate the weighted feature amount by applying a weight corresponding to a part of the face to the movement speed included in each of the feature amounts of the at least some of the plurality of feature points, and the avatar generation unit may generate the avatar with each part changing by moving each of the at least some of the plurality of feature points according to the movement speed to which the weight has been applied. The image processing device may include a range storage unit that stores a movable range of each of the plurality of feature points of the user's face, and the avatar generation unit may generate the avatar with each part changing by moving each of the at least some of the plurality of feature points within the movable range according to the movement speed to which the weight has been applied.

[0005] Any of the image processing devices may include a mode management unit that switches between a normal mode and an eye emphasis mode that emphasizes the eyes of the user, and when the normal mode is set, the avatar generation unit may generate the avatar based on the imaging data, and when the eye emphasis mode is set, the second calculation unit may calculate weighted features by applying weights corresponding to the eyes to the feature amounts of each of a plurality of feature points of the user's face that correspond to the user's eyes, and the avatar generation unit may generate the avatar using the weighted features of the plurality of feature points that correspond to the user's eyes. The mode management unit may switch between the normal mode, the eye emphasis mode, and an expression emphasis mode that emphasizes the user's expression, and when the expression emphasis mode is set, the second calculation unit may calculate weighted feature amounts by applying weights corresponding to parts of the face to the feature amounts of each of a plurality of feature points corresponding to the user's eyes, eyebrows, cheeks, and mouth among the plurality of feature points of the user's face, and the avatar generation unit may generate the avatar using the weighted feature amounts of the plurality of feature points corresponding to the user's eyes, eyebrows, cheeks, and mouth.

[0006] Any of the image processing devices may include a mode management unit that switches between a normal mode and an expression enhancement mode that enhances the user's expression, and when the normal mode is set, the avatar generation unit may generate the avatar based on the imaging data. When the expression enhancement mode is set, the second calculation unit may calculate weighted features by applying weights corresponding to parts of the face to the feature amounts of each of a plurality of feature points corresponding to the user's eyes, eyebrows, cheeks, and mouth among the plurality of feature points of the user's face, and the avatar generation unit may generate the avatar based on the weighted features of the plurality of feature points corresponding to the user's eyes, eyebrows, cheeks, and mouth and the feature amounts of a plurality of feature points other than the user's eyes, eyebrows, cheeks, and mouth.

[0007] According to one embodiment of the present invention, there is provided a program for causing a computer to function as the image processing device.

[0008] According to one embodiment of the present invention, there is provided an image processing method executed by a computer. The image processing method may include a first calculation step of analyzing imaging data capturing an image of a user's face and calculating feature amounts including a movement direction and a movement speed of each of a plurality of feature points on the user's face. The image processing method may include a second calculation step of calculating weighted feature amounts by applying weights corresponding to parts of the face to the feature amounts of at least some of the feature points on the user's face. The image processing method may include an avatar generation step of using an avatar generation unit to generate an avatar that moves in accordance with the movement of the user's face, generating the avatar based on the weighted feature amounts of each of the plurality of feature points. The image processing method may include a display control step of controlling the avatar generated by the avatar generation unit to be displayed superimposed on the portion of the user's face in the imaging data.

[0009] The above summary of the invention does not list all of the necessary features of the present invention, and subcombinations of these features may also constitute inventions. [Brief explanation of the drawings]

[0010] [Figure 1] 1 illustrates a schematic diagram of an example avatar display system 10. [Figure 2] 1 shows another example of an avatar display system 10 in a simplified manner. [Figure 3] 1 shows an example of a functional configuration of an image processing device 100. [Figure 4] FIG. 10 is an explanatory diagram for explaining eye enhancement. [Figure 5] FIG. 10 is an explanatory diagram for explaining facial expression emphasis. [Figure 6] 1 shows an example of a processing flow by the image processing device 100. [Figure 7] 1 shows an example of a hardware configuration of a computer 1200 that functions as the image processing device 100 or the communication device 200. DETAILED DESCRIPTION OF THE INVENTION

[0011] The present invention will be described below through embodiments of the invention, but the following embodiments do not limit the scope of the invention according to the claims. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.

[0012] User convenience may be improved by having an avatar perform movements that are different from the user's facial movements, rather than the user's facial movements themselves. For example, some users may wish to prevent others from understanding their facial expressions, and this can be accommodated. For example, for a user who has difficulty expressing themselves due to a facial defect, having an avatar perform movements that resolve the defect can enable the user to express expressions similar to those of a person without a facial defect. Furthermore, having an avatar perform movements that emphasize the user's facial movements can give others a good impression or an interesting experience.

[0013] Fig. 1 schematically illustrates an example of an avatar display system 10. In the example illustrated in Fig. 1, the avatar display system 10 includes an image processing device 100. The image processing device 100 includes an imaging unit 102. The image processing device 100 may be, but is not limited to, a smartphone, a tablet terminal, a PC (Personal Computer), or the like.

[0014] The imaging unit 102 provides imaging data generated by imaging an imaging target to the image processing device 100. The imaging unit 102 has an imaging camera. The imaging data may include captured images generated by the imaging camera. The captured images may be RGB images. The captured images may be moving images. The captured images may also be still images captured continuously.

[0015] The imaging unit 102 may further include a depth camera. The depth camera is a camera capable of generating a depth information image including depth information, which is the distance to the imaging target. The depth information image includes depth information for each pixel. The depth camera is sometimes called a depth camera. A known technique can be used to measure the distance, and for example, LiDAR (Light Detection And Ranging), triangulation, or TOF (Time Of Flight) may be used. The imaging data may include the depth information image.

[0016] The imaging unit 102 may include a stereo camera. The stereo camera may generate a captured image and a depth information image.

[0017] The imaging unit 102 may be built into the image processing device 100. The imaging unit 102 may also be attached externally to the image processing device 100.

[0018] The image processing device 100 may acquire imaging data of the face of the user 40 captured by the imaging unit 102 from the imaging unit 102, and use the imaging data to generate an avatar of the face of the user 40. For example, the image processing device 100 may generate an avatar of the face of the user 40 using the captured image and depth information image included in the imaging data. The depth information image generated by capturing an image of the face of the user 40 may be an image that represents the depth (three-dimensional shape) of the user 40 using discrete point cloud data.

[0019] The image processing device 100 generates an avatar by, for example, using a depth information image to generate a 3D model of the face of the user 40, using a captured image to generate texture data of the face of the user 40, and performing a process of pasting the texture data onto the 3D model. The image processing device 100 may generate an avatar that moves in accordance with the movement of the face of the user 40 by continuously generating avatars from continuously acquired captured image data.

[0020] The image processing device 100 according to this embodiment may generate an avatar that emphasizes part or all of the face of the user 40. For example, the image processing device 100 generates an avatar that emphasizes the facial expression of the user 40. For example, the image processing device 100 generates an avatar that emphasizes the eyes of the user 40.

[0021] The image processing device 100 may perform control to display a display image in which an avatar is placed in an area corresponding to the face of the user 40 in an RGB image included in the captured data. For example, the image processing device 100 may perform control to display the display image on a display provided in the image processing device 100. For example, the image processing device 100 may transmit the display image to another device and perform control to display the display image on the display of the other device.

[0022] The image processing device 100 may analyze imaging data capturing an image of the face of the user 40 and calculate a feature amount for each of a plurality of feature points on the face of the user 40. The feature amount may include a movement direction and a movement speed of the feature point. The image processing device 100 may calculate weighted feature amounts by applying weights corresponding to parts of the face to the feature amounts of at least some of the feature points on the face of the user 40. The image processing device 100 may generate an avatar using the weighted feature amounts for feature points to which weights have been applied, and using the feature amounts for feature points to which no weights have been applied.

[0023] For example, the image processing device 100 calculates weighted features by applying weighted values ​​corresponding to facial parts to the movement speeds included in the features of at least some of the feature points on the face of the user 40, and generates an avatar in which each part changes by moving each of the at least some of the feature points according to the movement speed to which the weighted value has been applied.

[0024] For example, the image processing device 100 generates an avatar that emphasizes the movement of a part of the face of the user 40. As a specific example, the image processing device 100 calculates weighted feature amounts by applying weighted values ​​corresponding to the eyes to the movement speeds included in the feature amounts of each of a plurality of feature points on the face of the user 40 that correspond to the eyes of the user 40, and generates an avatar in which the movement of the eyes is emphasized by moving each of the plurality of feature points corresponding to the eyes according to the movement speed to which the weighted value has been applied.

[0025] For example, the image processing device 100 generates an avatar that emphasizes the movements of multiple parts of the face of the user 40. As a specific example, weighted feature amounts are calculated by applying weighted values ​​corresponding to the parts to the movement speeds included in the feature amounts of each of multiple feature points corresponding to the eyes, nose, mouth, and cheeks of the face of the user 40, and an avatar with emphasized facial expressions is generated by moving each of the multiple feature points of the eyes, nose, mouth, and cheeks of the face of the user 40 according to the movement speed to which the weighted value has been applied. Of the parts of the face of the user 40, the parts whose movements the avatar should be emphasized may be set arbitrarily.

[0026] In this way, by emphasizing part or all of the face of the user 40, for example, it is possible to prevent the user's facial expression from being revealed to the other party communicating through the avatar. It is also possible to give the other party a good impression and provide an interesting experience. It is also possible for a user 40 who has difficulty expressing themselves due to some kind of facial disability to express their facial expressions in the same way as a person without a disability.

[0027] Fig. 2 schematically illustrates another example of the avatar display system 10. In the example illustrated in Fig. 2, the avatar display system 10 includes an image processing device 100 and a communication device 200. The communication device 200 includes an imaging unit 202. The communication device 200 may be, but is not limited to, a smartphone, a tablet terminal, a PC (Personal Computer), or the like.

[0028] The imaging unit 202 may be similar to the imaging unit 102. The imaging unit 202 provides the communication device 200 with imaging data generated by capturing an image of an imaging target. The imaging unit 202 has an imaging camera. The imaging data may include captured images generated by the imaging camera. The captured images may be RGB images. The captured images may be moving images. The captured images may be continuously captured still images. The imaging unit 202 may further have a depth camera. The imaging unit 202 may have a stereo camera. The stereo camera may generate captured images and depth information images. The imaging unit 202 may be built into the communication device 200. The imaging unit 202 may be external to the communication device 200.

[0029] The image processing device 100 receives imaging data generated by the imaging unit 202 capturing an image of the user 40 from the communication device 200 via the network 20.

[0030] The network 20 may include the Internet. The network 20 may include a LAN (Local Area Network). The network 20 may include a mobile communication network. The mobile communication network may conform to any of the following communication methods: LTE (Long Term Evolution), 5G (5th Generation), 3G (3rd Generation), and 6G (6th Generation) or later.

[0031] The image processing device 100 may be connected to the network 20 by wire. The image processing device 100 may be connected to the network 20 by wireless. The image processing device 100 may be connected to the network 20 via a wireless base station. The image processing device 100 may be connected to the network 20 via a Wi-Fi (registered trademark) access point.

[0032] The communication device 200 may be connected to the network 20 by wire. The communication device 200 may be connected to the network 20 by wireless. The communication device 200 may be connected to the network 20 via a wireless base station. The communication device 200 may be connected to the network 20 via a Wi-Fi access point.

[0033] The image processing device 100 may generate an avatar of the face of the user 40 using the imaging data received from the communication device 200. The image processing device 100 may control to display a display image in which the avatar is arranged in an area corresponding to the face of the user 40 in an RGB image included in the imaging data. For example, the image processing device 100 may control to display the display image on a display provided in the image processing device 100. The image processing device 100 may transmit the display image to the communication device 200 and control to display the display image on the display of the communication device 200. The image processing device 100 may transmit the display image to a device other than the communication device 200 and control to display the display image on the display of the device.

[0034] 3 schematically illustrates an example of the functional configuration of the image processing device 100. The image processing device 100 includes an imaging data acquisition unit 112, a first calculation unit 114, a second calculation unit 116, an avatar generation unit 118, a mode management unit 120, a range storage unit 122, and a display control unit 124. It is not essential that the image processing device 100 include all of these units.

[0035] The imaging data acquisition unit 112 acquires imaging data capturing an image of the face of the user 40. For example, the imaging data acquisition unit 112 acquires imaging data generated by the imaging unit 102 from the imaging unit 102. For example, the imaging data acquisition unit 112 receives imaging data generated by the imaging unit 202 from the communication device 200.

[0036] The first calculation unit 114 analyzes the imaging data acquired by the imaging data acquisition unit 112 and calculates a feature amount for each of a plurality of feature points on the face of the user 40. The feature amount includes a movement direction and a movement speed of the feature point. The first calculation unit 114 may calculate the feature amount for each of the plurality of feature points using a depth information image included in the imaging data.

[0037] The second calculation unit 116 calculates weighted feature amounts by applying weighted values ​​corresponding to facial features to the feature amounts of at least some of the feature points calculated by the first calculation unit 114. The weighted values ​​corresponding to facial features may be different values ​​depending on the facial feature. For example, the weighted value corresponding to the eyes may be different from the weighted value corresponding to the mouth. The weighted values ​​corresponding to the facial features may be the same value. For example, the weighted value corresponding to the eyes may be the same as the weighted value corresponding to the mouth. The weighted values ​​corresponding to the facial features may be arbitrarily set or may be changeable.

[0038] The weights corresponding to the facial features may be changeable as appropriate. The weights may be changeable, for example, by a slide bar displayed on the display of the image processing device 100. The weights may also be changeable by a method other than a slide bar. The second calculation unit 116 may calculate weighted feature amounts by applying weights corresponding to the facial features, which are set at that time, to the feature amounts of at least some of the feature points calculated by the first calculation unit 114. The range of the weights may be different or the same for different facial features. An example of the range of the weights is 1.0 to 1.5, but is not limited to this.

[0039] The second calculation unit 116 may calculate a weighted feature amount by applying a weighting value corresponding to a part of the face to the movement speed included in each feature amount of at least some of the feature points calculated by the first calculation unit 114.

[0040] The avatar generation unit 118 generates an avatar that moves in accordance with the facial movements of the user 40. For example, the avatar generation unit 118 generates a 3D model of the face of the user 40 using a depth information image included in the imaging data, generates texture data of the face of the user 40 using a captured image included in the imaging data, and performs a process of pasting the texture data onto the 3D model, thereby generating the avatar. By continuously performing such a process, it is possible to generate an avatar that moves in accordance with the facial movements of the user 40.

[0041] The avatar generation unit 118 further generates an avatar that emphasizes the movement of part or all of the face of the user 40. The avatar generation unit 118 generates an avatar based on the weighted feature amounts of each of the plurality of feature points calculated by the second calculation unit 116. The avatar generation unit 118 may generate an avatar in which each part changes by moving the feature points at a moving speed to which a weighted value is applied for each of the plurality of feature points for which the second calculation unit 116 has calculated the weighted feature amount.

[0042] For example, when generating an avatar that emphasizes the eye movement of user 40, second calculation unit 116 calculates weighted feature amounts by applying weighted values ​​corresponding to the eyes to the feature amounts of multiple feature points corresponding to the eyes of user 40, among multiple feature points on the face of user 40. Then, avatar generation unit 118 generates an avatar using the weighted feature amounts of the multiple feature points corresponding to the eyes of user 40. Avatar generation unit 118 generates an avatar in which parts of user 40 other than the eyes move in accordance with the movement of user 40's face, and the part of user 40's eyes moves at a movement speed increased by the weighted value. Increasing the movement speed results in an increased amount of movement, and the eye movement is emphasized.

[0043] The avatar generation unit 118 may generate an avatar using BlendShape. For example, the avatar generation unit 118 may generate an avatar whose facial expression changes by adjusting shape keys, based on an avatar generated by attaching texture data to a 3D model of the face of the user 40. The avatar generation unit 118 may then generate an avatar that moves in accordance with the movement of the face of the user 40 by adjusting the shape keys in accordance with the feature amounts calculated by the first calculation unit 114. The avatar generation unit 118 may also generate an avatar that emphasizes the movement of the user 40 by adjusting the shape keys in accordance with the weighted feature amounts calculated by the second calculation unit 116.

[0044] It should be noted that the avatar generation unit 118 does not need to generate a facial model using all depth data of the user 40's face. The avatar generation unit 118 may generate an avatar using one or more pieces of depth data selected from the facial depth data: depth data corresponding to the eyes, depth data corresponding to the nose, and depth data corresponding to the mouth. The depth data corresponding to the eyes of the user 40 may be depth data of one or more feature points around the eyes, and the number of such feature points is not particularly limited. The one or more feature points around the eyes of the user 40 may include feature points corresponding to the eyebrows. The depth data corresponding to the nose of the user 40 may be depth data of one or more feature points around the nose, and the number of such feature points is not particularly limited. The depth data corresponding to the mouth of the user 40 may be depth data of one or more feature points around the mouth, and the number of such feature points is not particularly limited. The one or more feature points around the mouth of the user 40 may include a feature point corresponding to the chin. By generating an avatar using data on at least one of the eyes, nose, and mouth, which are easy to recognize in terms of position and orientation and which are easy to express facial features, it is possible to reduce the amount of avatar data compared to avatars generated using data on the entire face while maintaining a person's features, thereby speeding up image processing.

[0045] The method for generating a 3D model of user 40's face from one or more of depth data corresponding to the eyes, the nose, and the mouth is not particularly limited. For example, avatar generation unit 118 may generate a 3D model of user 40's face by using a learning model that is machine-learned using as training data a set of various 3D face models and one or more of depth data corresponding to the eyes, the nose, and the mouth of the 3D models.

[0046] The mode management unit 120 manages modes related to avatar emphasis. Examples of modes include a normal mode, an eye emphasis mode that emphasizes the eyes of the user 40, a mouth emphasis mode that emphasizes the mouth of the user 40, and an expression emphasis mode that emphasizes the facial expression of the user 40.

[0047] The mode management unit 120 switches modes. The mode management unit 120 switches modes, for example, in accordance with an instruction from the user 40. When a normal mode and an eye emphasis mode are prepared as modes, the mode management unit 120 switches between the normal mode and the eye emphasis mode. When a normal mode and an expression emphasis mode are prepared as modes, the mode management unit 120 switches between the normal mode and the expression emphasis mode. When a normal mode and a mouth emphasis mode are prepared as modes, the mode management unit 120 switches between the normal mode and the mouth emphasis mode. When a normal mode and more than one of the eye emphasis mode, the mouth emphasis mode, and the expression emphasis mode are prepared, the mode management unit 120 switches between these modes.

[0048] When the normal mode is set, the avatar generation unit 118 may generate an avatar based on the imaging data acquired by the imaging data acquisition unit 112. That is, the avatar generation unit 118 may generate a 3D model of the face of the user 40 using the depth information image, generate texture data of the face of the user 40 using the captured image, and continuously perform a process of pasting the texture data onto the 3D model, thereby generating an avatar that moves in accordance with the movement of the face of the user 40. In the normal mode, no emphasis is performed.

[0049] When the eye emphasis mode is set, the second calculation unit 116 calculates weighted feature amounts by applying weights corresponding to the eyes to feature amounts of a plurality of feature points corresponding to the eyes of the user 40, among a plurality of feature points on the face of the user 40, and the avatar generation unit 118 generates an avatar using the weighted feature amounts of the plurality of feature points corresponding to the eyes of the user 40. The avatar generation unit 118 may generate an avatar by applying weighted feature amounts to a plurality of feature points corresponding to the eyes of an avatar generated based on imaging data.

[0050] When the mouth emphasis mode is set, the second calculation unit 116 calculates weighted feature amounts by applying weights corresponding to the mouth to the feature amounts of multiple feature points corresponding to the mouth of user 40 among multiple feature points on the face of user 40, and the avatar generation unit 118 generates an avatar using the weighted feature amounts of the multiple feature points corresponding to the mouth of user 40. The avatar generation unit 118 may generate an avatar by applying weighted feature amounts to the multiple feature points corresponding to the mouth of an avatar generated based on imaging data.

[0051] When the facial expression enhancement mode is set, the second calculation unit 116 may calculate weighted feature amounts by applying weights corresponding to facial parts to feature amounts of a plurality of feature points corresponding to the user's eyes, eyebrows, cheeks, and mouth among a plurality of feature points on the face of the user 40, and the avatar generation unit 118 may generate an avatar using the weighted feature amounts of the plurality of feature points corresponding to the eyes, eyebrows, cheeks, and mouth of the user 40. The avatar generation unit 118 may generate an avatar by applying weighted feature amounts to a plurality of feature points corresponding to the eyes, eyebrows, cheeks, and mouth of an avatar generated based on imaging data.

[0052] The range storage unit 122 stores the movable range of each of the multiple feature points on the face of the user 40. The movable range of a feature point may indicate the movable range of the feature point from a reference position on the face of the user 40. The movable range of each of the multiple feature points is set based on a range that does not look unnatural on a human face. For example, if the feature point of the mouth were to move to the position of the eyes, it would look unnatural on a human face, so the range within which the mouth can actually move on the face is set as the movable range of the feature point of the mouth. When the image processing device 100 uses BlendShape, the movable range of each of the multiple feature points may be set so as not to exceed the maximum value that can be expressed by the key value of the shape key.

[0053] The avatar generation unit 118 may generate an avatar by further using the movable ranges of each of the plurality of feature points stored in the range storage unit 122. The avatar generation unit 118 may generate an avatar in which each part changes by moving each of the plurality of feature points for which the weighted feature amount has been calculated by the second calculation unit 116, within the movable range stored in the range storage unit 122, at a moving speed to which a weight value has been applied. This makes it possible to prevent the face of user 40 from looking unnatural by emphasizing part or all of the face of user 40.

[0054] The display control unit 124 performs control to superimpose and display the avatar generated by the avatar generation unit 118 over the face of the user 40 in the imaging data. The display control unit 124 may perform control to generate a display image in which the avatar generated by the avatar generation unit 118 is placed over the face of the user 40 in the imaging data, and display the display image.

[0055] 4 is an explanatory diagram for explaining eye enhancement. Avatar 400 shows a state in which the eyes are not enhanced, and avatar 402 shows a state in which the eyes are enhanced.

[0056] When user 40 has a facial expression similar to avatar 400, in normal mode, user 40 is displayed as avatar 400, and in eye-emphasis mode, user 40 is displayed as avatar 402.

[0057] In this way, by making it possible to emphasize the eyes of an avatar, it is possible to convey one's intentions more strongly to the other person and to provide the other person with an interesting experience.

[0058] 5 is an explanatory diagram for explaining facial expression enhancement. Avatar 300 shows a state in which the facial expression is not enhanced, avatar 302 shows a state in which the facial expression is weakly enhanced, and avatar 304 shows a state in which the facial expression is strongly enhanced.

[0059] When user 40 has an expression like avatar 300, in normal mode, user 40 is displayed as avatar 300, in expression emphasis mode, with a low weighted value set, user 40 is displayed as avatar 302, and in expression emphasis mode, with a high weighted value set, user 40 is displayed as avatar 304.

[0060] In this way, by making it possible to emphasize the facial expression of the avatar, it is possible to make a good impression on the other person. Also, by making the weighting value changeable, it is possible to change the degree of emphasis on the facial expression as desired by the user 40.

[0061] 6 shows an example of the flow of processing by the image processing device 100. Here, an example will be described in which a normal mode, an eye emphasis mode, and an expression emphasis mode are prepared.

[0062] In S102, the first calculation unit 114 analyzes the imaging data acquired by the imaging data acquisition unit 112, and calculates the feature amount of each of the plurality of feature points on the face of the user 40.

[0063] In S104, the second calculation unit 116 determines whether or not the eye emphasis mode is set. If the eye emphasis mode is set, the process proceeds to S106, and if not, the process proceeds to S108. In S106, the second calculation unit 116 calculates weighted feature amounts by applying weights corresponding to the eyes to the feature amounts of each of the plurality of feature points corresponding to the eyes of the user 40.

[0064] In S108, the second calculation unit 116 determines whether or not the facial expression enhancement mode is set. If the facial expression enhancement mode is set, the process proceeds to S110, and if not, the process proceeds to S116. In S110, the second calculation unit 116 calculates weighted feature amounts by applying weights corresponding to each part to the feature amounts of a plurality of feature points corresponding to the eyes, eyebrows, cheeks, and mouth of the user 40.

[0065] In S112, the avatar generation unit 118 determines whether the destination of a feature point whose weighted feature amount was calculated in S106 or S110 will exceed the movable range when the weighted value is applied. If there is a feature point whose destination exceeds the movable range, the process proceeds to S114. In S114, the avatar generation unit 118 changes the destination of the feature point whose destination exceeds the movable range to within the movable range. For example, the avatar generation unit 118 changes the destination of the feature point whose destination exceeds the movable range to the intersection of the straight line connecting the pre-movement position and the destination position with the end of the movable range.

[0066] In S116, the avatar generation unit 118 generates an avatar. If the determination in S108 is NO, that is, if the normal mode is set, the second calculation unit 116 generates an avatar based on the imaging data. For example, the avatar generation unit 118 generates a 3D model of the face of the user 40 using a depth information image included in the imaging data, generates texture data of the face of the user 40 using a captured image included in the imaging data, and performs a process of pasting the texture data onto the 3D model, thereby generating the avatar.

[0067] In the eye emphasis mode, the avatar generation unit 118 generates an avatar using the weighted feature amounts of the multiple feature points corresponding to the eyes of the user 40 calculated in S106. In the facial expression emphasis mode, the avatar generation unit 118 generates an avatar using the weighted feature amounts of the multiple feature points corresponding to the eyes, eyebrows, cheeks, and mouth of the user 40 calculated in S110.

[0068] In S118, the display control unit 124 generates a display image in which the avatar generated in S116 is placed in the area of ​​the face of the user 40 in the imaging data, and performs control so that the display image is displayed.

[0069] If the avatar generation process is to be ended by an instruction from the user 40 or the like (YES in S120), the process ends, and if not (NO in S120), the process returns to S102 and continues.

[0070] 7 schematically illustrates an example of the hardware configuration of a computer 1200 functioning as the image processing device 100 or the communication device 200. A program installed on the computer 1200 can cause the computer 1200 to function as one or more "units" of the device according to the present embodiment, or can cause the computer 1200 to perform operations associated with the device according to the present embodiment or one or more "units," and / or can cause the computer 1200 to perform a process according to the present embodiment or steps of the process. Such a program can be executed by the CPU 1212 to cause the computer 1200 to perform specific operations associated with some or all of the blocks in the flowcharts and block diagrams described herein.

[0071] The computer 1200 according to this embodiment includes a CPU 1212, a RAM 1214, and a graphics controller 1216, which are interconnected by a host controller 1210. The computer 1200 also includes input / output units such as a communications interface 1222, a storage device 1224, a DVD drive, and an IC card drive, which are connected to the host controller 1210 via an input / output controller 1220. The DVD drive may be a DVD-ROM drive, a DVD-RAM drive, or the like. The storage device 1224 may be a hard disk drive, a solid-state drive, or the like. The computer 1200 also includes a ROM 1230 and legacy input / output units such as a keyboard, which are connected to the input / output controller 1220 via an input / output chip 1240.

[0072] The CPU 1212 operates according to programs stored in the ROM 1230 and the RAM 1214, thereby controlling each unit. The graphics controller 1216 acquires image data generated by the CPU 1212 into a frame buffer or the like provided in the RAM 1214 or into the graphics controller itself, and causes the image data to be displayed on the display device 1218.

[0073] The communication interface 1222 communicates with other electronic devices via a network. The storage device 1224 stores programs and data used by the CPU 1212 in the computer 1200. The DVD drive reads programs or data from a DVD-ROM or the like and provides them to the storage device 1224. The IC card drive reads programs and data from an IC card and / or writes programs and data to an IC card.

[0074] The ROM 1230 stores therein a boot program or the like that is executed by the computer 1200 upon activation, and / or programs that depend on the hardware of the computer 1200. The input / output chip 1240 may also connect various input / output units to the input / output controller 1220 via a USB port, a parallel port, a serial port, a keyboard port, a mouse port, etc.

[0075] The programs are provided by a computer-readable storage medium such as a DVD-ROM or an IC card. The programs are read from the computer-readable storage medium, installed in the storage device 1224, RAM 1214, or ROM 1230, which are also examples of computer-readable storage media, and executed by the CPU 1212. Information processing described in these programs is read by the computer 1200, and causes cooperation between the programs and the various types of hardware resources described above. An apparatus or method may be configured by implementing operations or processing of information in accordance with the use of the computer 1200.

[0076] For example, when communication is performed between the computer 1200 and an external device, the CPU 1212 may execute a communication program loaded into the RAM 1214 and instruct the communication interface 1222 to perform communication processing based on the processing described in the communication program. Under the control of the CPU 1212, the communication interface 1222 reads transmission data stored in a transmission buffer area provided in the RAM 1214, the storage device 1224, a DVD-ROM, or a recording medium such as an IC card, and transmits the read transmission data to the network, or writes reception data received from the network to a reception buffer area or the like provided on the recording medium.

[0077] Furthermore, the CPU 1212 may cause all or a necessary portion of a file or database stored in an external recording medium such as the storage device 1224, a DVD drive (DVD-ROM), an IC card, etc. to be read into the RAM 1214, and may perform various types of processing on the data on the RAM 1214. The CPU 1212 may then write back the processed data to the external recording medium.

[0078] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and may undergo information processing. The CPU 1212 may perform various types of processing on data read from the RAM 1214, including various types of operations, information processing, conditional judgment, conditional branching, unconditional branching, information search / replacement, etc., as described throughout this disclosure and specified by the instruction sequences of the programs, and write the results back to the RAM 1214. The CPU 1212 may also search for information in a file, database, etc. on the recording medium. For example, if multiple entries, each having an attribute value of a first attribute associated with an attribute value of a second attribute, are stored on the recording medium, the CPU 1212 may search for an entry whose attribute value of the first attribute matches a specified condition from among the multiple entries, read the attribute value of the second attribute stored in the entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.

[0079] The above-described programs or software modules may be stored in a computer-readable storage medium on or near the computer 1200. A recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can also be used as a computer-readable storage medium, thereby providing the programs to the computer 1200 via the network.

[0080] The blocks in the flowcharts and block diagrams in the present embodiments may represent stages of a process in which an operation is performed or "parts" of an apparatus responsible for performing the operation. Particular stages and "parts" may be implemented by dedicated circuitry, programmable circuitry provided with computer-readable instructions stored on a computer-readable storage medium, and / or a processor provided with computer-readable instructions stored on a computer-readable storage medium. The dedicated circuitry may include digital and / or analog hardware circuits, including integrated circuits (ICs) and / or discrete circuits. The programmable circuitry may include reconfigurable hardware circuits, such as field programmable gate arrays (FPGAs) and programmable logic arrays (PLAs), including AND, OR, XOR, NAND, NOR, and other logical operations, flip-flops, registers, and memory elements.

[0081] A computer-readable storage medium may include any tangible device capable of storing instructions that are executed by an appropriate device, such that a computer-readable storage medium having instructions stored thereon comprises an article of manufacture, including instructions that can be executed to create means for performing the operations specified in the flowcharts or block diagrams. Examples of computer-readable storage media may include electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, etc. More specific examples of computer-readable storage media may include floppy disks, diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), electrically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), Blu-ray disc, memory stick, integrated circuit card, etc.

[0082] The computer readable instructions may include either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk®, JAVA®, C++, etc., and conventional procedural programming languages ​​such as the “C” programming language or similar programming languages.

[0083] Computer-readable instructions may be provided locally or over a wide area network (WAN) such as a local area network (LAN), the Internet, etc. to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, or programmable circuitry, such that the processor or programmable circuitry executes the computer-readable instructions to generate means for performing the operations specified in the flowcharts or block diagrams. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc.

[0084] Although the present invention has been described above using the embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications and improvements can be made to the above embodiments. It is clear from the claims that such modifications and improvements can also be included within the technical scope of the present invention.

[0085] It should be noted that the execution order of each process, such as operations, procedures, steps, and stages, in the devices, systems, programs, and methods shown in the claims, specifications, and drawings is not specifically stated as "before," "prior to," etc., and that the processes can be performed in any order unless the output of a previous process is used in a later process. Even if the operational flow in the claims, specifications, and drawings is described using "first," "next," etc. for convenience, this does not mean that the processes must be performed in this order. [Explanation of symbols]

[0086] 10 Avatar display system, 20 Network, 40 User, 100 Image processing device, 102 Imaging unit, 112 Imaging data acquisition unit, 114 First calculation unit, 116 Second calculation unit, 118 Avatar generation unit, 120 Mode management unit, 122 Range memory unit, 124 Display control unit, 200 Communication device, 202 Imaging unit, 300, 302, 304, 400, 402 Avatar, 1200 Computer, 1210 Host controller, 1212 CPU, 1214 RAM, 1216 Graphics controller, 1218 Display device, 1220 Input / output controller, 1222 Communication interface, 1224 Storage device, 1230 ROM, 1240 Input / output chip

Claims

1. a first calculation unit that analyzes image data of a user's face and calculates feature amounts including a moving direction and a moving speed of each of a plurality of feature points of the user's face; a second calculation unit that calculates weighted feature amounts by applying weights corresponding to parts of the face to the feature amounts of at least some of the feature points of the user's face; an avatar generation unit that generates an avatar that moves in accordance with a movement of the user's face, the avatar generation unit generating the avatar based on the weighted feature amounts of the plurality of feature points; a display control unit that controls the avatar generated by the avatar generation unit to be superimposed on a face portion of the user in the imaging data; An image processing device comprising:

2. the second calculation unit calculates the weighted feature amount by applying a weight corresponding to a part of the face to the moving speed included in the feature amount of each of the at least some of the plurality of feature points; The image processing device according to claim 1 , wherein the avatar generation unit generates the avatar in which each part changes by moving the avatar according to the moving speed to which the weight value is applied for each of the at least some of the plurality of feature points.

3. a range storage unit that stores movable ranges of the plurality of feature points of the user's face; Equipped with 3. The image processing device according to claim 2, wherein the avatar generation unit generates the avatar in which each part changes by moving at least some of the plurality of feature points within the movable range at the movement speed to which the weighted value is applied.

4. a mode management unit that switches between a normal mode and an eye enhancement mode that enhances the user's eyes; Equipped with When the normal mode is set, the avatar generation unit generates the avatar based on the imaging data, 4. The image processing device according to claim 1, wherein when the eye emphasis mode is set, the second calculation unit calculates the weighted feature amounts by applying weights corresponding to the eyes to the feature amounts of each of a plurality of feature points corresponding to the user's eyes among the plurality of feature points of the user's face, and the avatar generation unit generates the avatar using the weighted feature amounts of the plurality of feature points corresponding to the user's eyes.

5. the mode management unit switches between the normal mode, the eye emphasis mode, and an expression emphasis mode that emphasizes an expression of the user; 5. The image processing device according to claim 4, wherein, when the facial expression emphasis mode is set, the second calculation unit calculates weighted feature amounts by applying weights corresponding to parts of the face to the feature amounts of each of a plurality of feature points corresponding to the user's eyes, eyebrows, cheeks, and mouth among the plurality of feature points of the user's face, and the avatar generation unit generates the avatar using the weighted feature amounts of the plurality of feature points corresponding to the user's eyes, eyebrows, cheeks, and mouth.

6. a mode management unit that switches between a normal mode and an expression enhancement mode that enhances the user's expression; Equipped with When the normal mode is set, the avatar generation unit generates the avatar based on the imaging data, 4. The image processing device according to claim 1, wherein, when the facial expression enhancement mode is set, the second calculation unit calculates weighted feature amounts by applying weights corresponding to parts of the face to the feature amounts of each of a plurality of feature points corresponding to the user's eyes, eyebrows, cheeks, and mouth among the plurality of feature points of the user's face, and the avatar generation unit generates the avatar based on the weighted feature amounts of the plurality of feature points corresponding to the user's eyes, eyebrows, cheeks, and mouth and the feature amounts of a plurality of feature points other than the user's eyes, eyebrows, cheeks, and mouth.

7. A program for causing a computer to function as the image processing device according to any one of claims 1 to 3.

8. 1. A computer-implemented image processing method comprising: a first calculation step of analyzing image data of a user's face and calculating feature amounts including a movement direction and a movement speed of each of a plurality of feature points of the user's face; a second calculation step of calculating weighted feature amounts by applying weights corresponding to parts of the face to the feature amounts of at least some of the feature points of the user's face; an avatar generation step of generating an avatar that moves in accordance with a movement of the user's face, the avatar being generated based on the weighted feature amounts of the plurality of feature points; a display control step of controlling the avatar generated by the avatar generation unit to be superimposed and displayed on a portion of the user's face in the imaging data; An image processing method comprising:

Citation Information

Patent Citations

  • Wearable terminal device and program

    JP2016126500A

  • Communication support system and communication support method

    JP2023180598A

  • Program, method, and information processing device

    JP2024006906A

  • Avatar control device

    JP2024008130A