Information processing device, information processing method, and program
The information processing apparatus addresses discomfort in user expression displays by gradually adjusting user expressions in mixed reality environments, providing a smoother and more natural transition.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Existing technologies for displaying user expressions, such as avatars, often cause discomfort due to abrupt or unnatural transitions.
An information processing apparatus that processes user expressions by determining whether they are processing targets and gradually adjusting them based on past and target expressions, using a combination of real and virtual images to create a mixed reality environment.
This approach reduces user discomfort by ensuring smoother and more natural transitions between user expressions, enhancing the overall experience.
Smart Images

Figure 2026074524000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus, and more particularly to an information processing apparatus that processes user expressions.
Background Art
[0002] Techniques for controlling the display of user avatars are known (Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] It is desired to output user expressions (such as avatars) with less discomfort.
Means for Solving the Problems
[0005] The information processing apparatus of the present disclosure is an information processing apparatus that processes a user expression in at least one of a CG and a real photographed image, and includes a determination unit that determines whether the user expression is a processing target, and a processing unit that processes the user expression based on the determination result of the determination unit. The processing unit processes the user expression based on the past user expression and the target user expression.
Effects of the Invention
[0006] According to the information processing apparatus of the present disclosure, a user expression with less discomfort can be output.
Brief Description of the Drawings
[0007] [Figure 1] It is a block diagram showing the configuration of an information processing system. [Figure 2] This flowchart shows the controls used when processing user expressions. [Figure 3] This is a settings screen for when to start and stop processing the user's expression. [Figure 4] This is a conceptual diagram illustrating how to process user expressions. [Figure 5] This is a conceptual diagram illustrating how to process user expressions. [Modes for carrying out the invention]
[0008] (First embodiment) (composition) Figure 1 is a block diagram showing an example of the system configuration according to the first embodiment. The system according to the first embodiment is an image processing system for presenting a mixed reality (MR) space, which fuses the real space and the virtual space, to the system user. In the first embodiment, we will describe a case in which the MR space is presented to the user by displaying a composite image that combines an image of the virtual space drawn by computer graphics (CG) and an image of the real space (real-life image).
[0009] The system according to the first embodiment includes a display device 101, an information processing device 102, and an operating device 103.
[0010] The information processing device 102 is composed of the following components. The functions of each component are configured, for example, by one or more CPUs (central processing units) that function as the control unit of the information processing device 102 executing programs. The components of the information processing device 102 may be composed of integrated circuits or the like, as long as they perform similar functions.
[0011] The information processing device 102 combines an image of the real world captured from the display device 101 with an image of the virtual world generated in the information processing device 102 to generate a composite image. The information processing device 102 outputs the composite image as a mixed reality image (MR image) to the display device 101. The first embodiment relates to an information processing device that displays an image of the virtual world, and the system according to the first embodiment is not limited to an MR system that displays an MR image (an image created by combining an image of the real world with an image of the virtual world). In other words, the system according to the first embodiment may be a so-called XR system, such as a VR (virtual reality) system that presents only an image of the virtual world to the user, or an AR (augmented reality) system that presents an image of the virtual world to the user by making the real world transparent.
[0012] The display device 101 includes a recording unit 104, an audio output unit 105, an imaging unit 106, and a display unit 107.
[0013] The recording unit 104 captures sounds such as sounds from the vicinity of the display unit 101 and the user's own voice as audio data and outputs it to the information processing device 102. Depending on the application, the recording unit 104 may be a directional microphone or a movable lavalier microphone, etc. Multiple such recording devices may also be provided.
[0014] The audio output unit 105 outputs audio data output from the information processing device 102. The audio output unit 105 may be a speaker, earphones, or headphones. Furthermore, there may be multiple such audio output devices, and each audio output device may be wired or wireless. Alternatively, for example, a device that includes both a recording unit 104 and an audio output unit 105, such as wireless earphones with a microphone function, may communicate with the display device 101 or the information processing device 102.
[0015] The imaging unit 106 continuously images the real space in a time series and outputs the captured images of the real space (captured images) to the information processing device 102. The imaging unit 106 may include a stereo camera consisting of two cameras fixed to each other so as to be able to capture the real space in the direction of the user's line of sight. In addition to a camera that captures the outside of the display device 101 as a reference, a camera that captures the inside may also be provided so as to be able to acquire the user's expressions such as facial expressions and body language. Furthermore, a camera other than the display device 101 may be used in conjunction to capture the user's expressions. For example, a camera fixed to a tripod or the wall of a building may be used in conjunction. In that case, the captured images from the camera other than the display device 101 may be transmitted to the information processing device 102 via the display device 101, or they may be transmitted directly to the information processing device 102. The user's expressions are transmitted to the information processing device 102 as captured images and stored in the data storage unit 113 via the control unit 111. In the above, the means for acquiring the user's expressions were described as imaging means, but other means may be used as long as the user's expressions can be acquired. For example, to obtain user expressions, multiple sensors may be placed on the user's body, including their face and limbs, and the user's expressions may be obtained by acquiring the values from these sensors.
[0016] The display unit 107 displays the MR image output from the information processing device 102. The display unit 107 may include two displays arranged to correspond to the user's left and right eyes, respectively. In this case, the display corresponding to the user's left eye displays the MR image for the left eye, and the display corresponding to the user's right eye displays the MR image for the right eye.
[0017] The display device 101 is, for example, a head-mounted display device (HMD). However, the display device 101 is not limited to an HMD; it may also be a handheld display (HHD). An HHD is a handheld display. In other words, the display device 101 may be a display that the user holds in their hand and looks through like binoculars to observe images. Alternatively, the display device 101 may be a display terminal such as a tablet or a smartphone.
[0018] The operating device 103 has a vibration unit 117 and an operation instruction unit 118. The operating device 103 is, for example, an operating device for a video game that can acquire values according to the direction of the tilt of a stick, values indicating the pressed state of buttons, etc. However, the operating device 103 is not limited to an operating device held by the user's hand like an operating device for a video game, and may be an operating device worn on the user's body or hand. The operating device 103 may be, for example, a ring-shaped (ring-shaped) device that can be worn on the user's finger. Also, there may be a plurality of buttons, or an optical track pad (hereinafter referred to as "OTP") may be incorporated. When an OTP is incorporated, the user can place a finger on the OTP and rub it in an arbitrary direction to align the pointer with a desired item. Then, the user can perform a determination operation such as determining the selection of a virtual object (virtual object) such as a menu item by pressing the button of the OTP. Also, while a virtual object is selected, it is possible to move the object by pointing to another location, twisting the selected arm or finger, or performing a predetermined operation.
[0019] The vibration unit 117 vibrates in accordance with a vibration instruction from the vibration instruction unit 116 to give the user a vibration sensation. In order to give a vibration sensation, it is desirable to use a vibration element with a short response time. For example, a piezo element can be cited, but other vibration elements such as a linear vibrator (LRA) or an eccentric motor (ERM) may also be used.
[0020] The operation instruction unit 118 outputs a value according to the user's operation (a value according to the direction of the tilt of a stick, a value indicating the pressed state of a button) as operation information to the operation processing unit 115. Here, the operations that can be considered include the selection, determination, and movement of virtual objects. Also, the movement of virtual objects may include not only simple parallel movement but also various affine transformations such as rotation, enlargement, and reduction.
[0021] In addition, the operation device 103 may have a position and orientation calculation unit (not shown). The position and orientation calculation unit calculates (acquires) the position and orientation (position and posture) of the operation device 103 in the world coordinate system. Then, the position and orientation calculation unit outputs the calculated position and orientation information to the operation processing unit 115. The position and orientation calculation unit has, for example, sensors (such as an angular velocity sensor, an acceleration sensor, or a geomagnetic sensor) for calculating the position and orientation of the operation device 103. Note that the position and orientation calculation unit may have a plurality of sensors. Also, the position and orientation calculation unit may include an imaging unit and an imaging processing unit, and perform Simultaneous Localiation and Mapping (SLAM) processing based on the feature points of the captured image to calculate the position and orientation of the operation device 103. Further, the position and orientation calculation unit may calculate the position and orientation in conjunction with an optical sensor installed in the real space. Note that the position and orientation calculation unit may calculate only one of the values of the position component and the orientation component instead of both values. Also, the position and orientation calculation unit may output values (values output from sensors), images, etc. for calculating the position and orientation of the operation device 103 to the information processing device 102, and a position and orientation calculation unit (not shown) of the information processing device 102 may calculate the position and orientation of the operation device 103.
[0022] The information processing device 102 and the display device 101 are connected so as to be able to communicate with each other. The information processing device 102 and the operation device 103 are connected so as to be able to communicate with each other. Note that the data communication may be wired communication or wireless communication. Also, the display device 101 may have the information processing device 102 inside. When using a camera separate from the display device 101, the camera separate from the display device 101 and the information processing device 102 are connected so as to be able to communicate with each other. Or the camera separate from the display device 101 and the display device 101 are connected so as to be able to communicate with each other.
[0023] The information processing device 102 includes a recording processing unit 108, an audio output instruction unit 109, an image synthesis unit 110, a control unit 111, an image generation unit 112, a data storage unit 113, a processing target discrimination unit 114, an operation processing unit 115, and a vibration instruction unit 116. It may also include a position and orientation calculation unit (not shown).
[0024] The position and orientation calculation unit (not shown) of the information processing device 102 calculates the position and orientation of the imaging unit 106 in the world coordinate system. Specifically, the position and orientation calculation unit extracts markers assigned to the world coordinate system from the image of real space captured by the imaging unit 106. Then, the position and orientation calculation unit calculates the position and orientation of the imaging unit 106 in the world coordinate system based on the position and orientation of the extracted markers. The position and orientation calculation unit then stores the calculated position and orientation information of the imaging unit 106 (position and orientation information) in the data storage unit 113 via the control unit 111.
[0025] Furthermore, the position and orientation calculation unit of the information processing device 102 calculates the position and orientation of the operating device 103 in the world coordinate system using the position and orientation information of the operating device 103 obtained from the position and orientation calculation unit of the operating device 103 or images of the real space captured by the imaging unit 106. At this time, depending on the calculation method of the position and orientation calculation unit of the operating device 103 (not shown), a difference (error) may occur between the calculated position and orientation of the operating device 103 and the actual position and orientation. For example, in a method that calculates the position and orientation using a combination of an angular velocity sensor, an acceleration sensor, and a geomagnetic sensor, the cumulative error of each sensor may result in a calculated position and orientation that is an error relative to the actual position and orientation. Alternatively, this method may not be able to calculate the position and orientation of the operating device 103. Also, in a method that calculates the position and orientation using an optical sensor installed in the real space, the optical sensor may be obscured by other real objects, resulting in a calculated position and orientation that is an error relative to the actual position and orientation. Alternatively, this method may also not be able to calculate the position and orientation. In such a case, the position and orientation calculation unit of the information processing device 102 can accurately calculate the position and orientation of the operating device 103 based on the position and orientation of the markers by extracting the markers attached to the operating device 103 from the image of the real space captured by the imaging unit 106. At this time, the position and orientation calculation unit of the information processing device 102 (not shown) may use all or part of the position and orientation calculation result obtained from the position and orientation calculation unit of the operating device 103 (not shown). However, the calculation method of the position and orientation calculation unit of the information processing device 102 (not shown) is not limited to the method using markers, but may also be a method calculated by SLAM processing. Alternatively, the orientation may be detected by the orientation sensor unit of the display device 101 (not shown). The orientation sensor unit of the display device 101 (not shown) has an inertial measurement unit (IMU) and outputs orientation information of the display device 101 (attitude information) to the information processing device 102.
[0026] The position and orientation calculation unit (not shown) of the information processing device 102 stores the calculated position and orientation information of the operating device 103 in the data storage unit 113 via the control unit 111.
[0027] If the operating device 103 is not visible in the image of the real space captured by the imaging unit 106, the position and orientation calculation unit of the information processing device 102 calculates the position and orientation of the operating device 103 based on the value obtained from the position and orientation calculation unit of the operating device 103. At this time, a position and orientation calculation unit (not shown) of the information processing device 102 may calculate the position and orientation of the operating device 103 based on the value obtained from the position and orientation calculation unit (not shown) of the operating device 103 and other information. For example, if the operating device 103 is not visible in the image of the real space, and the operating device 103 is equipped with an acceleration sensor or the like, the position and orientation calculation unit (not shown) of the information processing device 102 may calculate the position and orientation of the operating device 103 based on the detection result of the acceleration sensor or the like. For example, the position and orientation calculation unit (not shown) of the information processing device 102 adds the position of the operating device 103 at a past point in time when the operating device 103 was visible in the image of the real space, and the amount of movement of the operating device 103 from that point in time calculated from the acceleration. As a result, the position and orientation calculation unit (not shown) of the information processing device 102 may calculate the current position and orientation of the operating device 103. Furthermore, if the operating device 103 is not visible in the image of real space, the position and orientation calculation unit (not shown) of the information processing device 102 may only detect the orientation of the operating device 103, or it may not be able to calculate the position and orientation of the operating device 103.
[0028] The control unit 111 controls the entire information processing device 102. For example, the control unit 111 controls the position and orientation of the UI (graphics) displayed on the display unit 107 based on the position and orientation information of the operating device 103 stored in the data storage unit 113. The position of the UI is indicated by three-dimensional coordinate information following a three-axis orthogonal coordinate system, such as the X, Y, and Z axes. Alternatively, the three-dimensional coordinate information could follow a polar coordinate system. If the UI is a virtual ray emitted from the user's hand, the position of the UI is, for example, the start or end point of the ray. The orientation of the UI corresponds to the orientation of the UI in a three-dimensional virtual space. If the UI is a ray, the orientation of the UI corresponds to, for example, the direction in which the ray extends.
[0029] The above describes the case where the control unit 111 controls based on the position and orientation information of the operating device 103, but control may also be performed based on the position and orientation information of the imaging unit 106 stored in the data storage unit 113. Control may also be performed based on either one or both of the position and orientation information.
[0030] The user's representation (user representation) is acquired from the data storage unit 113 and transmitted to the processing target discrimination unit 114. Control is then performed according to the discrimination result of the processing target discrimination unit 114. If necessary, the image generation unit 112 processes the user's representation and constructs a virtual space that reflects the processed user representation, and stores the processed user representation in the data storage unit 113. At this time, the processed user representation may be stored in the data storage unit 113 by updating the acquired user representation stored in the data storage unit 113. Alternatively, if the image generation unit 112 processes the user's representation while referring to multiple histories of user representations, multiple user representations may be stored in the data storage unit 113 in a manner that indicates the storage order. In that case, to avoid storing an inexhaustible number of user representations, it is conceivable to automatically delete user representations after a predetermined time has elapsed since storage, or to overwrite older user representations when a predetermined number of them have been stored, similar to a ring buffer.
[0031] Furthermore, the system acquires sound source data that the user should be able to hear from the data storage unit 113, synthesizes the sound source data as needed, and outputs sound based on the sound source data from the sound output unit 105 of the display device 101 via the voice output instruction unit 109. The control unit 111 may also control the entire information processing device 102 according to voice commands. In that case, it analyzes the voice data acquired from the recording unit 104 via the recording processing unit 108, and if it is a voice command, it performs control according to the command content.
[0032] The operation processing unit 115 receives operation instructions from the operation instruction unit 118, and, if necessary, causes the image generation unit 112 to construct a virtual space that reflects the operation on the virtual object. This enables the user to operate the virtual object. The virtual object to be operated does not have to be a single object; for example, it may be a group composed of multiple objects. In the case of control operations such as confirming or canceling, the control unit 111 performs the control according to the operation instructions.
[0033] Vibration instructions are transmitted to the vibration unit 117 via the vibration instruction unit 116, providing the user with a vibration sensation.
[0034] The processing target discrimination unit 114 acquires the user's expression from the data storage unit 113 and determines whether the acquired expression is a target for processing. For example, if an expression is inappropriate to record or communicate to others, it may be determined to be a target for processing. User expressions include voice and nonverbal expressions, and these may be combined. When making the determination, expressions to be processed and those not to be processed may be registered in advance, or a machine learning model may be created and the determination may be made using that model. For example, when determining whether to process an expression, if an angry face is to be processed, suppose the user changes from a normal expression to an angry face and then back to a normal expression. In this case, when the expression is normal, it will be determined not to be a target for processing, then when the expression becomes angry, it will be determined to be a target for processing, and then it will be determined not to be a target for processing again.
[0035] The operation processing unit 115 modifies the data related to each virtual object stored in the data storage unit 113 via the control unit 111 based on the operation information obtained from the operation instruction unit 118, and stores it in the data storage unit 113. Then, via the control unit 111, it transmits the data related to each virtual object stored in the data storage unit 113 and the position and orientation information of the imaging unit 106 to the image generation unit 112. The virtual objects are then manipulated by having the image generation unit 112 construct a virtual space that takes the position and orientation into account.
[0036] The image generation unit 112 constructs a virtual space based on the virtual space data stored in the data storage unit 113 via the control unit 111. The virtual space data includes data related to each virtual object that constitutes the virtual space, data related to light sources that illuminate the virtual space, and user representations. The user representation in the virtual space constructed by the image generation unit 112 according to the instructions of the control unit 111 may be represented with an appearance similar to that of a real user, or it may be represented after processing based on the appearance of a real user. For example, it is conceivable to change skin color including skin beautification, change the appearance of hair, eyelashes, eyebrows, and eyes, change clothing, or change height and proportions. Such processing is performed while reflecting the user's representation. Alternatively, instead of basing it on the appearance of a real user, it may be based on an appearance similar to a 3D avatar, and the user's representation may be reflected. The method of reflecting the user's representation is also according to the instructions of the control unit 111. More details will be described later. The image generation unit 112 then obtains the position and orientation information of the imaging unit 106, calculated by the position and orientation calculation unit (not shown) of the information processing device 102, from the data storage unit 113 via the control unit 111. The image generation unit 112 also obtains the position and orientation information of the UI controlled by the control unit 111 from the data storage unit 113 via the control unit 111. The image generation unit 112 then generates a virtual space image (virtual space image) corresponding to the position and orientation of the imaging unit 106.
[0037] Regarding the technology for generating images in a virtual space according to the position and orientation of the imaging unit 106, it is a well-known technology, so a detailed explanation will be omitted.
[0038] The image synthesis unit 110 generates an MR image by combining the image of the virtual space generated by the image generation unit 112 with the image of the real space captured by the imaging unit 106. At this time, the image of the virtual space generated by the image generation unit 112 may be an image representing the entire virtual space or an image representing a part of the virtual space. The image synthesis unit 110 may also generate the MR image by performing an affine transformation on the image, or by assigning it to a parametric surface. The image synthesis unit 110 then outputs the generated MR image to the display unit 107. Although this example illustrates displaying it to the user, it may also be transmitted externally instead of being displayed immediately, or it may be stored as a 3D video on a storage medium for later display.
[0039] As described above, the data storage unit 113 stores various types of information. The data storage unit 113 includes RAM and a hard disk drive device. In addition to the information described above as information to be stored in the data storage unit 113, the data storage unit 113 also stores information that is described as known information in the first embodiment.
[0040] Furthermore, in the system according to the first embodiment, the user's hand may be used instead of the operation instruction unit 118 of the operating device 103. In this case, the processing target determination unit 114 recognizes the user's hand movements (gestures) from the image of the real space captured by the imaging unit 106 and stores the recognized gestures in the data storage unit 113. Based on the recognized gestures, the processing target determination unit 114 determines whether or not the object is a target for processing. If processing is not required, the operation processing unit 115 modifies the data related to each virtual object stored in the data storage unit 113 via the control unit 111 and stores it in the data storage unit 113. Subsequently, the data related to each virtual object stored in the data storage unit 113 and the position and orientation information of the imaging unit 106 are transmitted to the image generation unit 112 via the control unit 111. The virtual objects are then manipulated by having the image generation unit 112 construct a virtual space that takes position and orientation into account.
[0041] (Processing of user-generated content) Figure 2 is a flowchart for processing the user's expression. Here, we will explain the control process when processing the user's facial expression. Note that the operation shown in Figure 2 is realized, for example, by the CPU of the information processing device 102 executing a program stored in ROM.
[0042] In step S201, the machining status, which indicates the machining state, is initialized. The machining status is stored in the data storage unit 113, and the control unit 111 updates its value. The initial value is set to indicate no machining. For example, if the value indicating no machining is set to 0 and the value indicating machining is set to 1, the machining status stored in the data storage unit 113 is updated to 0 to perform the initialization.
[0043] In step S202, the system waits until a user representation is detected. The determination of whether a user representation has been detected is made by checking whether a user representation has been acquired by the imaging unit 106 or other means for acquiring user representations (not shown) and whether the control unit 111 has received the user representation. When the control unit 111 receives a user representation, it stores the received user representation in the data storage unit 113.
[0044] In step S203, it is determined whether the user's representation is a target for processing. The control unit 111 transmits the user's representation stored in the data storage unit 113 to the processing target determination unit 114 along with the processing target determination instruction. Upon receiving the processing target determination instruction and the user's representation, the processing target determination unit 114 determines whether the received user's representation is a target for processing using the method described above, and transmits the determination result to the control unit 111. If there are representations from multiple users, the processing target determination unit 114 may determine them all together and transmit the determination results to the control unit 111 all at once, or it may transmit the determination result to the control unit 111 for each user's representation. Considering the data traffic within the control unit 111, the former is preferable. In the former case, for example, when transmitting the user's representation from the control unit 111 to the processing target determination unit 114, multiple user representations may be transmitted as actual data. Alternatively, index values or pointers for identifying multiple user representations stored in the data storage unit 113 may be transmitted in data format such as an array or list. When the processing target discrimination unit 114 transmits multiple discrimination results to the control unit 111, it is preferable to transmit the multiple discrimination results in a data format such as an array or list. After the image generation unit 112 processes the user's expression, the control unit 111 instructs the image synthesis unit 110 to synthesize the image of the virtual space generated by the image generation unit 112 with the image of the real space captured by the imaging unit 106 to generate an MR image. The image synthesis unit 110 then outputs the generated MR image to the display unit 107. After processing the audio, the control unit 111 instructs the audio output instruction unit 109 to transmit the processed audio to the audio output unit 106, causing the audio output unit 105 to output the processed audio.
[0045] If it is determined in step S203 that the user's expression is to be processed, in step S204 the control unit 111 transmits the user's expression, which was determined to be to be processed in step S203, to the image generation unit 112 along with an image generation instruction. Upon receiving the image generation instruction and the user's expression, the image generation unit 112 processes the received user's expression based on the image generation instruction and stores the difference between the user's expression before and after processing (difference in user expression). The difference between the user's expression before and after processing may be calculated, for example, as a value based on the volume of the mismatched region in 3D space, or as a value based on the difference between numerical values of emotions that can be inferred from voice, facial expressions, and poses. For example, it is conceivable to quantify emotions using a pre-trained machine learning model that quantifies emotions such that the value becomes larger the closer the emotion is to anger and smaller the closer it is to happiness. In that case, the user's expression before and after processing are each quantified, and the difference between the two values is used as the difference in user expression. The difference between the user's expression before and after processing may be stored in the temporary memory used by the image generation unit 112, or it may be stored in the storage unit 113. Regarding image processing, for example, an angry face could be identified as the target for processing and processed into a smiling face. Image generation for user expressions other than those targeted for processing is as described above. Here, facial expressions (visual information) are used as an example of user expression, but nonverbal expressions of the user can be processed in a similar manner if they can be handled by image processing. When using voice (auditory information) as the user expression, in step S204, the control unit 111 transmits the voice identified as the target for processing along with a voice generation instruction to a voice generation unit (not shown) of the information processing device 102. The voice generation unit (not shown) of the information processing device 102 receives the voice generation instruction and the user expression (auditory information) and processes the received user expression based on the voice generation instruction. For example, an angry tone of voice could be identified as the target for processing and processed into a gentle tone of voice. Naturally, it is also easy to consider combining the user's nonverbal expressions and voice as the user expression. For example, an angry face and an angry tone of voice could be processed into a smiling face and a gentle tone of voice, respectively, using the methods described above.Furthermore, other possibilities include altering the voice of an avatar, such as changing its appearance from a high-pitched voice to a low-pitched voice.
[0046] In step S205, the machining status is updated to "in progress". The machining status stored in the data storage unit 113 is updated with a value indicating "in progress". If the value indicating "in progress" is set to 1, as in the example above, then the machining status stored in the data storage unit 113 will be updated to 1. The process is then repeated from step S202.
[0047] If it is determined in step S203 that the user's expression is not the object to be processed, then in step S206, it is determined whether the processing status is in progress. In the example above, the control unit 111 retrieves the processing status stored in the data storage unit 113 and determines whether its value is 1.
[0048] If the processing status is determined to be "processing" in step S206, in step S207, the user's expression is processed so that the difference between the user's expression before and after processing is smaller than the difference between the previous version before and after processing. Here, the current real user's expression is not the target of processing, so ideally the user's expression should be left unprocessed. However, because the processing status is "processing," suddenly switching to the unprocessed version would cause the user's expression to change abruptly and unnaturally. Therefore, the expression is gradually moved towards the unprocessed version to make it feel natural. The control unit 111 transmits the user's expression detected in step S202 to the image generation unit 112 along with an image generation instruction. If the difference between the previous user's expression before and after processing is stored in the data storage unit 113, it is retrieved from the data storage unit 113 and transmitted together. When the image generation unit 112 receives the image generation instruction and the user's expression, it processes the received user's expression so that the difference in the current user's expression is smaller than the difference in the previous user's expression stored based on the image generation instruction. For example, if the difference in user expression is a value based on the volume of the discrepancy region in 3D space, the data is processed so that this value is smaller than the previous value, thereby reducing the volume of the discrepancy region in 3D space. Similarly, if the difference in user expression is a value based on the difference between numerical values representing emotions that can be inferred from voice, facial expressions, and poses, the data is processed so that this value is smaller than the previous value, thereby reducing the difference between numerical values representing inferable emotions. Examples of numerical representations of emotions will be discussed later. The processing algorithm should ideally be one that obtains the processing result by providing the desired difference in user expression or the percentage by which the difference should be reduced. Other possible algorithms include one that evaluates the difference in user expression after processing and repeats the processing if it deviates from the desired difference by a predetermined value, but considering the processing of 3D moving images, it is desirable that the processing is fast and the processing time is constant. After processing, the difference in user expression before and after processing is stored. The previously stored difference in user expression used during processing may also be updated.Furthermore, if the processing uses a history of differences in user representations from multiple instances, not just the previous one, it is acceptable to store the differences in multiple user representations in a way that indicates the storage order. In that case, to avoid storing an inexhaustible number of differences in user representations, it is conceivable to automatically delete user representations after a predetermined time has elapsed since they were stored, or to overwrite older user representation differences after a predetermined number of them have been stored, similar to a ring buffer.
[0049] In step S208, it is determined whether the user difference is greater than a predetermined value, and if so, the process is repeated from step S202.
[0050] If the user's difference is determined to be less than or equal to a predetermined value in step S208, the processing status is updated to "no processing" in step S209. In the example above, the control unit 111 updates the processing status stored in the data storage unit 113 to 0. Also, since it is considered that the series of processing operations has been completed, the user's expressed difference is deleted to prepare for subsequent processing. If processing is performed while referring to the difference history, the difference history is also deleted. Naturally, if the difference history up to the current time is referred to during subsequent processing, the difference history will be retained.
[0051] Figure 4 shows a user expression using facial expressions as an example, with the horizontal axis representing time. Below each expression, a numerical value representing emotion (emotion value) is listed. Figure 4A shows the actual user expression, and Figure 4B shows the processed user expression. For example, the first and second expressions from the left in Figure 4A have emotion values of -1 and -0.5, respectively. In step S203, these are identified as targets for processing, and in step S204, they are processed to have an emotion value of 1, as seen in the first and second expressions from the left in Figure 4B. In this case, the difference in user expressions is 2 and 1.5, respectively. As time progresses, the expression in the middle of Figure 4A has an emotion value of 0. In step S203, this is not identified as a target for processing, and the process proceeds to step S206. Since the processing status is "processing in progress," the process proceeds to step S207. Ideally, the expression in the middle of Figure 4A does not need to be processed anymore, but if it is left unprocessed, the difference in user expressions suddenly becomes 0, which could result in an unnatural expression. Therefore, the user expression difference in this step is processed so that it is slightly smaller than the difference in the user expression from the previous step, which was the second expression from the left, i.e., 1.5. In this case, the emotion value of the middle expression in Figure 4B, 0.5, is the processed result, and the difference in user expression is 0.5. If, for example, the threshold was set to 0.5 in step S208, the process proceeds to step S209, and the processing status is updated to unprocessed. The second and first expressions from the right in Figure 4A are not identified as targets for processing in step S203, so the process proceeds to step S206, and since the processing status is unprocessed, the process proceeds to step S210. In step S210, no processing is performed, so the control unit 111 performs normal control.
[0052] This approach makes it easier to achieve natural user representation. Alternatively, to gradually approach the unprocessed state, the difference in the previous user representation may be calculated by multiplying it by a predetermined coefficient as the number of processing iterations and the passage of time. Or, a pre-trained machine learning model may be used to make the transition of the user representation differences appear natural, and the target difference in user representation may be determined by inference from the model. Another approach is to use a pre-trained machine learning model to gradually approach the unprocessed state, and perform the processing of the user representation by inference from the model.
[0053] It should be noted that users may not want the process to gradually return to an unprocessed state at the end of processing. To address this, a selection screen can be presented to the user, as shown in Figure 3B. The selected value can be stored in the data storage unit 113, and for example, if OFF is selected for suppression at the end of processing, the system can be controlled to always proceed to step S210 after step S206, thereby achieving an immediate unprocessed state. Even if the selection screen shown in Figure 3B is not presented, similar settings can be pre-stored in the data storage unit 113.
[0054] (Second embodiment) In the first embodiment, the main objective was to gradually approach the unprocessed state at the end of the processing, but it is also acceptable to gradually process from the start of the processing. Since the current real-world user expression is the target of processing, ideally we would want to process the user's expression all at once to the desired user expression. However, if we process it drastically all at once, the user's expression will change abruptly and become unnatural. Therefore, we process it gradually to make it feel natural.
[0055] In that case, when the user's expression is determined to be the target of processing in step S203, the image generation unit 112 first determines the desired user expression from the current real user expression. The difference between the desired user expression and the processed user expression is processed so that the difference is smaller than the previous time, as in step S207, and after processing, the difference between the user expression before and after processing is stored. Also, if the desired user expression is used as the processing result during the first processing, the difference from the user expression after the previous processing (=output result) will increase sharply, which may result in an unnatural expression. Therefore, during the first processing, processing is performed so that the difference from the user expression after the previous processing is less than or equal to a predetermined value. To determine whether it is the first processing, for example, the processing status can be referred to as in step S206. In this case, if the processing status is "no processing", it can be determined to be the first processing. Alternatively, a flag indicating whether it is the first processing or not can be prepared separately, initialized with a value indicating that it is the first processing in step S201, updated with a value indicating that it is not the first processing after the first processing, and updated with a value indicating that it is the first processing in step S209. Furthermore, after processing, the differences in the previously stored user representation used during processing may be updated. If the history of differences in multiple user representations is used during processing, not just the previous one, multiple user representation differences may be stored in a way that indicates the storage order. In this case, to avoid endlessly storing user representation differences, it is conceivable to automatically delete user representations after a predetermined time has elapsed since they were stored, or to overwrite older user representation differences after a predetermined number of them have been stored, similar to a ring buffer.
[0056] Subsequently, in a process similar to step S208, it is determined whether the difference between the expected user representation and the processed user representation is greater than a predetermined value. If it is greater, step S205 is executed, and the process is repeated from step S202. If it is less than or equal to the predetermined value, the expected user representation is used as the processed user representation, and the process is repeated from step S202. Step S205 is optional at this time, but if the processing status may change to a value other than "processing in progress," step S205 is executed to update the processing status to "processing in progress," and the process is repeated from step S202.
[0057] Figure 5 shows a user expression using facial expressions as an example, with the horizontal axis representing time, and the emotion value is written below each expression. Figure 5A shows the actual user expression, and Figure 5B shows the processed user expression. For example, the first to third facial expressions from the left in Figure 5A have emotion values of 1, 0.5, and 0, respectively. They are not identified as targets for processing in step S203, and the process proceeds to step S206. Since the processing status is "unprocessed," the process proceeds to step S210. In step S210, no processing is performed, so the control unit 111 performs normal control. As time progresses, when the second facial expression from the right in Figure 5A is reached, the emotion value becomes -0.5. In step S203, this is identified as a target for processing, and for example, the desired user expression is changed to an expression with an emotion value of 1, as shown in the leftmost facial expression in Figure 5B. Ideally, we would like to process it to match the desired user expression, but doing so would result in a large difference from the previously processed user expression, potentially leading to an unnatural expression. Therefore, processing is performed in a way that prevents a sudden increase in the difference from the previously processed user expression. For example, if we want the difference between the user's expression after processing and the previous processing result to be 0.5 or less, we set the emotion value of the second expression from the right in Figure 5B to 0.5 as the result of the current processing, so the difference from the user's expression after processing is 0.5. Also, the difference from the ideal user expression is 0.5. We compare the difference from the ideal user expression with a predetermined value, for example, 0.3, and since it is larger, we update the processing status to "processing" in step S205, and then repeat from step S202. As time progresses, when we reach the expression on the far right of Figure 5A, the emotion value becomes -1, and in step S203 we identify this as a target for processing, and the ideal user expression is set to the same emotion value of 1 as before. We process so that the difference between the ideal user expression and the processed user expression is smaller than the previous difference of 0.5. As a result, if the difference between the ideal user expression and the processed user expression is, for example, 0.3, the difference is less than or equal to the predetermined value of 0.3, so the ideal user expression is set to the processed user expression, and we repeat from step S202.
[0058] It should be noted that users may not want the process to gradually approach the no-process state at the start of processing. To address this, a selection screen can be presented to the user, as shown in Figure 3A. The selected value can be stored in the data storage unit 113, and for example, if OFF is selected for suppression at the start of processing, the system can be controlled to always proceed to step S210 after step S206, thereby enabling a direct transition to the no-process state. Even if the selection screen shown in Figure 3B is not presented, similar settings can be pre-stored in the data storage unit 113. (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.
[0059] While preferred embodiments of the present disclosure have been described above, the present disclosure is not limited to these embodiments, and various modifications and changes are possible within the scope of its essence.
[0060] This embodiment includes the following configurations, methods, and programs. (Composition 1) An information processing device for processing user representations in at least one of CG and live-action images, A determination means for determining whether the user's expression is subject to processing, The system includes a processing means for processing the user's expression based on the determination result of the determination means, The processing means processes the user's expression based on the user's past expression and the user's target expression. An information processing device characterized by the following: (Configuration 2) The aforementioned target user expression is an expression based on the current user expression. The information processing device according to configuration 1, characterized by the above. (Composition 3) The user's expression includes at least one of visual information and / or auditory information. An information processing apparatus according to configuration 1 or 2, characterized by the above. (Composition 4) The system has output means for outputting the user's expression processed by the processing means. An information processing device according to any one of configurations 1 to 3, characterized by the above. (Composition 5) The output means is a display means that displays the user's expression. The information processing apparatus according to configuration 4, characterized by the features described above. (Composition 6) The output means is a recording means for recording the user's expression. The information processing apparatus according to configuration 4, characterized by the features described above. (Composition 7) The output means is a communication means that transmits the user's expression via communication. The information processing apparatus according to configuration 4, characterized by the features described above. (Method 1) An information processing method for processing user representations in at least one of CG and live-action images, A determination step to determine whether the user's expression is subject to processing, The process includes a processing step for processing the user's expression based on the determination result of the determination step, The processing step involves processing the user's expression based on the user's past expressions and the target user's expression. An information processing method characterized by the following: (program) A program that causes a computer to perform each step of the information processing method described in Method 1. [Explanation of Symbols]
[0061] 101 Display device 102 Information Processing Device 103 Operating device 107 Display section 110 Image Synthesis Unit 111 Control Unit 112 Image generation unit 113 Data Storage Unit 114 Processing target discrimination unit
Claims
1. An information processing device for processing user representations in at least one of CG and live-action images, A determination means for determining whether the user's expression is subject to processing, The system includes a processing means for processing the user's expression based on the determination result of the determination means, The processing means processes the user's expression based on the user's past expression and the user's desired expression. An information processing device characterized by the following:
2. The aforementioned target user expression is an expression based on the current user expression. The information processing apparatus according to feature 1.
3. The user's expression includes at least one of visual information and / or auditory information. The information processing apparatus according to feature 1.
4. The system has output means for outputting the user's expression processed by the processing means. The information processing apparatus according to feature 1.
5. The output means is a display means that displays the user's expression. The information processing apparatus according to feature 4.
6. The output means is a recording means for recording the user's expression. The information processing apparatus according to feature 4.
7. The output means is a communication means that transmits the user's expression via communication. The information processing apparatus according to feature 4.
8. An information processing method for processing user representations in at least one of CG and live-action images, A determination step to determine whether the user's expression is subject to processing, The process includes a processing step for processing the user's expression based on the determination result of the determination step, The processing step involves processing the user's expression based on the user's past expressions and the target user's expression. An information processing method characterized by the following:
9. A program for causing a computer to perform each step of the information processing method described in claim 8.
Citation Information
Patent Citations
System and method for controlling the same
JP2024077887A