Generation system and generation method

The system generates facial images with changing expressions based on user input, addressing misinterpretation issues in online communication by dynamically representing emotions, thus improving interaction clarity.

JP7807785B2Active Publication Date: 2026-01-28THE RITSUMEIKAN TRUST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2021196757
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2026-01-28
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

Face icons that do not change with emotions can lead to misinterpretation in online communication, hindering smooth interaction.

Method used

A system and method that generates facial images whose expressions change over time based on user input, using an emotion model to represent temporal changes in emotions, allowing for dynamic expression of emotions through facial expressions.

Benefits of technology

Facial images that dynamically express emotions facilitate clearer communication by accurately representing emotional changes, enhancing online interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007807785000001
    Figure 0007807785000001
  • Figure 0007807785000002
    Figure 0007807785000002
  • Figure 0007807785000003
    Figure 0007807785000003
Patent Text Reader

Abstract

To provide a generation system for generating a facial image which shows a change to express of the emotion of a user, by a change of the expression.SOLUTION: A generation system 100 includes: an input unit for receiving a user operation; and processors 11 and 51 for performing content generation processing on the basis of the user operation input by the input unit. The processor receives an input of a first operation signal showing a user operation of inputting the change with time of a value showing the emotion of an emotion model in the content generation processing, and generates a facial image in which the expression changes with time according to the change with time of the value of the emotion shown by the first operation signal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a production system and a production method. [Background technology]

[0002] For example, Japanese Patent No. 6664757 (Patent Document 1) discloses an example in which face icons are used as a method for expressing human emotions on a display. A face icon refers to a simplified representation of a face image. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 6664757 Summary of the Invention

[0004] However, face icons, which are unchanging images as shown in Patent Document 1, can sometimes lead to misinterpretation of emotions. Such misunderstandings can be a factor that hinders smooth online communication. For smooth communication, it is preferable that face icons change according to the emotions that a user wants to express. Therefore, it is desirable to be able to express changes in the emotions that a user wants to express as changes in the facial expression of a face image.

[0005] According to one embodiment, the generation system includes an input unit that accepts user operations, and a processor that executes content generation processing based on the user operations accepted by the input unit, and the processor is configured to accept, in the content generation processing, input from the input unit a first operation signal that indicates a user operation that inputs a temporal change in a value indicating an emotion in an emotion model, and to generate a facial image whose expression changes over time in accordance with the temporal change in the value indicating the emotion indicated by the first operation signal.

[0006] According to one embodiment, a generation method is a method for generating a facial image video, comprising: receiving an input of a temporal change in a value indicating an emotion in an emotion model; and generating a facial image whose facial expression changes over time in accordance with the temporal change in the value indicating the emotion. By this method, a facial image whose facial expression changes over time is generated.

[0007] Further details will be described in the following embodiments. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a schematic diagram showing an overview of a generation system according to an embodiment and a specific example of the configuration of each device included in the generation system. [Figure 2] FIG. 2 is a diagram showing an example of the flow of a content generation method in the generation system. [Figure 3] FIG. 3 is a diagram showing an outline of the input screen. [Figure 4] FIG. 4 is a diagram for explaining the image generation process and the video generation process. [Figure 5] FIG. 5 is a diagram for explaining another example of the image generation process and the moving image generation process. [Figure 6] FIG. 6 is a diagram for explaining the storage process. [Figure 7] FIG. 7 is a diagram for explaining the replacement process. [Figure 8] FIG. 8 is a diagram illustrating an example of a method for recommending a candidate image using recommendation data. [Figure 9] FIG. 9 is a diagram illustrating another example of the replacement process. [Figure 10] FIG. 10 is a diagram showing an example of the flow of a face image video search method in the generation system. [Figure 11] FIG. 11 is a diagram showing an overview of the search screen. [Figure 12] FIG. 12 is a diagram showing an outline of the setting screen. [Figure 13]FIG. 13 is a diagram for explaining another example of a method for recommending a candidate image using recommendation data. [Figure 14] FIG. 14 is a diagram showing an example of a method for storing candidate images. DETAILED DESCRIPTION OF THE INVENTION

[0009] <1. Overview of the generation system and generation method>

[0010] (1) A generation system according to an embodiment includes an input unit that accepts user operations and a processor that executes content generation processing based on the user operations accepted by the input unit. The processor is configured to accept, in the content generation processing, input from the input unit a first operation signal that indicates a user operation that inputs a temporal change in a value indicating an emotion in an emotion model, and to generate a facial image whose expression changes over time in accordance with the temporal change in the value indicating the emotion indicated by the first operation signal.

[0011] A facial image can be any image of a face that can express emotions through facial expressions, including human or animal faces, as well as images or photographs of faces that mimic them. By generating a facial image whose facial expression changes over time, changes in the emotions the user wants to express can be expressed as changes in the facial expression in the facial image. This makes it easier to interpret the user's emotions. As a result, this effectively supports smoother communication using facial images, such as online communication.

[0012] (2) Preferably, the emotion model is a model represented by a coordinate system. The coordinate system is, for example, a two-dimensional coordinate system. This makes it possible to easily represent emotions using coordinate values.

[0013] (3) Preferably, the emotion model is the Russell Circumplex model. By using the Russell Circumplex model, emotions can be expressed by combining emotional changes from "sad" to "happy" and changes from "arousal" to "non-arousal."

[0014] (4) Preferably, inputting the change over time in the value indicating the emotion includes inputting coordinate values ​​at a plurality of points in time, whereby the change in emotion is represented by the coordinate values.

[0015] (5) Preferably, the step of inputting the change in the value indicating the emotion over time includes inputting a line drawn in a coordinate system. Continuous changes in emotion can be easily input.

[0016] (6) Preferably, the facial image whose facial expression changes over time includes a facial image in which at least a part of the form of the facial image changes according to a value indicating an emotion, whereby a change in emotion is represented by a change in facial expression over time.

[0017] (7) Preferably, at least a portion of the facial image includes at least one of the mouth and eyebrows. For example, in Russell's circular model, pleasant and unpleasant feelings are expressed by changes in the mouth, and arousal levels are expressed by changes in the shape of the eyebrows. Therefore, emotions can be accurately expressed by changes in the shape of at least one of the mouth and eyebrows.

[0018] (8) Preferably, the facial image whose facial expression changes over time includes a display position of the facial image that changes according to a value indicating an emotion, whereby a change in emotion is represented by a change in facial expression over time.

[0019] (9) Preferably, the processor is configured to present candidate images for replacing at least a portion of the facial image whose expression changes over time, thereby generating a facial image that more accurately expresses the emotion the user wants to express.

[0020] (10) Preferably, presenting the candidate images includes selecting the candidate images from among stored images based on a temporal change in the value indicated by the first operation signal, thereby increasing the likelihood that a facial image that is close to the user's emotion will be selected as the candidate image.

[0021] (11) Preferably, the processor is configured to receive, from the input unit, a second operation signal indicating a user operation to designate one of the facial images whose expressions change over time, and to associate content with the facial image designated by the second operation signal. The content may be any content that can be played back together with the facial image whose expressions change over time, such as text, images, audio, or a combination thereof. This allows the user to express the emotion they wish to express through the content in addition to the facial image whose expressions change over time. This makes it possible to accurately express emotional changes, such as in a review of an object.

[0022] (12) Preferably, the processor is configured to store in the memory a temporal change in the value indicating the emotion indicated by the first operation signal and the generated facial image whose expression changes over time in association with each other, which can be used for, for example, searching, which will be described later.

[0023] (13) Preferably, the processor is configured to retrieve, from the memory, facial images whose expressions change over time based on the change over time of the input emotion value, thereby enabling the search for facial images whose expressions change over time using the tendency of emotion change.

[0024] (14) A generation method according to an embodiment is a method for generating a facial image video, which includes receiving input of a temporal change in a value indicating an emotion in an emotion model, and generating a facial image whose facial expression changes over time in accordance with the temporal change in the value indicating the emotion. This method generates a facial image whose facial expression changes over time. By using such a facial image, changes in the emotion that a user wants to express are expressed as changes in the facial expression in the facial image. This makes it easier to interpret the user's emotions. As a result, this effectively supports smoother communication using facial images, such as online communication.

[0025] <2. Examples of generation systems and generation methods>

[0026] FIG. 1 is a schematic diagram showing an overview of a generation system 100 according to the present embodiment and a specific example of the configuration of each device included in the generation system 100. The generation system 100 generates a video of facial images whose expressions change over time. The facial images may be any facial images that can express emotions through facial expressions, such as human or animal faces, or images or photographs of faces that resemble them. Referring to FIG. 1, the generation system 100 includes, as an example, a server 1 that functions as a backend and a terminal device 5 that functions as a frontend.

[0027] The terminal device 5 is, for example, a tablet or a personal computer, and is used as a user interface. Referring to Fig. 1, the terminal device 5 is configured by a general computer having a processor 51 and a memory 52. ​​The processor 51 is, for example, a CPU (Central Processing Unit).

[0028] The memory 52 may be a primary storage device or a secondary storage device. The memory 52 stores a computer program 521 executed by the processor 51. The processor 51 executes the computer program 521 stored in the memory 52 to perform a content generation process 510. The content generation process 510 executed by the processor 51 includes at least a part of a process in the generation system 100 to generate a video of facial images whose expressions change over time.

[0029] The terminal device 5 includes a communication device 53. The communication device 53 is, for example, a communication module. The communication device 53 communicates with the server 1 via a communication network 3 such as the Internet. The communication device 53 transmits instructed information to the server 1 in accordance with a control signal from the processor 51. The communication device 53 also inputs an electrical signal received from the server 1 to the processor 51.

[0030] The terminal device 5 has a display unit that displays images to the user, and a touch panel 54 that is an example of an input unit that accepts user operations. The touch panel 54 displays information based on the specified information in accordance with a control signal from the processor 51. The touch panel 54 also accepts user operations and inputs operation signals indicating the user operations to the processor 51.

[0031] The content generation process 510 executed by the processor 51 includes an image generation process 511. The image generation process 511 includes generating a facial image using an operation signal input from the touch panel 54. The processor 51 stores an emotion model 512, and uses the emotion model 512 in the image generation process 511. The emotion model is a model represented by a coordinate system. The coordinate system is, for example, a two-dimensional coordinate system. This makes it possible to easily represent emotions using coordinate values.

[0032] The processor 51 passes the generated face image to the communication device 53, causing it to be transmitted to the server 1. This allows the server 1 to perform content generation processing using the face image.

[0033] The processor 51 executes the computer program 521 to perform display processing 513. The display processing 513 includes accessing another device in accordance with access information received from the server 1 and causing the touch panel 54 to perform display based on the acquired data. This allows a video generated by the server 1, which will be described later, to be displayed on the touch panel 54.

[0034] 1, the server 1 is configured as a computer having a processor 11 and a memory 12. The processor 11 is, for example, a CPU. The memory 12 includes a flash memory, an EEPROM, a ROM, a RAM, etc. Alternatively, the memory 12 may be a primary storage device or a secondary storage device.

[0035] The memory 12 stores a computer program 121 that is executed by the processor 11. The processor 11 executes the computer program 121 to perform a content generation process 110. The content generation process 110 executed by the processor 11 includes at least a part of a process in the generation system 100 to generate a video of facial images whose expressions change over time.

[0036] The memory 12 has a moving image storage unit 122, which is an area for storing moving images generated by the content generation process 110. The memory 12 also has a recommendation data storage unit 123, which is a storage area for storing recommendation data. The recommendation data will be described later.

[0037] The server 1 includes a communication device 13. The communication device 13 is, for example, a communication module. The communication device 13 communicates with the terminal device 5 via the communication network 3. The communication device 13 transmits instructed information to the terminal device 5 in accordance with a control signal from the processor 11. The server 1 also inputs an electrical signal received from the communication device 13 to the processor 11.

[0038] The content generation process 110 executed by the processor 11 includes a moving image generation process 111. The moving image generation process 111 includes generating a moving image of facial images whose expressions change over time, using a plurality of facial images transmitted from the terminal device 5.

[0039] Preferably, the moving image generation process 111 includes a replacement process 112. The replacement process 112 includes replacing at least some of the facial images in a moving image generated based on a plurality of facial images transmitted from the terminal device 5 with different images.

[0040] The processor 11 executes the computer program 121 to perform the storage process 113. The storage process 113 includes storing the video generated by the video generation process 111 in the memory 12. The storage process 113 also includes causing the communication device 13 to transmit, to the terminal device 5, access information for accessing the stored memory 12.

[0041] Preferably, the content generation process 110 executed by the processor 11 includes an addition process 114. The addition process 114 includes adding other content to at least a part of the video generated by the video generation process 111. The other content is, for example, text. The video to which the text has been added may be stored in the memory 12.

[0042] Preferably, the processor 11 executes the computer program 121 to perform the search process 115. The search process 115 includes searching the memory 12 for a corresponding video according to a signal from the terminal device 5.

[0043] At least a part of the content generation process 110 executed by the processor 11 of the server 1 may be performed by the processor 51 of the terminal device 5. That is, at least one of the video generation process 111, the replacement process 112 therein, and the addition process 114 may be performed by the processor 51 of the terminal device 5. In this case, the generated content may be passed from the terminal device 5 to the server 1, and the server 1 may perform the storage process 113 and store the content in the memory 12. The content generation process 110 and the content generation process 510 are merely examples of a method for generating content in the generation system 100.

[0044] Fig. 2 is a diagram showing an example of the flow of a content generation method in the generation system 100. Referring to Fig. 2, the generation system 100 first receives a user operation indicating a change in emotion via the terminal device 5 (step S1). The change in emotion is represented by a temporal change in a value indicating the emotion in an emotion model. One example of an emotion model is the Russell Circumplex model.

[0045] FIG. 3 is a diagram illustrating an example of an input screen 41 displayed on the terminal device 5 to receive a user operation indicating a change in emotion in step S1. Referring to FIG. 3, the input screen 41 includes a coordinate system 60. The coordinate system 60 is set in accordance with Russell's circumplex model. Specifically, the coordinate system 60 is defined by a first axis 61 representing a change from "sad" to "happy" and a second axis 62 representing a change from "arousal" to "relaxed." As an example, the first axis 61 is the horizontal axis, and the second axis 62 is the vertical axis. Using Russell's circumplex model, it is possible to express emotions by combining the change from "sad" to "happy" with the change from "arousal" to "relaxed."

[0046] Each position on the coordinate system 60 is expressed by a coordinate value defined by a first axis 61 and a second axis 62. By specifying a position on the coordinate system 60 corresponding to an emotion, the emotion can be expressed by a coordinate value. That is, the value representing the emotion is, for example, a coordinate value on the coordinate system 60.

[0047] The input screen 41 receives input of values ​​indicating emotions by receiving a user's touch operation on the coordinate system 60. Inputting temporal changes in emotions corresponds to inputting coordinate values ​​at multiple points in time. Inputting coordinate values ​​at multiple points in time may also be inputting coordinate values ​​at discrete points in time. In this way, changes in emotions are represented by coordinate values.

[0048] Inputting coordinate values ​​at multiple points in time may be, for example, inputting an emotion curve 66, which is a line drawn in a coordinate system. This makes it easy to input continuous changes in emotion.

[0049] The emotion curve 66 is input by the movement trajectory of the touch position 65 relative to the coordinate system 60 while maintaining the touch state. The example of Figure 3 shows that the touch position 65 starts from point P1, moves through points P2, P3, P4... to point PN while maintaining the touch state, and the touch state is released at point PN.

[0050] Processor 51 receives input of an operation signal (first operation signal) representing emotion curve 66 from touch panel 54. The first operation signal is continuously input to processor 51 as touch position 65 moves while maintaining the touch state until the touch state is released. When the first operation signal is input, processor 51 executes image generation processing 511 (step S3). Processor 51 repeatedly executes image generation processing 511 until the continuous input of the first operation signal is terminated.

[0051] FIG. 4 is a diagram for explaining the image generation process 511 executed by the processor 51 of the terminal device 5 and the video generation process 111 executed by the processor 11 of the server 1. As shown in FIG.

[0052] In image generation processing 511, processor 51 generates a facial image according to the coordinate values ​​indicated by the first operation signal. That is, each time a first operation signal indicating each of multiple points P1, P2, P3, P4...PN included in emotion curve 66 is input, multiple corresponding facial images 71, 72, 73 are generated. In this way, a group of facial images 70 is generated in real time according to the movement trajectory of touch position 65.

[0053] The generated facial images 71, 72, and 73 have at least a portion of their morphology corresponding to the value indicating the emotion. As a result, the morphology of at least a portion of the facial image changes with a change in the value indicating the emotion. The morphological change includes a change in at least one of shape, pattern, color, size, and a combination thereof, or a morphological change associated with the direction of the face. As a result, changes in emotion are accurately represented by temporal changes in facial expression.

[0054] At least one of the mouth and eyebrows is included in the at least part of the facial expression, for example, both the mouth and the eyebrows. For example, in Russell's circular model, pleasant and unpleasant feelings are expressed by changes in the mouth, and arousal is expressed by changes in the shape of the eyebrows. Therefore, emotions can be accurately expressed by changes in the shape of at least one of the mouth and the eyebrows.

[0055] As one example, the shape of the mouth and the shape of the eyebrows are set in advance according to a value (coordinate value) indicating an emotion. In this case, in image generation processing 511, processor 51 generates, for each input value, a facial image in which the shape of the mouth and the shape of the eyebrows correspond to the input value. As another example, processor 51 may store in advance a function that calculates the shape of the mouth and the shape of the eyebrows using the value (coordinate value) indicating the emotion as a variable, and generate a facial image by calculating the shape of the mouth and the shape of the eyebrows for each input value by substituting the input value into the function.

[0056] 2, when a face image is generated in terminal device 5, the generated face image together with the coordinate values ​​indicated in the first operation signal is passed from terminal device 5 to server 1 (step S5). Transmission from terminal device 5 to server 1 may be performed every time a face image is generated, at short time intervals, every time a predetermined number of face images are generated, or at the timing when the touch state at touch position 65 is released.

[0057] The server 1 executes a moving image generation process 111 (step S7). The moving image generation process 111 may be executed every time a face image is transmitted from the terminal device 5, or may be executed after the touch state is released.

[0058] 4, in the moving image generation process 111, as an example, an image is generated based on a plurality of facial images 71, 72, and 73 generated in the terminal device 5. The method of generating a moving image in the moving image generation process 111 is not limited to a specific method, and may be a general method of generating a moving image using a plurality of images. For example, in the specific moving image generation process 111, the plurality of facial images 71, 72, and 73 may or may not be interpolated. As a result, a facial image moving image 80 is generated, which is a facial image whose facial expression changes over time according to a value indicating an emotion along time t.

[0059] By generating the facial image video 80, changes in the emotions the user wants to express are expressed as changes in facial expression in the facial image. This makes it easier to interpret the user's emotions. As a result, this effectively supports smoother communication using facial images, such as online communication.

[0060] In addition, in a facial image whose expression changes over time, the display position of the facial image may change according to the value indicated by the emotion. The change in the display position of the facial image according to the value indicated by the emotion may be such that the speed of the change changes according to the value indicated by the emotion. Furthermore, in a facial image whose expression changes over time, both the shape and display position of at least a part of the facial image may change according to the value indicated by the emotion. In this way, changes in emotion are represented by changes in the facial expression over time in the facial image.

[0061] FIG. 5 is a diagram illustrating another example of the image generation process 511 executed by the processor 51 of the terminal device 5 and the video generation process 111 executed by the processor 11 of the server 1. Referring to FIG. 5, in the image generation process 511 as another example, the processor 51 generates a facial image group 75 including a plurality of facial images 76, 77, and 78 corresponding to a plurality of points P1, P2, P3, P4, ...PN included in the emotion curve 66. The generated facial images 76, 77, and 78 have display positions corresponding to values ​​indicating emotions. As an example, the display positions of the facial images are set in advance according to values ​​(coordinate values) indicating emotions. In this case, in the image generation process 511, the processor 51 generates, for each input value, a facial image whose display position corresponds to the input value.

[0062] In the video generation process 111, the processor 11 of the server 1 may, for example, interpolate between the plurality of facial images 76, 77, and 78 included in the facial image group 75. In this case, the processor 11 interpolates a facial image that is to be displayed at a position between the plurality of facial images 76, 77, and 78. This generates a facial image video 87 in which the display position of the facial image changes over time according to the value of the facial expression indicating the emotion along time t.

[0063] 2, in generation system 100, when a facial image moving image is generated in server 1, the generated facial image moving image is stored in moving image storage unit 122 of memory 12 (step S9), and access information to memory 12 is passed from server 1 to terminal device 5 (step S10). Terminal device 5 accesses memory 12 in accordance with the access information to acquire moving image data, thereby displaying the moving image on touch panel 54 (step S13).

[0064] In step S9, the processor 11 of the server 1 associates the temporal change in the value indicating the emotion indicated by the first operation signal with the generated facial image video in which the facial expression changes over time, and stores the associated video in the video storage unit 122. FIG. 6 is a diagram for explaining the storage process 113 in the server 1. In the example of FIG. 6, each of the generated facial image videos 32A-32D is stored in the video storage unit 122 in association with the change in coordinate values ​​indicated in the first operation signal transmitted from the terminal device 5, that is, the emotion curves 31A-31D used to generate the respective videos. Storing the facial image videos in memory 12 in this manner in association with the emotion curves allows them to be used, for example, in the search process 115, which will be described later.

[0065] Preferably, the process of generating a facial image video in which facial expressions change over time includes replacing at least a portion of the facial images. That is, the processor 11 of the server 1 executes a replacement process 112 for at least a portion of the facial image video in which facial expressions change over time. At least a portion of the facial images are, for example, facial images in accordance with a user operation accepted by the terminal device 5.

[0066] 7 is a diagram illustrating the replacement process 112. As an example, after the facial image video 80 is generated, the processor 51 of the terminal device 5 accepts a user operation to specify point Q on the emotional curve 66 as an operation to specify a facial image to be replaced. In response to this user operation, the terminal device 5 passes a signal to the server 1 requesting the replacement process 112. As a result, the processor 11 executes the replacement process 112 and replaces the facial image corresponding to point Q in the facial image video 80 with another facial image.

[0067] As another example, the point Q may be a predefined point. The predefined point may be, for example, the end point PN of the emotion curve, the start point P1, or an intermediate point.

[0068] Preferably, in the replacement process 112, the processor 11 passes candidate images 82, 83, ... for replacing the facial image included in the facial image video 80 to the terminal device 5 to have it presented. This allows the user to select another facial image from the presented candidate images to replace the facial image at the specified point Q, making the operation easier. By replacing the facial image corresponding to point Q in the facial image video 80 with the selected facial image, the processor 11 can generate a facial image video that more accurately expresses the emotion the user wants to express.

[0069] The replacement process 112 includes determining candidate images 82, 83, ... by referring to the recommendation data storage unit 123. Fig. 8 is a diagram for explaining an example of a method for recommending candidate images using recommendation data stored in the recommendation data storage unit 123.

[0070] The recommendation data is data showing one or more facial images corresponding to each emotional curve. As an example, the recommendation data is data in which the replaced facial images are associated with the emotional curves by the replacement process 112 performed earlier.

[0071] 8, for example, recommendation data 30A is data representing the correspondence between emotional curve 31A and facial image 31A1 replaced by previous replacement process 112. Recommendation data 30B, 30C, ... 30N is data representing the correspondence between emotional curves 31B, 31C, ... 31N and facial images 31B1, 31C1, ... 31N1 replaced by previous replacement process 112, respectively.

[0072] In replacement processing 112, processor 11 selects a candidate image to recommend from among the stored facial images based on temporal changes in values ​​representing emotions expressed by user operations on terminal device 5 that instruct the generation of facial image video 80. Specifically, processor 11 calculates the similarity between each of emotional curves 31A, 31B, 31C, ... 31N included in the recommendation data and emotional curve 66, and selects a facial image associated with one of emotional curves 31A, 31B, 31C, ... 31N obtained by calculation based on the similarity as the candidate image to recommend. As an example, among emotional curves 31A, 31B, 31C, ... 31N whose similarity to emotional curve 66 is higher than a certain level, those with the most associated facial images may be selected as candidate images using recommendation data.

[0073] In the example of Fig. 8, the emotional curves 31A, 31B, 31C, ... 31N have the highest similarity to the emotional curve 66, in that order, meaning that the emotional curve 31A is most similar in shape to the emotional curve 66. In the example of Fig. 8, one facial image is associated with each of the emotional curves 31A, 31B, 31C, ... 31N. Therefore, as an example, the processor 11 determines that the facial image 31A1 associated with the emotional curve 31A in the recommendation data 30A is first in the recommendation order.

[0074] As another example, candidate images may be selected using a partial curve 66A from the starting point P1 to point Q of the emotion curve 66 shown in Fig. 7. As shown in Fig. 13, the recommendation data may further include a facial image associated with a point used for replacement on the emotion curve. In this case, the processor 11 may determine the recommendation order based on the similarity between the partial curve of the emotion curve 66 and the partial curve up to the point used for replacement on one of the emotion curves 31A, 31B, 31C, ... 31N.

[0075] As an example, the processor 11 reads out a plurality of pieces of recommendation data 30A, 30B, 30C, ... 30N from the recommendation data storage unit 123, and causes the terminal device 5 to present, as candidate images, facial images 31A1, 31B1, 31C1, ... 31N1 associated with emotional curves in each piece of recommendation data in the determined recommendation order. As a result, the facial images used in the replacement process 112 with facial image video that shows emotional changes close to the emotional curve 66 are selected as candidate images in order. This increases the likelihood that a facial image that is close to the user's emotion will be selected as a candidate image.

[0076] As another example of a method for selecting candidate images from stored face images based on temporal changes in values ​​representing emotions, candidate images may be stored for each coordinate value in the recommendation data storage unit 123, as shown in Fig. 14. Fig. 14 is a diagram showing an example of a method for storing candidate images in the recommendation data storage unit 123. In the example of Fig. 14, candidate recommendation face images are stored on a coordinate system 60.

[0077] In this case, in the replacement process 112, the processor 11 may select a candidate image to recommend based on the distance between the associated coordinate value and the point Q. As an example, the processor recommends candidate images in descending order of the distance between the associated coordinate value and the point Q.

[0078] For example, the candidate images for each coordinate value may be a plurality of face images prepared in advance, which may be obtained using face images to which a plurality of users have subjectively assigned coordinate values ​​in the coordinate system 60. In this case, for example, a face image having coordinate values ​​closest to the coordinate values ​​of point Q or a face image within a predetermined range may be determined as the candidate image.

[0079] As another example, the replacement process 112 may be to deform a face image corresponding to point Q in the face image video 80 in accordance with a user operation. Fig. 9 is a diagram for explaining another example of the replacement process 112. With reference to Fig. 9, the processor 11 receives a user operation from the terminal device 5 instructing deformation of face image 81 corresponding to point Q in the face image video 80, and deforms the face image 81 in accordance with the user operation.

[0080] 9, the deformation may be, for example, a change in the size of face image 81. In the example of Fig. 9, processor 11 enlarges face image 81 to generate face image 84, which replaces face image 81 in face image video 80.

[0081] As another example of the transformation, a mask image may be added to the facial image 81. The mask image may be stored in advance, or the processor 11 may obtain the mask image by accessing another device. The example in Fig. 9 shows an example in which the processor 11 generates a facial image 86 by adding a mask image 85 specified by a user operation to the facial image 81, and replaces the facial image 81 in the facial image video 80.

[0082] In the generation system 100, the generated facial image moving image is associated with the emotional curve and stored in the memory 12. Therefore, the generation system 100 preferably performs processing that utilizes the stored facial image moving image.

[0083] The process of utilizing facial image videos includes, for example, accepting input of an emotional curve from a user and retrieving facial image videos from memory 12 based on the input emotional curve. Retrieving facial image videos from memory 12 based on the input emotional curve may be retrieving facial image videos generated by an emotional curve similar to the input emotional curve, or retrieving facial image videos generated by an emotional curve that has a specific relationship with the input emotional curve. The specific relationship may be, for example, a relationship in which the curves change inversely, or a relationship in which only a specific range is similar and other ranges are different.

[0084] FIG. 10 is a diagram illustrating an example of the flow of a method for searching for facial image videos in the generation system 100. Referring to FIG. 10, the generation system 100 first receives, via the terminal device 5, a user operation indicating a change in emotion for search (step S15). FIG. 11 is a diagram illustrating an overview of a search screen 45 as an example of a screen displayed on the terminal device 5 for receiving a user operation indicating a change in emotion for search in step S15. Referring to FIG. 11, the search screen 45 includes a coordinate system 60, and receives a user's touch operation on the coordinate system 60 to input a value indicating an emotion. The change in emotion for search is input, for example, by a search curve 67 represented by the movement trajectory of a touch position 65.

[0085] The search screen 45 includes a button 46 for instructing a search for a corresponding face image video. Referring to Fig. 10, when the button 46 on the search screen 45 is touched, the input search curve 67 is sent from the terminal device 5 to the server 1 (step S17). In the server 1, a search process 115 is executed to search for a corresponding face image video (step S19).

[0086] 6, generated facial image videos and the emotional curves used for their generation are stored in association with each other in the video storage unit 122. In the search process 115, the processor 11 searches for facial image videos stored in association with emotional curves similar to the input search curve 67, for example. This results in the search for facial image videos with similar trends in the changes in values ​​representing emotions.

[0087] When the face image video is searched for in the server 1, access information to the memory 12 where the face image video is stored is passed from the server 1 to the terminal device 5 (step S21). The terminal device 5 accesses the memory 12 in accordance with the access information to acquire the video data, thereby displaying the video on the touch panel 54 (step S23).

[0088] By performing such a search, the user can obtain a corresponding facial image video by inputting the tendency of emotional changes as a search curve into the terminal device 5. Therefore, it is possible to easily search for a facial image video.

[0089] As another example, the process of utilizing the facial image video may include adding content to at least some of the facial images in the generated facial image video. The content may be, for example, text, images, audio, or a combination thereof, as long as it can be played back together with the facial image video. As an example, it is assumed that text input to the terminal device 5 is added.

[0090] 12 is a diagram showing an overview of a setting screen 21 as an example of a screen displayed on the terminal device 5 for setting content to be added to the facial image video 80. As an example, the setting screen 21 accepts designation of facial images 91, 92, and 93 to which text is to be added from the target facial image video 80. The designation of the facial images 91, 92, and 93 may be performed by, for example, a touch operation on the facial image video 80. Note that the facial images 91, 92, and 93 may be automatically set by the processor 11, or may be configured to accept changes to the automatically set images.

[0091] The setting screen 21 accepts input of text for each of the facial images 91, 92, and 93. Referring to Fig. 12, the setting screen 21 includes input fields 22, 23, and 24 for text to be added to the facial images 91, 92, and 93, respectively.

[0092] Adding text to the facial image video 80 is expected, for example, when creating a review about an object. More specifically, the facial image video 80 changes its display as the emotion changes from negative to positive over time, as shown by the emotion curve 66. When creating a review about an object in which such an emotion change has occurred, it is expected that the facial image video 80 is generated using the generation system 100, and text expressing the change in emotion is added to the generated facial image video 80.

[0093] In the example of Fig. 12, facial images 91, 92, and 93 correspond to negative, neutral, and positive emotions, respectively. In this case, in the example of Fig. 12, texts expressing negative, neutral, and positive impressions about an object (e.g., a product) are entered in input fields 22, 23, and 24, respectively.

[0094] The setting screen 21 includes a button 25 for instructing the setting of the facial images 91, 92, and 93 and the addition of text. When the button 25 is touched, an operation signal (second operation signal) indicating a user operation for designating the facial images 91, 92, and 93, and the text entered in the input fields 22, 23, and 24 are transferred from the terminal device 5 to the server 1. In the server 1, an addition process 114 is executed, and the entered text is added to each of the facial images 91, 92, and 93 in the facial image video 80.

[0095] By adding content such as text to the facial image video 80, it becomes possible to express the emotions that the user wants to express through the content in addition to the facial image video, making it possible to accurately express changes in emotions such as reviews of an object.

[0096] The text added to the facial image video 80 may be stored in association with the facial image video in the video storage unit 122, as shown in Fig. 6. In the example of Fig. 6, contents 33A, 33B, and 33C, which are text added to the facial image video 32A, are stored in association with the facial image video 32A. This allows a review of the object to be generated as a facial image video and stored in the memory 12.

[0097] Facial image videos with added content such as text can also be stored in memory 12 and can be subject to search processing 115 shown in Figures 10 and 11. This allows, for example, when creating a review, to easily obtain facial image videos with added content that are generated based on similar emotional changes by inputting changes in emotions toward an object as a search curve 67.

[0098] <3. Notes> The present invention is not limited to the above-described embodiment, and various modifications are possible. [Explanation of symbols]

[0099] 1: Server 3: Communication network 5: Terminal device 11: Processor 12: Memory 13: Communication equipment 21: Settings screen 22: Input field 23: Input field 24: Input field 25: Button 30A: Recommendation data 30B: Recommendation data 30C: Recommendation data 30N: Recommendation data 31A: Emotional curve 31B: Emotional curve 31C: Emotion curve 31N: Emotion curve 31A1:Facial image 31B1:Facial image 31C1:Facial image 31N1:Facial image 32A:Facial image video 32B:Facial image video 32C:Facial image video 32D:Facial image video 33A: Content 33B: Content 33C: Content 41: Input screen 45: Search screen 46: Button 51: Processor 52: Memory 53: Communication equipment 54: Touch panel 60: Coordinate system 61: First axis 62: Second axis 65: Touch position 66: Emotion curve 66A: Partial curve 67: Search curve 70: Facial images 71:Facial image 72:Facial image 73:Facial image 75: Facial images 76:Facial image 77:Facial image 78:Facial image 80: Face image video 81:Facial image 82: Candidate image 83: Candidate image 84:Facial image 85: Mask image 86:Facial image 87: Face image video 91:Facial image 92:Facial image 93:Facial image 100: Generator System 110: Content generation processing 111: Video generation processing 112: Replacement process 113: Storage process 114: Additional processing 115: Search processing 121: Computer Programs 122: Video storage unit 123: Recommendation data storage unit 510: Content generation process 511: Image generation processing 512: Emotional Model 513: Display processing 521: Computer Programs

Claims

1. an input unit that accepts user operations; a processor that executes a content generation process based on the user operation received by the input unit, In the content generation process, the processor receiving, from the input unit, an input of a first operation signal indicating the user operation of inputting a temporal change in a value indicating an emotion in an emotion model expressed in a coordinate system; a facial image video in which facial expressions change over time in accordance with a change over time in the value indicating the emotion indicated by the first operation signal is generated and stored; Inputting the temporal change in the emotion-indicating value includes inputting coordinate values ​​at a plurality of points in time. Generation system.

2. The emotion model is Russell's circumplex model. The production system of claim 1 .

3. Inputting the temporal change in the value indicating the emotion includes inputting a line drawn in a coordinate system. A generating system according to claim 1 or 2.

4. The facial image whose expression changes over time includes a facial image in which at least a part of the form of the facial image changes according to the value indicating the emotion. A production system according to any one of claims 1 to 3.

5. At least a portion of the facial image includes at least one of a mouth and eyebrows. The production system of claim 4 .

6. The facial image whose expression changes over time includes changing the display position of the facial image according to the value indicating the emotion. A production system according to any one of claims 1 to 5.

7. The processor is configured to present candidate images for replacing at least a portion of the facial image of the facial image having a time-varying expression. A production system according to any one of claims 1 to 6.

8. The processor is configured to receive, from the input unit, an input of a second operation signal indicating the user operation to designate any one of the facial images whose expressions change over time, and to associate content with the facial image designated by the second operation signal. A production system according to any one of claims 1 to 7.

9. The processor is configured to store in a memory a temporal change in the value indicating the emotion indicated by the first operation signal and the generated facial image whose expression changes over time in association with each other. A production system according to any one of claims 1 to 8.

10. The processor is configured to search the memory for a face image whose expression changes over time based on a change over time of the input emotion-indicating value. The production system of claim 9.

11. An input unit that accepts user operations; a processor that executes a content generation process based on the user operation received by the input unit, In the content generation process, the processor receiving, from the input unit, an input of a first operation signal indicating the user operation of inputting a temporal change in a value indicating an emotion in an emotion model; a facial image whose expression changes over time in accordance with a change over time in the value indicating the emotion indicated by the first operation signal; The facial image whose expression changes over time includes changing the display position of the facial image according to the value indicating the emotion. Generation system.

12. An input unit that accepts user operations; a processor that executes a content generation process based on the user operation received by the input unit, In the content generation process, the processor receiving, from the input unit, an input of a first operation signal indicating the user operation of inputting a temporal change in a value indicating an emotion in an emotion model; a facial image whose expression changes over time in accordance with a change over time in the value indicating the emotion indicated by the first operation signal; The processor is configured to present candidate images for replacing at least a portion of the facial image of the facial image having a time-varying expression. Generation system.

13. Presenting the candidate image includes selecting the candidate image from among stored images based on a temporal change in the value indicated by the first operation signal. A production system according to claim 7 or 12.

14. A method for generating a facial image video, comprising: Accepts input of temporal changes in values ​​indicating emotions in an emotion model expressed in a coordinate system; generating and storing a facial image video in which facial expressions change over time in accordance with the change over time of the emotion-indicating value; Inputting the temporal change in the emotion-indicating value includes inputting coordinate values ​​at a plurality of points in time. Generation method.

Citation Information

Patent Citations

  • Sales support device, sales support method, and sales support program

    JP6664757B1