Information processing device, non-verbal information conversion system, information processing method and program
The system uses user-specific models to estimate and convert non-verbal information, addressing individual and cultural differences, ensuring clear intention conveyance in dialogue communication.
Patent Information
- Application Number
- JP2021044286
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-03-18
- Publication Date
- 2025-11-12
- Estimated Expiration
- 2041-03-18
AI Technical Summary
Conventional methods struggle to accurately convert non-verbal information considering individual and cultural differences in expression and recognition, leading to misinterpretation in dialogue communication.
An information processing device and system that utilizes user-specific non-verbal expression and recognition models to estimate intentions and convert non-verbal information for improved understanding between communicators.
Enables effective conversion of non-verbal information to convey intended intentions clearly, addressing individual and cultural differences, thereby enhancing dialogue communication accuracy.
Smart Images

Figure 0007767723000001 
Figure 0007767723000002 
Figure 0007767723000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, a non-verbal information conversion system, an information processing method, and a program. [Background technology]
[0002] In recent years, with the advancement of deep learning, it has become possible to accurately recognize non-verbal information such as a person's gaze and facial expressions from video footage in real time, and this has been applied to a variety of applications, such as automatic analysis of surveillance camera footage and health monitoring. In addition, non-verbal information conversion technology, which has developed in conjunction with non-verbal information recognition technology, has been attracting attention in recent years, and by using these technologies, it is possible to give the desired impression to the other party in a conversation using video calls, etc.
[0003] Here, correct expression and recognition of non-verbal information is important in dialogue-based communication. However, problems may arise in dialogue if non-verbal information is not handled correctly or if both parties misinterpret the non-verbal information. As a solution to such problems, for example, Patent Document 1 discloses a method for estimating emotions from speech and generating and displaying character images that can easily convey those emotions to the other party. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-330958 [Non-patent literature]
[0005] [Non-Patent Document 1] Tabas Baltrusaitis, et al., “OpenFace: an open source facial behavior analysis toolkit”, ICCV 2016. [Non-patent document 2] A. Pumarola,et al.,“Ganimation: Anatomically-aware facial animation from a single image”,ECCV,2018. Summary of the Invention [Problem to be solved by the invention]
[0006] However, because there are individual and cultural differences in the recognition and expression of non-verbal information, it is necessary to make estimations and expressions that are appropriate for each individual.However, conventional methods have had the problem that there is room for improvement in converting non-verbal information that takes into account the individual characteristics of how non-verbal information is expressed and recognized. [Means for solving the problem]
[0007] In order to solve the above-mentioned problem, the invention of claim 1 provides an intention estimation means for estimating the intention of a first user indicated in first non-verbal information based on first non-verbal information which is non-verbal information of a first user and a first non-verbal model indicating a relationship between the first non-verbal information and an intention; and a second non-verbal model indicating a relationship between second non-verbal information, which is non-verbal information of a second user, and the intention; and Based on The first non-verbal information, and non-verbal information conversion means for converting the non-verbal information into second non-verbal information to be output to a second user. [Effects of the Invention]
[0008] According to the present invention, it is possible to convert non-verbal information in dialogue communication to convey the intention to be conveyed to the other party in an easily understandable manner. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a diagram illustrating an example of the overall configuration of a non-linguistic information conversion system. [Figure 2]FIG. 1 illustrates an example of a hardware configuration of a computer. [Figure 3] FIG. 2 is a diagram illustrating an example of a functional configuration of an information processing apparatus. [Figure 4] 1 is a schematic diagram illustrating an example of a non-language information conversion system according to a first embodiment. [Figure 5] 10 is a flowchart illustrating an example of a non-linguistic information conversion process. [Figure 6] 10 is a flowchart illustrating an example of a process for extracting features of non-language information. [Figure 7] 10 is a flowchart illustrating an example of a process for estimating a sender's intention. [Figure 8] FIG. 10 is a conceptual diagram illustrating an example of an intention-feature database corresponding to a non-linguistic expression model. [Figure 9] 10 is a flowchart illustrating an example of a feature conversion process. [Figure 10] FIG. 1A is a diagram showing an example of parameters of feature amounts of input non-language information, and FIG. 1B is a diagram showing an example of parameters of feature amounts of non-language information after conversion. [Figure 11] 10 is a flowchart illustrating an example of a conversion process of video data. [Figure 12] FIG. 10 is a schematic diagram illustrating an example of a non-language information conversion system according to a second embodiment. [Figure 13] FIG. 1A is a diagram showing an example of parameters of feature amounts of input non-language information, and FIG. 1B is a diagram showing an example of parameters of feature amounts of non-language information after conversion. [Figure 14] FIG. 10 is a schematic diagram illustrating an example of a non-language information conversion system according to a third embodiment. [Figure 15] FIG. 10 is a schematic diagram illustrating an example of a non-language information conversion system according to a fourth embodiment. [Figure 16] 13 is a flowchart illustrating an example of a feature conversion process according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In the description of the drawings, the same elements are given the same reference numerals, and duplicated explanations will be omitted.
[0011] ●Embodiment● Overview of the non-verbal information conversion system First, an outline of the configuration of a non-verbal information conversion system according to an embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the overall configuration of a non-verbal information conversion system. The non-verbal information conversion system 1 shown in Fig. 1 is a system that converts non-verbal information exchanged in interactive communication using video communication or the like.
[0012] As shown in FIG. 1, the non-verbal information conversion system 1 includes an information processing device 10 that converts non-verbal information, a transmitting device 70 used by a sender of the non-verbal information, and a receiving device 90 used by a receiver of the non-verbal information. The information processing device 10, the transmitting device 70, and the receiving device 90 that constitute the non-verbal information conversion system 1 can communicate via a communication network. The communication network is constructed using the Internet, a mobile communication network, a LAN (Local Area Network), or the like. Note that the communication network may include not only wired communication but also wireless communication networks such as 3G (3rd Generation), 4G (4th Generation), 5G (5th Generation), Wi-Fi (Wireless Fidelity) (registered trademark), WiMAX (Worldwide Interoperability for Microwave Access), or LTE (Long Term Evolution).
[0013] The information processing device 10 is a computer that converts non-verbal information so that the intention of a first user who is a sender is conveyed to a second user who is a receiver in an easily understandable manner. The information processing device 10 acquires data including the non-verbal information of the sender, converts the non-verbal information so that the intention of the sender is conveyed to the receiver in an easily understandable manner, and outputs converted data in which the non-verbal information of the acquired data has been converted.
[0014] Here, non-verbal information includes features such as the user's gaze, facial expression, upper limb posture, hand shape, arm or leg shape or posture, or voice tone or intonation. The sender's intention includes what the sender wants to convey to the receiver, from among the sender's state (pleasure, concentration, activity, etc.), the sender's emotion (joy, anger, sadness, pleasure, confusion, disgust, etc.), and the sender's will (command, refusal, request, etc.).
[0015] Furthermore, the non-verbal information conversion system 1 has a non-verbal expression model and a non-verbal recognition model constructed for each user to improve the conversion accuracy of non-verbal information. The non-verbal expression model stores the relationship between the user's non-verbal expression and intention, and is a model that learns the individuality of the user's non-verbal expression. The non-verbal recognition model stores the relationship between the user's non-verbal recognition and expression, and is a model that learns the individuality of the user's non-verbal recognition.
[0016] The non-verbal information conversion system 1 uses a non-verbal expression model and a non-verbal recognition model trained for each user to convert non-verbal information, thereby converting non-verbal information taking into account the individual characteristics of the expression and recognition of non-verbal information. For example, as shown in Fig. 1, the information processing device 10 reads the sender's non-verbal expression model and the receiver's non-verbal recognition model, and converts data including the acquired non-verbal information of the sender into non-verbal information that is easy for the receiver to recognize.
[0017] The information processing device 10 may be configured by one computer or by multiple computers. The information processing device 10 may be a computer that exists in a cloud environment or a computer that exists in an on-premise environment.
[0018] The transmitting device 70 and the receiving device 90 are computers such as laptop PCs (Personal Computers) used by the sender and the receiver, respectively, for interactive communication. The transmitting device 70 transmits, for example, video data of the sender photographed from the front to the information processing device 10. The receiving device 90 displays the video data (converted data) of the sender converted by the information processing device 10 on a display. Note that the transmitting device 70 and the receiving device 90 are not limited to laptop PCs and may be, for example, smartphones, tablet devices, wearable devices, or desktop PCs. Furthermore, while FIG. 1 illustrates an example of interactive communication between two people, the non-verbal information conversion system 1 is also applicable to interactive communication between three or more people. Furthermore, in interactive communication, the sender and the receiver may each play the role of both sender and receiver.
[0019] In the conversion of non-verbal information in conversational communication, there are known technologies that convert a user's facial expression captured by a camera, or user attributes such as gender or age, in real time, or that adjust the intonation and tone of a machine voice to change the impression given to the other person.Non-verbal information has a stronger influence on impressions than verbal information, and it is believed that using these technologies will make it possible to give the other person the desired impression.
[0020] Furthermore, in dialogue communication, it is important to correctly express and recognize non-verbal information. Problems can arise in dialogue when non-verbal information is not handled correctly or when the parties misinterpret the non-verbal information. In particular, in nursing care and special education settings, conflicts are likely to arise between staff and individuals who are unable to handle non-verbal information well. In such situations, there is often a discrepancy between the intended message and the non-verbal expression, or differences in the rules for non-verbal expression and recognition between the parties. In such situations, for example, technologies are known that allow a sender who does not know sign language to simply input voice input to synthesize an easy-to-understand sign language image for the other party, or to convert visual information into an easy-to-understand image for the visually impaired or partially sighted.
[0021] However, since there are individual and cultural differences in the recognition and expression of non-verbal information, estimation and expression should be tailored to each individual. Therefore, the conventional methods described above have room for improvement in converting non-verbal information while taking into account the individual differences in the expression and recognition of non-verbal information.
[0022] Therefore, the non-verbal information conversion system 1 receives as input video data representing the sender's non-verbal information, and estimates the sender's intention based on the sender's non-verbal expression model using the information processing device 10. The information processing device 10 then converts the sender's non-verbal information based on the estimated intention and a set conversion rule (for example, a conversion rule based on the receiver's non-verbal recognition model). The information processing device then outputs video data representing the converted non-verbal information, which is the conversion data, to the receiver. This allows the non-verbal information conversion system 1 to convey the sender's intention in a way that is easy to understand during a dialogue between the sender and the receiver.
[0023] ●Hardware configuration Next, the hardware configuration of each device constituting the non-language information conversion system according to the embodiment will be described with reference to Fig. 2. Each device constituting the non-language information conversion system 1 has the configuration of a general computer. Here, an example of the hardware configuration of a general computer will be described.
[0024] 2 is a diagram showing an example of the hardware configuration of a computer. The hardware configuration of the computer shown in FIG. 2 may have the same configuration in each embodiment, and components may be added or deleted as necessary. The computer includes a CPU (Central Processing Unit) 101, a ROM (Read Only Memory) 102, a RAM (Random Access Memory) 103, a HD (Hard Disk) 104, an HDD (Hard Disk Drive) controller 105, a display 106, an external device connection I / F (Interface) 107, a communication I / F 108, a bus line 110, a keyboard 111, a pointing device 112, an audio input / output I / F 113, a microphone 114, a speaker 115, a camera 116, a DVD-RW (Digital Versatile Disk Rewritable) drive 117, and a media I / F 119.
[0025] Of these, the CPU 101 controls the overall operation of the computer. The ROM 102 stores programs, such as an IPL, used to drive the CPU 101. The RAM 103 is used as a work area for the CPU 101. The HD 104 stores various data, such as programs. The HDD controller 105 controls the reading and writing of various data from and to the HD 104 under the control of the CPU 101. The display 106 is a type of display means for displaying various information, such as a cursor, menus, windows, characters, or images. The display 106 may also be a touch panel display equipped with an input means. The external device connection I / F 107 is an interface for connecting various external devices. The communication I / F 108 is an interface for transmitting and receiving data to and from other computers, electronic devices, etc. The communication I / F 108 is, for example, a communication interface such as a wired or wireless LAN. The communication I / F 108 may also include a communication interface for mobile communication, such as 3G, 4G, 5G, or LTE, Wi-Fi, WiMAX, etc. The bus line 110 is an address bus, a data bus, etc. for electrically connecting the components such as the CPU 101 shown in FIG.
[0026] The keyboard 111 is a type of input means having multiple keys for inputting characters, numbers, various instructions, etc. The pointing device 112 is a type of input means for selecting and executing various instructions, selecting a processing target, moving a cursor, etc. The input means may be not only the keyboard 111 and the pointing device 112, but also a touch panel, a voice input device, etc. The display means such as the display 106, and the input means such as the keyboard 111 and the pointing device 112 may be a UI (User Interface) external to the computer. The sound input / output I / F 113 is a circuit that processes input and output of sound signals between the microphone 114 and the speaker 115 under the control of the CPU 101. The microphone 114 is a type of built-in sound collection means that inputs sound. The speaker 115 is a type of built-in output means that outputs sound signals. The camera 116 is a type of built-in imaging means that captures an image of a subject and obtains image data. The microphone 114, the speaker 115, and the camera 116 may be external devices rather than built into the computer. The DVD-RW drive 117 controls reading and writing of various data from and to a DVD-RW 118, which is an example of a removable recording medium. The medium is not limited to a DVD-RW, and may be a DVD-R, a Blu-ray (registered trademark) Disc, or the like. The media I / F 119 controls reading and writing (storing) of data from and to a recording medium 121, such as a flash memory.
[0027] Each of the above programs may be recorded as an installable or executable file on a computer-readable recording medium and distributed. Examples of the recording medium include a CD-R (Compact Disc Recordable), a DVD (Digital Versatile Disk), a Blu-ray Disc, an SD card, or a USB memory. The recording medium may be provided domestically or internationally as a program product. For example, the information processing device 10 executes the program according to the present invention to realize the information processing method according to the present invention.
[0028] ●Function configuration Next, the functional configuration of the non-verbal information conversion system according to the embodiment will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of the functional configuration of an information processing device. The information processing device 10 has a data acquisition unit 11, a non-verbal information processing unit 12, and a data output unit 13. Each of these units is a function or a means for performing the function that is realized when any of the components shown in Fig. 2 operates in response to an instruction from the CPU 101 in accordance with a program for the information processing device deployed on the RAM 103.
[0029] The data acquisition unit 11 is mainly realized by the processing of the CPU 101 on the communication I / F 108 or the external device connection I / F 107, and acquires various data transmitted from the transmission device .
[0030] The non-language information processing unit 12 is mainly realized by the processing of the CPU 101 , and converts the non-language information indicated in the data acquired by the data acquisition unit 11 .
[0031] The data output unit 13 is mainly realized by the processing of the CPU 101 on the communication I / F 108 or the external device connection I / F 107, and outputs various data processed by the non-language information processing unit 12 to the receiving device 90.
[0032] Here, we will explain in detail the configuration of the non-verbal information processing unit 12. The non-verbal information processing unit 12 has a feature extraction unit 21, an intention estimation unit 22, a feature conversion unit 23, a conversion rule setting unit 24, a video conversion unit 25, a model learning unit 26, a data storage unit 31, an intention storage unit 32, a synthesis information storage unit 33, and a memory / readout unit 29. The non-verbal information processing unit 12 also has a storage unit 2000 constructed by the ROM 102, HD 104, or recording medium 121 shown in FIG.
[0033] The feature extraction unit 21 receives an image of a predetermined person as an input and extracts features of non-verbal information shown in the image.
[0034] The intention estimation unit 22 estimates the intention of the sender based on the non-verbal information of the sender indicated in the data acquired by the data acquisition unit 11 and the non-verbal expression model of the sender.
[0035] The feature conversion unit 23 converts the features of the sender's non-verbal information shown in the data acquired by the data acquisition unit 11 based on the intention estimated by the intention estimation unit 22 and the conversion rule set by the conversion rule setting unit 24.
[0036] The conversion rule setting unit 24 sets conversion rules for converting the feature quantities of the sender's non-verbal information.
[0037] The video conversion unit 25 converts the video data acquired by the data acquisition unit 11 based on the feature converted by the feature conversion unit 23 .
[0038] The model learning unit 26 performs learning of various learning models (non-language expression model 210, non-language recognition model 220, feature extraction model 230, and conversion model 240) used for converting non-language information.
[0039] The data storage unit 31 stores various data used for converting non-verbal information. The intention storage unit 32 stores the intention of the sender estimated by the intention estimation unit 22. The composite information storage unit 33 stores composite information corresponding to the video of the sender, which is the converted data converted by the video conversion unit 25.
[0040] The storage / readout unit 29 stores various data (or information) in the storage unit 2000 and reads various data (or information) from the storage unit 2000. The storage unit 2000 stores a non-language expression model 210, a non-language recognition model 220, a feature extraction model 230, and a conversion model 240, which are constructed by a conversion process of non-language information and learned by the model learning unit 26. Of these, the non-language expression model 210 and the non-language recognition model 220 are person-dependent, and therefore different models are stored for each user. The model parameters of the non-language expression model 210 and the non-language recognition model 220 can be adjusted, for example, based on parameters of extracted features, to perform desired recognition of non-language expressions and conversion of non-language recognition. Furthermore, the model parameters of the non-language expression model 210 and the non-language recognition model 220 can be adjusted, for example, based on the relationship between the sender and the receiver, based on parameters of the extracted features. On the other hand, the feature extraction model 230 and the conversion model 240 are not dependent on the person, and therefore, one model is stored for each. Note that the storage unit 2000 in which the various learning models are stored may be configured to be constructed in a storage device external to the information processing device 10.
[0041] First embodiment ○Outline○ Next, a non-verbal information conversion system according to the first embodiment will be described with reference to Figs. 4 to 11. Fig. 4 is a schematic diagram showing an example of the non-verbal information conversion system according to the first embodiment. The non-verbal information conversion system 1A according to the first embodiment is a system that converts non-verbal information using a non-verbal expression model of the sender and a non-verbal expression model of person A in interactive communication between a sender and a receiver. Person A is a person different from the sender and the receiver, and who behaves in a manner that gives a more favorable impression than the sender. Person A is an example of a third user.
[0042] First, the information processing device 10 acquires video data showing the sender and stores it in the data storage unit 31 of the non-verbal information processing unit 12. Next, the non-verbal information processing unit 12 reads the non-verbal expression model of the sender and estimates the sender's intention using the non-verbal expression model of the sender read from the video data of the sender stored in the data storage unit 31. Then, the non-verbal information processing unit 12 stores information indicating the estimated sender's intention in the intention storage unit 32.
[0043] Next, the non-verbal information processing unit 12 reads the non-verbal expression model of person A and converts the feature quantities of the non-verbal information shown in the sender's video data using the non-verbal expression model of person A read from the sender's intention stored in the intention storage unit 32. The non-verbal information processing unit 12 also converts the sender's video data based on the converted feature quantities to generate converted data that is composite information in which the video and label information are combined. The non-verbal information processing unit 12 stores the converted converted data, that is, composite information, in the composite information storage unit 33. The information processing device 10 then outputs the converted non-verbal information indicated by the composite information stored in the composite information storage unit 33 of the non-verbal information processing unit 12 to the recipient.
[0044] In this way, the non-verbal information conversion system 1A according to the first embodiment can improve the accuracy of non-verbal information conversion in interactive communication between a sender and a receiver by converting non-verbal information using the non-verbal expression model of person A, who is good at non-verbal expression, in addition to the non-verbal expression model of the sender.
[0045] Processing or Operation of the First Embodiment Next, the processing or operation of the non-language information conversion system according to the first embodiment will be described with reference to Fig. 5 to Fig. 11. First, the overall flow of the non-language information conversion processing executed by the non-language information conversion system 1A will be described with reference to Fig. 5. Fig. 5 is a flowchart showing an example of the non-language information conversion processing.
[0046] First, the data acquisition unit 11 of the information processing device 10 acquires video data captured by a sender, who is a first user (step S1). Specifically, the transmission device 70 used by the sender captures an image of the sender using the camera 116 and transmits the captured video data to the information processing device 10. Then, the data acquisition unit 11 of the information processing device 10 acquires the video data transmitted from the transmission device 70. The video of the sender is mainly, for example, a video of the sender's actions in everyday communication. The data acquisition unit 11 stores the acquired video data in the data storage unit 31 of the non-verbal information processing unit 12.
[0047] Next, the non-verbal information processing unit 12 of the information processing device 10 inputs the video data acquired in step S1 into the feature extraction model 230 and extracts features of non-verbal information (step S2). The features of non-verbal information include parameters such as the position of a person's facial landmarks, Action Unit (AU), the gaze angle of each eye, the position of skeletal landmarks, and the head rotation angle or distance.
[0048] Next, the non-verbal information processing unit 12 of the information processing device 10 inputs the feature amounts extracted in step S2 into the non-verbal expression model 210 and estimates the intention of the sender. The non-verbal information processing unit 12 acquires intention label information indicating the intention of the sender, for example, using the non-verbal expression model 210 to which the extracted feature amounts have been input. The intention label information includes the six basic emotions of "happy, angry, sad, surprised, disgusted, and composed," as well as information on the intensity of "Neutral, Deny, Accept, Arousal, and Interested," expressed as "0 to 1."
[0049] Next, the non-verbal information processing unit 12 of the information processing device 10 inputs the intention estimated in step S3 into the conversion rule set in the conversion rule setting unit 24 to convert the feature and obtain the desired feature (step S4).
[0050] Next, the non-verbal information processing unit 12 of the information processing device 10 inputs the features converted in step S4 into the conversion model 240, converts the video data acquired in step S1, and acquires the converted video as converted data (step S5).
[0051] Then, the data output unit 13 of the information processing device 10 outputs the converted video, which is the converted data converted in step S5, to the recipient, who is the second user (step S6). Specifically, the data output unit 13 transmits the converted video to the receiving device 90 used by the recipient, and the receiving device 90 displays the converted video transmitted (output) from the information processing device 10 on the display 106.
[0052] Feature extraction processing Here, the details of each process shown in Fig. 5 will be described with reference to Fig. 6 to Fig. 11. First, the feature extraction process executed by the non-language information processing unit 12 of the information processing device 10 in step S2 will be described with reference to Fig. 6. Fig. 6 is a flowchart showing an example of the process of extracting features of non-language information.
[0053] First, the feature extraction unit 21 inputs a general person video, which is a video of a person unrelated to the sender or person A (step S21). The general person video is a video of a third person (general person) different from the sender, person A, and the receiver. The general person video is a video of a general person making various changes to their facial expression, body or head orientation, gaze direction, or distance from the camera. Next, the feature extraction unit 21 performs an annotation process using the general person video input in step S21 (step S22). The feature extraction unit 21 defines non-verbal information features as label information, for example, by inputting data from the annotator using the keyboard 111 or the like. As a result, the feature extraction unit 21 creates a dataset to be used in the feature extraction process. The feature extraction unit 21 stores the created dataset in the data storage unit 31.
[0054] Next, the feature extraction unit 21 constructs a feature extraction model 230 used for feature extraction (step S23). The feature extraction model 230 is configured with a hierarchy of an input layer, a convolutional neural network (CNN) layer, a long short-term memory (LSTM) layer, and an estimation layer. The feature extraction model 230 extracts image features for each frame of the input video using the CNN layer. The feature extraction model 230 also extracts non-verbal feature and state information for each frame using the LSTM layer as input features of multiple frame images in the input video. Here, the state information for each frame corresponds to memory information transmitted to the next cell in LSTM processing and represents context information in LSTM, which performs text analysis. The state information represents, for example, the state of a person, such as whether they are stationary or moving. By combining a CNN layer and an LSTM layer, the feature extraction model 230 can simultaneously analyze each frame image and time-series information.
[0055] Next, the feature extraction unit 21 reads the data set created by the processes of steps S21 and S22 for training the feature extraction model 230 (step S24). As a result, the model training unit 26 trains the feature extraction model 230 used for feature extraction. Note that, because the extraction of features of non-verbal information from input video does not change depending on the person, training of the feature extraction model 230 needs to be performed only once, and if a trained feature extraction model 230 exists, the processes of steps S23 and S24 do not need to be performed.
[0056] Next, the feature extraction unit 21 inputs the video data acquired in step S1 to the feature extraction model 230 trained in the processes of steps S23 and S24 (step S25). Then, the feature extraction unit 21 acquires features of non-verbal information indicated by the video data acquired in step S1 (step S26).
[0057] It should be noted that the processes in steps S21 to S25 described above can be omitted by using a known technique that can acquire feature amounts in real time, such as OpenFace described in Non-Patent Document 1.
[0058] Intention estimation processing Next, the intention estimation process executed by the non-verbal information processing unit 12 of the information processing device 10 in step S3 will be described with reference to Figures 7 and 8. Figure 7 is a flowchart showing an example of the process of estimating the sender's intention.
[0059] First, the intention estimation unit 22 inputs the video data of the sender acquired in step S1 (step S31). Then, the intention estimation unit 22 executes an annotation process using the video data input in step S21 (step S32). The intention estimation unit 22 defines the intention corresponding to the video of the sender as intention label information, for example, by input operation of the annotator on the keyboard 111 or the like. An example of this intention label information L is shown below (Equation 1).
[0060] Intention label information L = {angry;0.1, composed;0.2, disgusted;0.2, happy;0.8, sad;0.4, surprised;0.6, neutral;0.2, deny;0.3, accept;0.3, arousal;0.7, interested;0.8} (Equation 1)
[0061] The annotation in step S32 is performed by the sender (annotator = sender), who defines the type and strength of intention as shown in (Equation 1). The annotator, for example, plays back the video data input in step S31 and inputs the numerical value of the strength of intention for each frame of the video. The annotation can be performed, for example, by using a dedicated application and dragging an input unit such as the pointing device 112 to specify the numerical value of the strength of intention for each frame, thereby reducing the burden of the annotation process. The annotation is not limited to the strength of intention, and may be configured to specify multiple types and dimensions of information, such as the type of intention or the degree of certainty, by dragging an input unit such as the pointing device 112. In this way, the intention estimation unit 22 creates a dataset to be used in the intention estimation process. The intention estimation unit 22 stores the created dataset in the data storage unit 31.
[0062] Next, the intention estimation unit 22 constructs a non-verbal expression model 210 for the sender to be used for estimating the intention (step S33). Here, because the expression of intention from the feature amounts of non-verbal information depends on the person, the intention estimation unit 22 constructs the non-verbal expression model 210 for the sender. The structure of the non-verbal expression model 210 is the same regardless of the person, and is composed of a hierarchy of an input layer, an LSTM layer, and an estimation layer. As preprocessing, the intention estimation unit 22 inputs the feature amounts of the non-verbal information extracted in step S26 into the non-verbal expression model 210. The intention estimation unit 22 inputs the feature amounts of multiple frame images in the input video using the LSTM layer of the non-verbal expression model, and outputs the intention and frame number for each frame. The frame number indicates the ordinal number of the input frame among the multiple frames indicating the intention.
[0063] Here, the non-language expression model 210 is a learning model that indicates the relationship between intentions and feature quantities of non-language information, and has, for example, a database-like structure. Here, for convenience, it is referred to as an intention-feature quantity database. FIG. 8 is a conceptual diagram showing an example of the intention-feature quantity database corresponding to the non-language expression model. As shown in FIG. 8, the feature quantities of non-language information have time-series values for each strength of intention (here, 1.0, 0.8, 0.5). Furthermore, for the feature quantities of non-language information ((1) to (7) shown in FIG. 8), there are four possible time-series values for each strength of intention (occurrence probability: 0.3, 0.25, 0.25, 0.2).
[0064] For example, parameter (1) corresponding to AU1 (ActionUnit1) is expressed as in the following (Equation 2). As shown in (Equation 2), parameter (1) includes values for N frames (for example, N=10) for each of four occurrence probabilities. Other parameters (2) to (7) are also expressed as in (Equation 2), similar to parameter (1). Thus, for example, if only one frame worth of parameters (1) to (7) with an occurrence probability of 0.3 is extracted, the result is as in the following (Equation 3).
[0065] (1)={{0.3,0.3,0.2,…} 0.3 ,{0.3,0.3,0.1,…} 0.25 ,{0.3,0.1,0.2,…} 0.25 ,{0.3,0.3,0.0,…} 0.2}...(Formula 2)
[0066] {(1),(2),(3),(4),(5),(6),(7)}={0.3,0.2,0.6,0.1,0.1,0.5,0.5}...(Formula 3)
[0067] Next, the intention estimation unit 22 reads the data set created by the processes of steps S31 and S32 for training the non-verbal expression model 210 (step S34). As a result, the model training unit 26 trains the non-verbal expression model 210 used for intention estimation. Since the expression of intention from the feature amount of non-verbal information depends on the person, the training of the non-verbal expression model 210 is performed for each person.
[0068] Next, since the expression of intention is person-dependent based on the feature quantities of non-verbal information, the intention estimation unit 22 reads the non-verbal expression model 210 for the sender trained in steps S33 and S34 (step S35). Then, the intention estimation unit 22 estimates the intention of the sender based on the non-verbal information indicated by the video data acquired in step S1 (step S36). The estimated intention also includes multidimensional information such as the type, strength, or certainty of the intention. For example, the intention estimation unit 22 acquires intention label information that serves as an estimate of the intention obtained by inputting the video data of the sender into the non-verbal expression model 210 for the sender. The intention estimation unit 22 stores the information of the estimated intention in the intention storage unit 32.
[0069] In this way, the intention estimation unit 22 can estimate the sender's intention based on the sender's video data and the sender's non-verbal expression model, thereby estimating the intention while taking into account individual differences in the expression of non-verbal information and differences due to cultural differences.
[0070] Feature conversion processing Next, the feature conversion process executed by the non-language information processing unit 12 of the information processing device 10 in step S4 will be described with reference to Figures 9 and 10. Figure 9 is a flowchart showing an example of the feature conversion process.
[0071] First, the conversion rule setting unit 24 sets a conversion rule for converting non-language information (step S41). The conversion rule set by the conversion rule setting unit 24 includes a conversion item, a conversion ratio, a conversion destination, and an item of the intention-feature database (see FIG. 8) corresponding to the non-language expression model 210 of the conversion destination.
[0072] Among these, the conversion items indicate the type of non-verbal information feature to be converted. The types of non-verbal information feature include, for example, posture information, gaze information, facial expression intensity for each emotion, and head rotation angle. Here, posture information refers only to the position of the skeletal landmarks of the non-verbal information feature that corresponds to the spine. By not including all facial landmarks and skeletal landmarks in the conversion items, it is possible to convert the spine and gaze while maintaining individual differences in face and physique.
[0073] The conversion ratio indicates how similar the target person should be, and is defined as a value between 0 and 1. For example, select "1" to make the target person look like Person A, or select "0" to maintain the original state before conversion.
[0074] Furthermore, the conversion destination indicates the person to whom the conversion is to be made. Here, for example, person A, who behaves in a manner that gives a more favorable impression than the sender, is set. Then, the association between person A's intention and the features of non-verbal information is defined as an intention-feature database corresponding to the non-verbal expression model 210. The definition method is performed by constructing and learning the non-verbal expression model 210 of person A, similar to the process shown in FIG. 7. The definition method is performed, for example, similar to the creation of the dataset in steps S31 and S32, by inputting a video of person A and having person A annotate the intention corresponding to each frame of the input video. Note that the structure of the intention-feature database corresponding to person A's non-verbal expression model 210 is similar to the example shown in FIG. 8, but the parameters are different from those of the sender's non-verbal expression model.
[0075] Then, the feature conversion unit 23 applies the conversion rule set in step S41 and converts the features through the processes of steps S42 to S44. Specifically, the feature conversion unit 23 selects the sender's intention estimated through the process of step S3 (step S42). The feature conversion unit 23 selects the type, strength, and frame number of the intention from the sender's video data based on the estimated intention. The feature conversion unit 23 selects, for example, the intention with the greatest strength as the intention of that frame. If there are multiple intentions with the greatest strength as in the above-mentioned (Equation 1), both of them are selected, and the feature of the non-language information intermediate between them is calculated by linear interpolation, which will be described later. In this case, the ratios are both set to 0.5. The feature conversion unit 23 selects the row (record) of the corresponding intention from the non-language expression model shown in FIG. 8, and selects the feature of the non-language information of the corresponding frame number from the corresponding row (record).
[0076] Next, the feature conversion unit 23 probabilistically selects (probabilistic selection) one of the four possible time-series values as shown in the above (Equation 2) (step S43). Then, the feature conversion unit 23 performs linear interpolation of the features of the non-verbal information of the sender and the conversion destination using the following (Equation 4) according to the conversion ratio indicated in the conversion rule set in step S41 (step S44). Here, in (Equation 4), X1 indicates the feature of the non-verbal information of the sender, X2 indicates the feature of the non-verbal information of the conversion destination, and α indicates the conversion ratio.
[0077] X=α×X1+(1-α)×X2...(Formula 4)
[0078] Fig. 10(A) is a diagram showing an example of parameters of feature quantities of input non-verbal information. Fig. 10(A) shows some of the parameters of feature quantities of non-verbal information extracted from video data of the sender before conversion and an example of frame numbers estimated therefrom. The feature quantity conversion unit 23 receives the parameters shown in Fig. 10(A), for example, and converts the features based on the intention-feature quantity database corresponding to the non-verbal expression model shown in Fig. 8.
[0079] For example, the feature conversion unit 23 selects {Neutral, 1.0} as the intention of the sender in step S42, and converts the time series value {} 0.3 is assumed to be probabilistically selected. The conversion items shown in the set conversion rule are only (4) to (7) shown in FIG. 8, and the conversion ratio is α=1.0. An example of the parameters of the feature quantities of non-verbal information after conversion in this case is shown in FIG. 10(B). As shown in FIG. 10(B), the parameters (4) to (7) have been converted from the parameters in FIG. 10(A), but the parameters (1) to (3) remain the same as before conversion. In other words, the feature quantity conversion unit 23 converts only the values of the posture and gaze direction of the sender without changing the facial expression of the sender.
[0080] In this way, the feature conversion unit 23 converts the parameters of the features of non-verbal information based on the estimated value of the intention estimated by the intention estimation unit 22 and the conversion rule set by the conversion rule setting unit 24 so as to increase the probability that the sender's intention will be correctly conveyed to the receiver.
[0081] Video data conversion processing Next, the video data conversion process executed by the non-language information processing unit 12 of the information processing device 10 in step S5 will be described with reference to Fig. 11. Fig. 11 is a flowchart showing an example of the video data conversion process.
[0082] First, the video conversion unit 25 inputs a video of a general person and the feature amounts of non-verbal information corresponding to the video of the general person (step S51). The video of the general person input here is mainly a video of a general person performing actions that change only the feature amounts of non-verbal information in various ways. Actions that change only the feature amounts include, for example, changing the direction of gaze or the direction of the head. In this case, label information is not required. As a result, the video conversion unit 25 creates a dataset for performing conversion processing on the video data. The video conversion unit 25 stores the created dataset in the data storage unit 31.
[0083] Next, the video conversion unit 25 constructs a conversion model 240 for converting the video data (step S52). The conversion model 240 is an extension of the GANimation method described in Non-Patent Document 2 to include not only facial expressions but also a person's posture and gaze information. GANimation is a technology that converts an input video into an video with a desired expression label by inputting not only an input video but also a set of AU feature intensities, which are expression labels, into an image generation network. The conversion model 240 is realized by extending the GANimation method to non-verbal information features other than AU features.
[0084] The video conversion unit 25 inputs a set of "pre-conversion video, pre-conversion non-language information features, and post-conversion desired non-language information features" into the conversion model 240, and outputs a set of "post-conversion video, post-conversion non-language information features." This differs from a typical GAN (Generative Adversarial Network) in that it also inputs desired post-conversion label information. The loss function is calculated as the mean square error between the post-conversion video and post-conversion non-language information features and the desired video and non-language information features.
[0085] Next, the video conversion unit 25 reads the data set created by the processing of step S51 for training the conversion model 240 (step S53). As a result, the model training unit 26 trains the conversion model 240 used for converting the video data. Note that since the conversion of video data does not change depending on the person, training of the conversion model 240 needs to be performed only once, and if a trained conversion model 240 exists, the processing of steps S52 and S53 does not need to be performed.
[0086] Next, the video conversion unit 25 converts the video data acquired in step S1 based on the conversion model read in step S53 (step S54). The video conversion unit 25 converts the video data for each frame. The video conversion unit 25 stores the converted data, which is synthesis information in which the video and the intention label information are synthesized, in the synthesis information storage unit 33.
[0087] In this way, the video conversion unit 25 converts the video data of the sender so as to increase the probability that the estimated value of the intention estimated by the intention estimation unit 22 will be correctly conveyed to the receiver based on the non-linguistic expression model 210.
[0088] As described above, the non-verbal information conversion system 1A according to the first embodiment estimates the intention of the sender based on the sender's video data and a non-verbal expression model for the sender, and converts the sender's video data based on the estimated intention of the sender and a conversion rule based on a non-verbal expression model of a person who is good at non-verbal expression. As a result, the non-verbal information conversion system 1A according to the first embodiment can improve the conversion accuracy of non-verbal information to convey the sender's intention to the receiver in an easy-to-understand manner during a dialogue between the sender and the receiver.
[0089] Second embodiment Next, a non-verbal information conversion system according to a second embodiment will be described with reference to FIGS. 12 and 13. The same configurations and functions as those of the above-described embodiment are denoted by the same reference numerals, and their description will be omitted. FIG. 12 is a schematic diagram showing an example of a non-verbal information conversion system according to the second embodiment. The non-verbal information conversion system 1B according to the second embodiment differs from the non-verbal information conversion system 1A according to the first embodiment in that it converts non-verbal information using a non-verbal expression model of the sender and a specific correction value used for conversion in dialogue communication between a sender and a receiver. When converting features in step S4, the non-verbal information conversion system 1B according to the second embodiment sets conversion rules by directly specifying the items to be corrected and the correction guideline values, rather than selecting a specific person image as the conversion destination.
[0090] Here, the feature conversion process of step S4 in the second embodiment, which differs from that in the first embodiment, will be described. In the second embodiment, the conversion rule set by the conversion rule setting unit 24 in step S41 includes a conversion item, its value, and a conversion ratio. Among these, the conversion item indicates the type of non-language information feature to be changed, and the value of the conversion item indicates the value of a correction guide value. The change item and correction guide value are, for example, posture information (0.0, 0.0) and gaze information (0.0, 0.0).
[0091] The conversion ratio indicates how similar the image should be to the destination image, and is set to a small value to ensure smooth image conversion. The conversion ratio is, for example, 0.5.
[0092] 13 shows an example in which the conversion rule set in step S41 is applied in the second embodiment. Fig. 13(A) is a diagram showing an example of parameters of the feature amounts of input non-language information, and Fig. 13(B) is a diagram showing an example of parameters of the feature amounts of the converted non-language information. The feature amount conversion unit 23 receives the parameters shown in Fig. 13(A), for example, and converts the feature amounts based on the conversion rule set in step S41.
[0093] As shown in Fig. 13(B), the parameters (4) to (7) approach the correction guideline values indicated in the conversion rule from the parameters in Fig. 13(A), but the parameters (1) to (3) remain the same as before conversion. In other words, the feature conversion unit 23 can always perform conversion such that the posture and gaze approach the correction guideline values, regardless of, for example, intention.
[0094] In this way, the non-verbal information conversion system 1B according to the second embodiment converts the sender's video data based on the conversion rule that uses the estimated intention of the sender and the correction guideline values of the conversion items. As a result, the non-verbal information conversion system 1B according to the second embodiment can convert non-verbal information by specifying specific numerical values of the items to be converted, thereby converting non-verbal information without being bound by the intention of the non-verbal information.
[0095] Third embodiment Next, a non-verbal information conversion system according to a third embodiment will be described with reference to FIG. 14. The same components and functions as those in the above-described embodiments are denoted by the same reference numerals, and their description will be omitted. FIG. 14 is a schematic diagram illustrating an example of a non-verbal information conversion system according to the third embodiment. The non-verbal information conversion system 1C according to the third embodiment differs from the non-verbal information conversion system 1A according to the first embodiment in that it converts non-verbal information using a non-verbal expression model of the sender and a non-verbal expression model of a general person during dialogue communication between a sender and a receiver. When converting features in step S4, the non-verbal information conversion system 1C according to the third embodiment sets a conversion rule by specifying a general person as the conversion destination. A general person is a person having an intention-feature database corresponding to an average non-verbal expression model.
[0096] Here, the feature conversion process of step S4 in the third embodiment, which differs from that in the first embodiment, will be described. In the third embodiment, the conversion rule set by the conversion rule setting unit 24 in step S41 includes a conversion item, a conversion ratio, a conversion destination, and an item in the intention-feature database (see FIG. 8) corresponding to the non-language expression model 210 of the conversion destination. Among these, the conversion item indicates the type of feature of non-language information to be converted. The type of feature of non-language information includes, for example, posture information, gaze information, facial expression intensity for each emotion, and head rotation angle.
[0097] The conversion ratio indicates how similar the target person should be, and is defined as a value between "0 and 1." For example, "1" is selected to make the target person resemble the target person, and "0" is selected to maintain the state before conversion. Here, the conversion ratio is defined as "1," for example, to make the target person resemble a general person.
[0098] Furthermore, the conversion destination indicates a person to be converted. Here, a generic person is set as the conversion destination. Then, the conversion rule setting unit 24 defines the association between the intention of a generic person and the feature of non-verbal information as an intention-feature database corresponding to the non-verbal expression model 210. The definition method is performed by constructing and learning the non-verbal expression model 210 of a generic person, similar to the process shown in FIG. 7. The definition method is performed, for example, similar to the creation of the dataset in steps S31 and S32, by inputting a video of an arbitrary person and having the person annotate the intention corresponding to each frame of the input video. The non-verbal information processing unit 12 performs this definition method multiple times and creates an intention-feature database corresponding to the non-verbal expression model 210 of a generic person by averaging each feature. The feature conversion unit 23 converts the feature by applying the conversion rule set by the conversion rule setting unit 24. The subsequent processes are the same as those in steps S42 to S44 in the first embodiment.
[0099] In this way, the non-verbal information conversion system 1C of the third embodiment can clearly convey the sender's intention to the receiver, even when converting the sender's video data based on the estimated sender's intention and conversion rules based on a non-verbal expression model of a typical person.
[0100] Fourth embodiment Next, a non-language information conversion system according to a fourth embodiment will be described with reference to FIGS. 15 and 16. The same configurations and functions as those of the above-described embodiments are denoted by the same reference numerals, and their description will be omitted. FIG. 15 is a schematic diagram showing an example of a non-language information conversion system according to the fourth embodiment. The non-language information conversion system 1D according to the fourth embodiment differs from the non-language information conversion system 1A according to the first embodiment in that it converts non-language information using a non-language expression model of the sender and a non-language recognition model of the receiver in interactive communication between the sender and the receiver. When converting features in step S4, the non-language information conversion system 1D according to the fourth embodiment converts the features of the sender's non-language information so that it is easily recognized by the receiver by using the receiver's non-language recognition model.
[0101] Here, the feature conversion process in step S4 in the fourth embodiment, which is different from that in the first embodiment, will be described in detail. Fig. 16 is a flowchart showing an example of the feature conversion process in the fourth embodiment.
[0102] First, the feature conversion unit 23 inputs a general person video, which is a video of a person unrelated to the sender and the receiver (step S101). The general person video is a video of a third person (general person) different from the sender and the receiver. The general person video is mainly a video of the actions that general people perform during everyday communication. Next, the feature conversion unit 23 performs an annotation process using the general person video input in step S101 (step S102). The feature conversion unit 23 defines an intention corresponding to the general person video as intention label information, for example, by input operation of the annotator on the keyboard 111 or the like. This intention label information is similar to the example shown in (Equation 1) above.
[0103] The annotation in step S102 is performed by the receiver (annotator = receiver), and defines the type and strength of intention as shown in (Equation 1). The annotator, for example, plays back the video data input in step S101 and inputs the numerical value of the strength of intention for each frame of the video. The annotation can be performed, for example, by using a dedicated application and dragging an input unit such as the pointing device 112 to specify the numerical value of the strength of intention for each frame, thereby reducing the burden required for the annotation process. The annotation is not limited to the strength of intention, and may be configured to specify multiple types and dimensions of information, such as the type of intention or the degree of certainty, by dragging an input unit such as the pointing device 112. As a result, the feature conversion unit 23 creates a dataset to be used in the feature extraction process. The feature conversion unit 23 stores the created dataset in the data storage unit 31.
[0104] Next, the feature conversion unit 23 constructs a non-language recognition model 220 for the recipient to be used for converting the features (step S103). Here, because recognizing intention from the features of non-language information depends on the person, the feature conversion unit 23 constructs the non-language recognition model 220 for the recipient. The structure of the non-language recognition model 220 is the same regardless of the person, and is composed of a hierarchy of an input layer, an LSTM layer, and an estimation layer. As preprocessing, the feature conversion unit 23 inputs the features of the non-language information extracted in step S26 to the non-language recognition model 220. The feature conversion unit 23 inputs the features of multiple frame images in the input video using the LSTM layer of the non-language recognition model 220, and outputs the intention and frame number for each frame. The frame number indicates the ordinal number of the input frame among multiple frames indicating the intention.
[0105] Here, the non-language recognition model 220 is a learning model that indicates the relationship between non-language recognition and expression, and has, for example, a database-like structure. Here, for convenience, it is referred to as an intention-feature database, and the structure of the intention-feature database corresponding to the non-language recognition model 220 is similar to the structure of the intention-feature database of the non-language expression model 210 shown in FIG. 8.
[0106] Next, the feature conversion unit 23 reads the data set created by the processes of steps S101 and S102 for training the non-language recognition model 220 (step S104). As a result, the model training unit 26 trains the non-language recognition model 220 used for feature conversion. Because recognition of intention from the feature of non-language information depends on the person, training of the non-language recognition model 220 is performed for each person. Next, because recognition of intention from the feature of non-language information depends on the person, the feature conversion unit 23 reads the non-language recognition model 220 for the recipient trained in steps S103 and S104 (step S105).
[0107] Next, the conversion rule setting unit 24 sets a conversion rule for converting non-language information (step S106). The conversion rule set by the conversion rule setting unit 24 includes items of the intention-feature database (see FIG. 8) corresponding to the conversion item, the conversion ratio, and the non-language recognition model 220 of the conversion destination.
[0108] Among these, the conversion item and conversion ratio indicate the type of feature of non-verbal information to be converted. The types of feature of non-verbal information include, for example, posture information, gaze information, facial expression intensity for each emotion, and head rotation angle. The conversion ratio indicates the degree to which the conversion should resemble the target, and is defined as a value between 0 and 1. For example, to resemble the receiver, select 1, and to maintain the state before conversion, select 0. Furthermore, an intention-feature database corresponding to the non-verbal recognition model 220 constructed by the processing of steps S101 to S105 defines the relationship between the intention that is easily recognized by the receiver and the feature of non-verbal information.
[0109] Then, the feature transform unit 23 performs the processes of steps S107 to S109 to transform the features by applying the transformation rule set in step S106. Note that the processes of steps S107 to S109 are the same as the processes of steps S42 to S44 in Fig. 9, respectively, and therefore description thereof will be omitted.
[0110] In this way, the non-verbal information conversion system 1D according to the fourth embodiment converts the sender's video data based on the estimated intention of the sender and the conversion rule based on the receiver's non-verbal recognition model. As a result, the non-verbal information conversion system 1D according to the fourth embodiment can improve the conversion accuracy of non-verbal information by using both the intention that the sender wants to convey and the recognition of the receiver's non-verbal expressions, thereby conveying the sender's intention to convey to the receiver in an easy-to-understand manner.
[0111] Effect of the embodiment As described above, the non-verbal information conversion system 1 (1A, 1B, 1C, 1D) converts the non-verbal information shown in the sender's video data using a non-verbal expression model and a non-verbal recognition model that differ for each person, thereby converting non-verbal information taking into account the individuality of the expression and recognition of non-verbal information.The non-verbal information conversion system 1 (1A, 1B, 1C, 1D) converts non-verbal information taking into account the individuality of each person in interactive communication, thereby making it possible to convey the sender's intention to communicate to the receiver in an easy-to-understand manner.
[0112] ●Additional Information● Each function of the above-described embodiments can be realized by one or more processing circuits. Here, the term "processing circuit" in the present embodiment includes a processor programmed to perform each function by software, such as a processor implemented by an electronic circuit, as well as devices designed to perform each function described above, such as an application specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), a system on a chip (SOC), a graphics processing unit (GPU), and a conventional circuit module.
[0113] So far, we have explained an information processing device, a non-verbal information conversion system, an information processing method, and a program according to one embodiment of the present invention, but the present invention is not limited to the above-mentioned embodiment, and other embodiments can be added, modified, or deleted, etc., within the scope that can be conceived by a person skilled in the art, and any aspect is included in the scope of the present invention as long as it achieves the functions and effects of the present invention. [Explanation of symbols]
[0114] 1(1A,1B,1C,1D) Non-verbal information conversion system 10. Information processing equipment 11 Data acquisition unit (an example of acquisition means) 12 Non-verbal information processing section 13 Data output unit (an example of an output means) 21 Feature extraction unit (an example of feature extraction means) 22 Intention estimation unit (an example of an intention estimation means) 23 Feature conversion unit (an example of non-linguistic information conversion means) 24 Conversion rule setting section 25 Video conversion unit (an example of non-verbal information conversion means) 70 Transmitting device 90 Receiving device 210 Nonverbal Expression Model 220 Non-verbal Cognition Model 230 Feature Extraction Model 240 conversion model
Claims
1. an intention estimation means for estimating an intention of a first user indicated in the first non-language information, based on first non-language information that is non-language information of a first user and a first non-language model that indicates a relationship between the first non-language information and an intention; a non-language information conversion means for converting the first non-language information into second non-language information to be output to a second user based on the estimated intention of the first user and a second non-language model indicating a relationship between second non-language information, which is non-language information of a second user, and the intention; An information processing device comprising:
2. 2. The information processing device according to claim 1, an acquisition means for acquiring video data captured by the first user; an output unit that outputs converted data obtained by converting the acquired video data, the intention estimation means estimates an intention of the first user based on the first non-verbal information shown in the video data; the non-language information conversion means converts the video data representing the first non-language information into the conversion data representing the second non-language information; The output means is an information processing device that outputs the video related to the converted data to a receiving device used by the second user.
3. the first non-verbal model is a model that learns the personality of the first user in non-verbal expressions, The information processing apparatus according to claim 2 , wherein the intention estimation means calculates an estimated value of the first user's intention obtained by inputting the video data acquired by the acquisition means into the first non-language model.
4. the second non-language model is a model that learns the personality of a second user in non-language recognition, The information processing device according to claim 3, wherein the non-verbal information conversion means converts the acquired video data so as to increase the probability that the calculated estimated value of the first user's intention will be correctly conveyed to the second user based on the second non-verbal model.
5. 5. The information processing device according to claim 3, a feature extraction means for extracting a feature of the first non-language information; An information processing apparatus in which the first non-language model and the second non-language model are adjusted based on parameters of the extracted feature amounts.
6. The information processing device according to claim 1 , wherein the first non-verbal information and the second non-verbal information include at least one feature of gaze or facial expression, hand, arm or leg shape, and posture.
7. The information processing device according to claim 1 , wherein the intention of the first user indicates a feeling or intention that the first user wants to convey to the second user.
8. The information processing device according to claim 7 , wherein the first user's intention includes a type or intensity of the emotion.
9. an intention estimation means for estimating an intention of a first user indicated in the first non-language information, based on first non-language information that is non-language information of a first user and a first non-language model that indicates a relationship between the first non-language information and an intention; a non-language information conversion means for converting the first non-language information into second non-language information to be output to a second user based on the estimated intention of the first user and a second non-language model indicating a relationship between second non-language information, which is non-language information of a second user, and the intention; A non-verbal information conversion system comprising:
10. An information processing method executed by an information processing device, an intention estimation step of estimating an intention of a first user indicated in first non-language information based on first non-language information that is non-language information of a first user and a first non-language model indicating a relationship between the first non-language information and an intention; a non-language information conversion step of converting the first non-language information into second non-language information to be output to a second user based on the estimated intention of the first user and a second non-language model indicating a relationship between second non-language information, which is non-language information of a second user, and the intention; An information processing method including:
11. an intention estimation step of estimating an intention of a first user indicated in first non-language information based on first non-language information that is non-language information of a first user and a first non-language model indicating a relationship between the first non-language information and an intention; a non-language information conversion step of converting the first non-language information into second non-language information to be output to a second user based on the estimated intention of the first user and a second non-language model indicating a relationship between second non-language information, which is non-language information of a second user, and the intention; A program that causes a computer to execute the following.
Citation Information
Patent Citations
Interface device
JP2001083984A
Information processing device, information processing method and program
JP2020021025A
Image processing apparatus, camera apparatus, and image processing method
JP2020048149A
JP330958A
Systems and methods for enhancement of facial expressions
US20140376785A1