Information processing device and program

The information processing apparatus uses emotion estimation and adaptive content generation to create personalized content that aligns with viewer emotions, addressing the risk of quality degradation in existing methods.

JP2026073732APending Publication Date: 2026-05-01吉元 行法
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
吉元 行法
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing content generation methods risk degrading the quality by adjusting emotions or thoughts of viewers, leading to inappropriate content.

Method used

An information processing apparatus that acquires user state data, estimates emotions, and generates partial content based on generation conditions to ensure emotional alignment with viewer preferences, using multimodal emotion estimation and adaptive content generation.

Benefits of technology

Ensures reliable generation of personalized content that aligns with viewer emotions, maintaining content quality and emotional integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073732000001_ABST
    Figure 2026073732000001_ABST
Patent Text Reader

Abstract

We provide technology to more reliably generate content that is appropriate for the viewer. [Solution] The user terminal 3 transmits state data representing the user's state, acquired by the camera 6, headgear 8, and wristband 9. The information processing device 1 uses this state data to perform multimodal emotion estimation for the user while they are viewing content and reflects this in the generation of content. The generation conditions for content generation refer to the target emotion targeted by that content, and generate the content so that the target emotion is actually evoked in the user, or so that the user is evoked with emotional intonation that matches the user's sensitivity. The target emotion may be determined by considering the content being viewed and the user's sensitivity tendencies identified through emotion estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus and a program.

Background Art

[0002] Currently, various contents are being distributed via a network. In recent years, there has also been an attempt to understand the thoughts and emotions of content viewers and adjust, for example, edit or improve, the content that the viewers actually view accordingly to generate new content (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] By adjusting the content, it is possible to personalize the content to be more desirable for the viewer. However, there is some content that must to some extent ignore the thoughts or emotions of the viewer. For example, in movies, it is common to produce while assuming the emotions evoked in the viewer by the scenes. In such content where the ebb and flow of emotions is assumed, there is a risk of significantly degrading the quality of the content by making adjustments according to either the thoughts or emotions of the viewer. As a result, there is also a high possibility that appropriate content for the viewer cannot be generated.

[0005] The present invention provides a technique for more reliably generating appropriate content for the viewer.

Means for Solving the Problems

[0006] An information processing apparatus according to one aspect of the present invention includes: a data acquisition unit that acquires state data representing the state of a user; an emotion estimation unit that estimates the user's emotions using the state data acquired by the data acquisition unit; a condition generation unit that generates generation conditions for the partial content when the user is allowed to view content in which there is one or more partial content to be generated, based on the emotion estimation result by the emotion estimation unit and the target emotion to be targeted for the user; and a content generation unit that generates the partial content based on the generation conditions generated by the condition generation unit. [Effects of the Invention]

[0007] This invention makes it possible to more reliably generate content that is appropriate for the viewer. [Brief explanation of the drawing]

[0008] [Figure 1] This diagram illustrates an overview of the services that can be realized by an information processing device according to one embodiment of the present invention. [Figure 2] This figure illustrates an example of a network system configuration using an information processing device according to one embodiment of the present invention. [Figure 3] This diagram illustrates an example of a playback screen sent from a web server to a user's terminal. [Figure 4] This figure shows an example of the hardware configuration of an AP server, which is an information processing device according to one embodiment of the present invention. [Figure 5] This figure shows an example of a functional configuration implemented on an AP server, which is an information processing device according to one embodiment of the present invention. [Figure 6] This flowchart shows an example of the target data generation process performed by the CPU. [Figure 7] This flowchart shows an example of the entire process executed by the CPU. [Figure 8] This flowchart shows an example of the scene content generation process. [Figure 9]This is a flowchart showing an example of emotion estimation processing. [Modes for carrying out the invention]

[0009] Embodiments of the present invention will be described below with reference to the drawings. Figure 1 is a diagram illustrating the outline of a service (hereinafter referred to as "this service") that can be realized by an information processing device according to one embodiment of the present invention.

[0010] The information processing device 1 targets users using a communication-enabled terminal (hereinafter referred to as "user terminal") 3 and provides this service using status data that indicates the user's status. This service is, for example, a membership-based service for registered members. In Figure 1, status data is acquired by a camera 6, headgear 8, and wristband 9. In addition to these, the user terminal 3 is connected to a display 4, speaker 5, keyboard 7A, and mouse 7B.

[0011] Camera 6 is primarily intended to capture the user's face. The video data output by Camera 6 is state data representing facial expressions. This video data is output to the user terminal 3.

[0012] The headgear 8 is equipped with an electroencephalograph 8A and a microphone (hereinafter referred to as "microphone") 8B. The brainwaves measured by the electroencephalograph 8A and the audio input to the microphone 8B are transmitted to the user terminal 3 as brainwave data and audio data, respectively. The wristband 9 is a device capable of measuring the pulse rate of the user wearing it. The measured pulse rate data, such as the number of pulses per minute, is transmitted to the user terminal 3. Pulse rate data, like electroencephalogram (EEG) data and voice data, is state data. Note that status data is not limited to those described above. For example, video footage of the entire user, the user's body temperature, respiratory rate, etc., can also be used as status data. There are no particular limitations on combinations of status data.

[0013] This state data is used to estimate the emotions of users viewing the content. Specifically, the emotions estimated include mental states such as joy, sadness, tension, unpleasantness, fear, anxiety, and relaxation. This service employs multimodal emotion estimation to achieve more accurate emotion estimation. The emotion estimation results are used to generate the content to be viewed. Content generation is carried out using the expected emotional intonation of the content as one guideline. Therefore, even when emotion estimation results are used in content generation, a decline in quality that would impair the emotional intonation of the generated content is avoided, or at least minimized. Furthermore, the existence of a guideline allows for more reliable creation of presentations, scenes, playback speed adjustments, sound effect changes, or music changes that are more desirable for the user. In this way, this service can more reliably generate content that is appropriate for the user (viewer).

[0014] To enable the generation of content as described above, the information processing device 1 is equipped with the following functional configuration: an emotion estimation unit 1A, a generation condition setting unit 1B, a content generation unit 1C, and a tendency estimation unit (profiling unit) 1D. The emotion estimation unit 1A performs emotion estimation (emotion recognition) using state data received from the user terminal 3. The generation condition setting unit 1B sets the generation conditions for the content to be viewed by the user. The content generation unit 1C generates content according to the set generation conditions. The tendency estimation unit 1D estimates the tendency and degree of sensitivity characteristics of the user from the content, based on the user's emotional changes (affect) or requests. These sensitivity characteristics are estimated for each category of emotion, such as anger, anxiety, tension, and relaxation.

[0015] The emotion estimation unit 1A performs emotion estimation continuously while the user is viewing the content. This allows for monitoring of changes in the user's emotions while they are viewing the content. The setting of generation conditions by the generation condition setting unit 1B is roughly classified into those performed before the start of content generation and those performed at any time during content generation. Hereinafter, the former will be referred to as "initial generation conditions" and the latter as "individual generation conditions" for distinction.

[0016] For ease of understanding, video content including audio such as movies will be specifically described as an example. Note that the content is not limited to video content. It may be only video (video or still image), only audio, text, scent, vibration, etc. For example, for video content with audio, it may be a combination with scent, vibration, etc. For this reason, the devices to be controlled during content playback are not limited to the display 4 and the speaker 5. A scent generator, an actuator, etc., may also be a control target.

[0017] In such video content, the whole can be divided by scenes. Therefore, as shown in FIG. 1, as the original video content, it is also possible to prepare one having the video content divided into a plurality of scenes and one or more partial contents of each scene. Hereinafter, the partial content will be referred to as "scene content".

[0018] The scene content in each scene is prepared in consideration of differences such as the sensitivity of the user. For example, in a stimulating scene, considering individual differences in the emotions evoked in that scene, a plurality of scene contents are prepared. Thereby, users with high sensitivity to stimuli can view scene contents with weaker stimuli, and users with low sensitivity to stimuli can view scene contents with stronger stimuli. In FIG. 1, "Scene 2-1", "Scene 2-2", "Scene 3-1", "Scene 3-2", etc., the number before the hyphen represents the playback order of the scene, and the number after it represents that the content is different even in the same scene.

[0019] The initial generation conditions include the transition path of the scene content to be played. The scene content of each scene is basically played along the set transition path. For users who have never viewed the content before, a predetermined transition path (hereinafter referred to as the "original transition path") is set. However, for users who have viewed the content before, the original transition path is modified as needed, and the modified path is set as the transition path. The modification of the original transition path is performed by referring to the sensitivity tendency characteristics estimated by the tendency estimation unit 1D.

[0020] Multiple scene contents prepared for the same scene are created considering the emotions to be evoked and their degree. Therefore, the transition path of the scene contents, i.e., the playback order, must be desirable for the user. By referring to the estimation results of sensitivity tendency characteristics by the tendency estimation unit 1D, a more appropriate playback order of scene contents can be set as the transition path for the user. Therefore, simply by determining the transition path, it is possible to more reliably generate content that is appropriate for the user. Hereafter, the emotions and their degree that are expected in a scene or scene contents will be collectively referred to as "target emotions" and distinguished from the user's emotions.

[0021] The emotions actually evoked in a user, and to what extent, can vary depending on the user's state at the time. For example, a user who has already viewed stimulating content may experience a lower level of emotion from other stimulating content. The opposite can also be true. For this reason, we have implemented individual generation conditions to enable the creation of more appropriate scene content.

[0022] These individual generation conditions are set considering factors such as the difference between the target emotion of the currently viewed scene content and the emotion actually evoked, the difference between the target emotion of the next viewed scene content and the actual emotion evoked, and the user's sensitivity to the next viewed scene content. As a result, individual generation conditions are generated and set to better match the user's current mental characteristics, ensuring that the user's emotions transition more appropriately. This makes it possible to generate more relevant content for the user, allowing them to enjoy more personalized content tailored to their individual needs.

[0023] Therefore, this service generates and delivers emotionally and adaptively content to users by performing emotion estimation using multimodal state data. As a result, the information processing device 1 operates as a multimodal adaptive emotional content generator (MAECG).

[0024] In this service, information processing device 1 performs multimodal emotion estimation and content generation. However, these functions could also be performed using another information processing device 2 equipped with emotion estimation AI (Artificial Intelligence) 2A and generation AI 2B. This would eliminate the need to equip information processing device 1 with all the functions necessary for providing this service. Including this, various modifications are possible in the provision of this service.

[0025] Hereafter, embodiments of the present invention will be described in detail with further reference to the drawings. Figure 2 illustrates an example of a network system configuration using an information processing device according to one embodiment of the present invention. In this network system, the information processing device 1 is implemented as an AP server installed within the facilities of the service provider 10 that provides this service. Therefore, the AP server is assigned the code "1". This AP server, along with the Web server 11 and the DB (Database) server 12, is connected to the internal network 13. Note that these servers 1, 11, and 12 may utilize cloud services.

[0026] The web server 11 is connected to an external network 15. This network 15 is a complex network that includes the internet. Users wishing to use this service simply need to connect their user terminal 3 to network 15 and access the web server 11. The user terminal 3 is an information processing device equipped with communication and content viewing functions, such as a PC (Personal Computer), smartphone, or tablet PC. The web server 11 responds to requests from user terminal 3 and performs the necessary processing. When it receives a request related to the provision of this service from user terminal 3, it requests the necessary processing from AP server 1 or DB server 12.

[0027] Figure 3 illustrates an example of a playback screen sent from a web server to a user terminal. This playback screen 30 is for when a user plays (views) video content including audio. As shown in Figure 3, the playback screen 30 includes a display area 31 for video playback, a display area 32 for text input, a "play / stop" button 33 for instructing playback or stopping of the content, an "end" button 34 for instructing the end of content playback, a "send" button 35 for instructing the sending of text entered in the display area 32, and a "cancel" button 36 for canceling (deleting) text input.

[0028] The text entered in display area 32 is assumed to represent requests (instructions) regarding content playback. For example, users who have requests such as "I want more exciting content," "I want more variety," or "The tempo is too slow" can enter their requests as text and operate the "Send" button 35. The requests (instructions) entered as text will be reflected in the subsequent content generation.

[0029] DB Server 12 stores various data necessary for providing this service. For example, video content including audio, and other content, are stored and managed in DB Server 12. AP Server 1 then retrieves the necessary data, including content, from DB Server 12 and generates the content to be distributed. Because adaptive content generation is performed, streaming is used for content playback. The transmission of content for streaming is handled by Web Server 11.

[0030] Figure 4 shows an example of the hardware configuration of an AP server, which is an information processing device according to one embodiment of the present invention. This hardware configuration example is just one example and is not particularly limited. For example, only one CPU (Central Processing Unit) 41 and one GPU (Graphics Processing Unit) 44 are shown, but multiple units of each may be installed. As shown in Figure 4, AP Server 1 has a configuration in which a CPU 41, ROM (Read Only Memory) 42, RAM (Random Access Memory) 43, GPU 44, NIC (Network Interface Card) 45, auxiliary storage device 46, media drive 47, and I / FC (Interface Controller) group 48 are connected to bus 49. VRAM (video RAM) 44A is connected to GPU 44.

[0031] The auxiliary storage device 46 is a device capable of permanently storing data, such as a hard disk drive or an SSD (Solid State Drive). The media drive 47 is a device on which the recording medium 47A can be attached and detached. The media 47A is such as a CD (Compact Disc)-ROM, DVD-ROM, DVD-RAM, etc.

[0032] The I / FC group 48 includes various I / FCs that enable communication with various peripheral devices, including the input device 48A and the display device 48B, or with external devices. The input device 48A and the display device 48B are temporarily connected to the I / FC group 48 as needed. The auxiliary storage device 46 stores the OS (Operating System) and various application programs that run on that OS as programs. Among these various application programs is an application program that enables the provision of this service. Hereafter, this application will be referred to as the "content generation app".

[0033] ROM42 is also a device capable of permanently storing data, such as firmware and various other data. The CPU41 reads the firmware stored in ROM42 into RAM43 and executes it. Subsequently, the firmware reads the OS stored in auxiliary storage device 46 into RAM43 and executes it. Various application programs, including some content generation applications, are read into RAM43 by the OS and executed. The GPU44 executes various application programs, including some content generation applications, that are stored in auxiliary storage device 46 and read into VRAM44A.

[0034] The content generation application may be stored on media 47A and distributed, or it may be distributed via network 15. When distributed via network 15, the content generation application should be stored on a recording medium that can be directly or indirectly accessed by the information processing device distributing it. In other words, the storage medium may be directly or indirectly accessible by another information processing device that can communicate with the information processing device distributing it.

[0035] Figure 5 shows an example of a functional configuration implemented on an AP server, which is an information processing device according to one embodiment of the present invention. This example of a functional configuration is mainly implemented by having the CPU 41 and GPU 44 execute different parts of the content generation application, respectively. As shown in Figure 5, the CPU 41 of the AP server 1 is functionally configured to include a transmission / reception processing unit 411, an audio processing unit 412, an image processing unit 413, a biometric data processing unit 414, a request processing unit 415, an audio recognition unit 416, a natural language processing unit 417, a target generation unit 418, an encoding unit group 419, a proofreading unit 420, a generation condition setting unit 421, a profiling unit 422, and an emotion comparison unit 423.

[0036] On the GPU44, the following functional configurations are implemented: a multimodal emotion estimation unit 441, a storyline generation unit 442, a scenario generation unit 443, a video generation unit 444, and an audio generation unit 445. AI is used in all of these; in other words, it is a generative AI. As a result of this functional configuration being realized on the CPU 41 and the GPU 44, the auxiliary storage device 46 has a personal history data storage unit 461, a content management data storage unit 462, a target data storage unit 463, a video data group storage unit 464, an audio data group storage unit 465, and a learning data group storage unit 466 reserved as data storage areas.

[0037] The various data stored in each memory unit 461-466 are actually read into RAM 43 or VRAM 44A for processing. Data transfer between CPU 41 and GPU 44 is also performed via RAM 43. Communication with the web server 11 is performed via NIC 45. These are conveniently ignored in Figure 5. The explanation will proceed accordingly.

[0038] The data stored in each of the storage units 461 to 466 allocated on the auxiliary storage device 46 is, for example, as follows: The personal history data stored in the personal history data storage unit 461 is history data that enables users to play more appropriate content, and is generated and stored each time a user views content. This personal history data includes user identification data such as member ID (IDentification), content ID which is the identification data of the viewed content, viewing date and time, goal data, sentiment difference, etc.

[0039] Target data is data generated before viewing content begins. Target data represents the transition patterns of emotions (target emotions) that should be evoked in the user during content viewing. For example, in movie content, the transition patterns represent the emotions to be evoked for each scene. In content containing multiple sub-contents, such as movie content, the transition path specifying the sub-content to be delivered is also included in the target data. Member ID and content ID are included in the target data regardless of the content type. The target data storage unit 463 is a storage area reserved for saving such target data. This target data corresponds to the initial generation conditions described above.

[0040] The emotional difference that constitutes personal history data represents the difference between the emotion represented by the target data and the actual emotion of the user. For example, in movie content, it represents the difference between the target emotion intended to be evoked in a scene and the user's estimated emotion in that scene. This emotional difference allows us to estimate the user's sensitivity tendencies to various types of video, such as scary, relaxing, or sad videos. These sensitivity tendencies enable us to more reliably generate content that is more appropriate for the user.

[0041] The content management data storage unit 462 stores content management data for each existing piece of content, which includes data representing the transition pattern of the emotions (target emotions) anticipated during the production of that content. In addition to the transition pattern, the content management data also includes the content ID, the emotional inflection pattern within each scene, and candidate transition destinations between scene contents based on the target emotion (Figure 1). The emotional inflection pattern within each scene can be used to extract features to be used for emotion estimation in each scene.

[0042] The video data group storage unit 464 stores video data and video management data. Each video data is the source video data for editing. The video management data is data prepared for each video data and includes data such as the identifier assigned to the corresponding video data, the playback location, the migration destination candidate (see Figure 1), and the storage location of the video data. This identifier is data that represents two numbers separated by a hyphen following each scene shown in Figure 1 (for example, "2-1", "2-2", "3-1", etc.).

[0043] The sound data group storage unit 465 stores sound data and sound management data. Each sound data is the source data for generating sounds such as character voices, various sounds from nature, sounds emitted by various machines, musical pieces, etc. The sound management data is data prepared for each sound data and includes data such as an identifier assigned to the corresponding sound data, sound type, target object, and storage location of the sound data.

[0044] The learning data storage unit 466 stores learning data and learning management data. Learning data is, for example, video data for the user to view or data for displaying problems. Learning management data is data prepared for each set of learning data and includes data such as the genre, difficulty level, scope, and storage location of the learning data. Data other than the target data stored in the target data storage unit 463 is managed, for example, by the DB server 12, and retrieved and stored from the DB server 12 as needed. In this way, the AP server 1 works in cooperation with the DB server 12 to provide this service.

[0045] As described above, various requests and data sent from the user terminal 3 are received by the Web server 11, and requests and data to be processed are sent from the Web server 11 to the AP server 1. The transmission / reception processing unit 411 performs processing for sending and receiving various data, including requests, with the Web server 11. Various status data received from the user terminal 3 via the Web server 11 are passed to the audio processing unit 412, the video processing unit 413, or the biometric data processing unit 414, depending on their type. Requests are passed to the request processing unit 415. Text entered in the display area 32 is received as text data by the transmission / reception processing unit 411 and passed to the natural language processing unit 417.

[0046] The voice processing unit 412 processes the voice data received from the user terminal 3, extracts the user's voice, and further extracts physical features such as the speed, intonation, and tone of the voice. The extracted features, i.e., the generated feature data, are passed to the encoding unit group 419, and the voice data is passed to the voice recognition unit 416.

[0047] The speech recognition unit 416 analyzes the received audio data and extracts the content of what the user is saying as text data. The extracted text data is passed from the speech recognition unit 416 to the natural language processing unit 417. The natural language processing unit 417 also receives text data from the transmission / reception processing unit 411.

[0048] The natural language processing unit 417 performs natural language processing on the received text data to determine the content of the text spoken by the user or entered in the display area 32, and performs sentiment analysis using the determined content to identify the user's sentiment class. The identified sentiment class is passed to the encoding unit group 419, and the determined content is passed to the request processing unit 415 as a request representing the user's wishes.

[0049] The video processing unit 413 recognizes, for example, the user's facial expression, pupil size, and eye movements represented by the video data, and extracts feature quantities for estimating emotions from the recognition results and passes them to the encoding unit group 419. The biometric data processing unit 414 extracts feature quantities for estimating emotions from the state of brain waves, pulse rate, etc., and their changes, as represented by the biometric data, and passes them to the encoding unit group 419. The audio data, text data, video data, and biometric data mentioned above are all state data. The processing of these data is based on well-known techniques. Therefore, a more detailed explanation will be omitted. Furthermore, the types of biometric data will not be discussed further.

[0050] The request processing unit 415 controls each part for processing the request. The content to be distributed can be specified directly by the user, for example. Alternatively, users can specify conditions and select content from those filtered by those conditions. Furthermore, users can automatically specify content by specifying conditions for generating original content (content without a source).

[0051] The conditions that can be specified include, for example, genre (action film, science fiction film, horror film, etc.), setting (sea, space, etc.), production method (live-action, animation, CG, etc.), target age, keywords, etc. For learning content, these include, for example, field, purpose (learning, test, etc.), target age, keywords, etc. These are also common to original content. For original content, it may also be possible to specify the pattern of emotional transitions to be evoked in the user, the degree of expression (range of emotional inflection), etc. These are all just examples, and there are no particular limitations on the items that can be specified as conditions.

[0052] The process of allowing the user to determine which content to play is performed by the Web server 11. Therefore, once the content to be played is determined, a command containing data specifying that content is sent as a request from the Web server 11 to the AP server 1. To avoid confusion, unless otherwise specified, this explanation assumes that the user directly specifies the content to be played.

[0053] When the request processing unit 415 receives a request specifying the content to be played, it instructs the target generation unit 418 to generate target data for generating that content. In response to this instruction, the target generation unit 418 generates the target data and stores the generated target data in the target data storage unit 463, which is secured on the auxiliary storage device 46.

[0054] The text data sent as a request, that is, the text data entered in the display area 32, is analyzed for its semantic content, and the instruction content is identified. Similarly, the instruction content is identified in the text data extracted from the audio data. The identified instruction content is passed from the request processing unit 415 to the profiling unit 422.

[0055] The profiling unit 422 analyzes the instructions it receives and either estimates the user's sensitivity traits or updates existing estimation results. By reflecting the requests notified by the user's request in the estimation of sensitivity traits, it is possible to estimate more appropriate sensitivity traits. This also enables the generation of more appropriate content, including editing.

[0056] As described above, the encoding unit 419 receives feature quantities (feature data) from the audio processing unit 412, the natural language processing unit 417, the video processing unit 413, and the biometric data processing unit 414 as needed. The encoding unit 419 selects the feature quantities to be used, for example, based on the type of state data, and encodes the selected feature quantities in a format suitable for input to the multimodal emotion estimation unit 441. The reason for selecting the feature quantities to be used is that emotion estimation is performed once per scene. It is also possible to enable multiple emotion estimations per scene and extract the most effective emotion estimation result from the multiple emotion estimation results.

[0057] In this service, the generation of content including audio (e.g., scene content) is performed in the following order, based on the source content: plot generation → scenario generation → video generation reflecting the generated scenario (video editing) → audio generation to match the generated video. To ensure consistency with previously viewed scene content, this service checks whether the generated plot is consistent, and proceeds to scenario generation only if consistency is confirmed. Proofreading unit 420 checks whether this consistency is maintained. Specifically, proofreading unit 420 checks the consistency between scenes, the consistency with the scenario in the previous scene, etc. Consistency between scenarios can also be checked by generating document vectors, for example.

[0058] The generation condition setting unit 421 generates and sets generation conditions for each scene. These generation conditions correspond to the individual generation conditions described above. Therefore, to avoid confusion, the generation conditions set by the generation condition setting unit 421 will be referred to as "individual generation conditions" from now on. The generation of individual generation conditions references target data and each sensitivity tendency characteristic. If the user provides instructions via voice or text, those instructions are also referenced. For content generation, these instructions apply only to the scene immediately following the scene the user is currently viewing, to all subsequent scenes, or to the relevant scene. In some cases, they may be ignored. For example, instructions requesting that subsequent scenes be made more tense when there is nothing tense in those scenes will be ignored.

[0059] The generation of individual generation conditions is performed, for example, by manipulating each sensitivity tendency characteristic and user instructions in relation to the target emotion of the next scene represented by the target data. This manipulation could involve representing each sensitivity tendency characteristic numerically and using a calculation formula based on those values ​​to calculate the numerical value of the emotion to be evoked. The method of manipulation is not particularly limited.

[0060] The individual generation conditions generated and set in this manner are used as input to the script generation unit 442. In scenario generation, the script generated by the script generation unit 442 is used as a generation condition, for example, as a prompt input. In video generation, the generated scenario is used as a generation condition, for example, as a prompt input. In audio generation, the generated video is used as a generation condition. Therefore, the video and audio are synchronized. In video content as shown in Figure 1, the user may be prompted to input identifiers for the partial content (scene content) to be generated.

[0061] The emotions recognized by the multimodal emotion estimation unit 441 are processed by the emotion comparison unit 423. The emotion comparison unit 423 compares the recognized emotions with the target emotions set in the individual generation conditions and passes the comparison result to the profiling unit 422. This comparison result is the emotion difference, and is stored in the personal history data storage unit 461 as data that constitutes the personal history data. The target data that constitutes the personal history data is stored in the personal history data storage unit 461 by, for example, the target generation unit 418. The profiling unit 422 refers to the emotion difference passed from the emotion comparison unit 423 and updates the sensitivity tendency characteristics as necessary.

[0062] The multimodal emotion estimation unit 441, implemented on the GPU 44, takes various feature quantities encoded by the encoding unit group 419 as input and performs emotion estimation, as described above. The storyline generation unit 442, scenario generation unit 443, video generation unit 444, and audio generation unit 445 all take individual generation conditions set by the generation condition setting unit 421 as input and perform generation processing. All units 441 to 445 function according to instructions from the CPU 41.

[0063] Figure 6 is a flowchart showing an example of the target data generation process executed by the CPU. The target generation unit 418 is realized when the CPU 41 executes this generation process. Next, with reference to Figure 6, this generation process will be explained in detail.

[0064] First, in step S101, the CPU 41 extracts the content management data specified in the request from the content management data storage unit 462. In the next step, S102, it is determined whether the user viewing the content has a past viewing history, that is, whether personal history data with the user's member ID exists in the personal history data storage unit 461. If such personal history data exists in the personal history data storage unit 461, the determination in step S102 is YES and the process proceeds to step S103. Otherwise, that is, if the user has never viewed the content before, the determination in step S102 is NO and the process proceeds to step S105.

[0065] In step S103, the CPU 41 generates and determines the emotional transition patterns that should be evoked in the user while viewing the content. These transition patterns are generated by manipulating the transition patterns represented by the content management data corresponding to the content to be delivered, as described above. This manipulation of transition patterns refers to the user's emotional tendencies and is intended to personalize the content to match the user's emotional tendencies.

[0066] In step S104, following step S103, the CPU 41 refers to the content management data and sets the transition path between each scene content according to the transition pattern. This setting of the transition path determines which scene content should be played in each scene, for example, Scene 1 → Scene 2-2 → Scene 3-2. After this, the target data generation process ends.

[0067] The target data generation process described above assumes content that includes multiple sub-contents, such as movie content, and for which source content exists. If source content does not exist, neither content management data extraction nor migration path determination will be performed. Based on the user's specified requests and past viewing history of the content, the content to be distributed is determined, and emotional transition patterns are generated according to that content. The generated emotional transition patterns are then manipulated. In this case, content generation may, for example, be performed by generating a scenario and then generating a video based on prompt input from the generated scenario. Since the prerequisites for content generation are different, the modeling and training content for the generation AI that needs to be prepared will also be different.

[0068] Figure 7 is a flowchart showing an example of the overall processing performed by the CPU. This overall processing includes the processing performed by the CPU 41 for content generation after content generation has started. By executing this overall processing, the following are realized on the CPU 41: audio processing unit 412, video processing unit 413, biometric data processing unit 414, speech recognition unit 416, natural language processing unit 417, target generation unit 418, encoding unit group 419, proofreading unit 420, generation condition setting unit 421, profiling unit 422, and sentiment comparison unit 423. Next, we will refer to Figure 7 and explain this overall processing in detail. Note that content generation is started when the user clicks the "Play / Stop" button 33 on the playback screen 30.

[0069] First, in step S201, the CPU 41 determines whether or not it has received status data. If it has received any of the following: audio data, biometric data, or video data, the determination in step S201 is YES and the process proceeds to step S202. If it has not received any of the following, the determination in step S201 is NO and the process proceeds to step S208. In step S202, the CPU 41 performs processing according to the type of status data received. This processing enables the audio processing unit 412, the video processing unit 413, and the biometric data processing unit 414. The processing itself is based on well-known technology.

[0070] In the following step S203, the CPU 41 determines whether or not there is audio data among the received state data. If the user makes any sound, audio data is sent from the user terminal. Therefore, the determination in step S203 is YES and the process proceeds to step S204. Otherwise, the determination in step S203 is NO and the process returns to step S201.

[0071] In step S204, the CPU 41 performs speech recognition processing on the received audio data and converts it into text data. In the following step S205, the CPU 41 performs natural language processing on the previously obtained text data to identify the user's instructions. After that, the process moves to step S206, where the CPU 41 determines whether or not there are instructions that should be valid for content generation. If such instructions and content that should be recognized are identified through natural language processing, the determination in step S206 is YES and the process moves to step S207. Otherwise, the determination in step S206 is NO and the process returns to step S201.

[0072] When a user gives instructions by voice, the processing loop formed in steps S201-S206 is repeatedly executed until the content of the instructions is identified. When instructions are given by sending text data (Figure 3), the content of the instructions is identified in a single natural language processing step.

[0073] In step S207, the CPU 41 saves the identified instruction (instruction content). The saved instruction will be reflected in the subsequent generation of partial content as needed. After saving, the process returns to step S201. If the judgment in step S201 is NO, the process proceeds to step S208, where the CPU 41 determines whether or not the timing for generating partial content has arrived. If the generation timing has arrived, the judgment in step S208 is YES and the process proceeds to step S209. If the timing has not arrived, the judgment in step S208 is NO and the process proceeds to step S210.

[0074] As described above, this service allows users to view content via streaming. Therefore, the individual content segments that make up the content are generated sequentially. The generation timing is automatically determined based on factors such as the start time of the currently viewed segment and its playback duration.

[0075] In step S209, the CPU 41 executes a scene content generation process to generate the next scene content to be distributed. After the execution of this generation process, the process returns to step S201.

[0076] Figure 8 is a flowchart showing an example of the scene content generation process. We will now refer to Figure 8 to explain the generation process in detail. First, in step S301, the CPU 41 generates and sets individual generation conditions. These individual generation conditions include, in addition to the target emotion, an identifier assigned to, for example, the scene content to be edited.

[0077] The target emotion specified by the emotion transition pattern represented by the target data is, as described above, an operation that takes into account sensitivity tendency characteristics when applying the emotion transition pattern represented by the content management data. Therefore, in the generation of individual generation conditions here, the target emotion is operated on by considering, for example, the difference between the target emotion set in the currently viewed scene content and the emotion actually recognized by the user, and / or the user's instructions. This difference is referenced to optimize the emotional intonation expected in the content.

[0078] For example, in a case where the currently playing scene content is intended to evoke fear in the user, and the next scene content is intended to evoke relaxation, it is desirable that the user actually experiences the expected difference in emotion (emotional inflection). However, if the user does not feel as much fear as expected during the currently playing scene content, the expected difference in emotion is unlikely to occur. Therefore, in this service, in order to more reliably create such a difference in emotion, the recognized emotions of the user are used to manipulate, or determine, the target emotion.

[0079] In step S302, following step S301, the CPU 41 passes the set individual generation conditions to the GPU 44 and instructs it to generate the storyline. This instruction causes the storyline generation unit 442 on the GPU 44 to function. In the next step, S303, CPU41 performs a proofreading process to check whether there are any discrepancies between the script generated by GPU44 and the script already generated. After that, the process moves to step S304, where CPU41 determines whether there are any discrepancies. If there are discrepancies, the determination in step S304 is YES, and the process moves to step S305. If there are no discrepancies, the determination in step S304 is NO, and the process moves to step S306.

[0080] In this context, "defects" primarily refer to inconsistencies in content. For example, a scene might take place in a location that doesn't fit the narrative, feature characters that don't fit the narrative, or be missing characters that should be present.

[0081] In step S305, CPU41 resets the individual generation conditions. This reset is done, for example, by changing the scene content identifier. For example, it changes the identifier to the scene content that is considered to better match the target emotion among the destinations of the currently viewed scene content. After resetting the individual generation conditions, including this, the process returns to step S302, and GPU44 is instructed to generate a storyline with the reset individual generation conditions.

[0082] In step S306, CPU 41 instructs GPU 44 to generate a scenario using the generated storyline as prompt input. In this instruction, GPU 44 is given the storyline along with a scene content identifier as a generation condition. CPU 41 then instructs GPU 44 to generate a video using the generated scenario and generation conditions including the scene content identifier (step S307), and further instructs GPU 44 to generate audio using the generated video as a generation condition (step S308). Finally, CPU 41 sends the generated data, i.e., the video content with audio, as scene content to the web server 11 (step S309). After this transmission, the scene content generation process ends.

[0083] In accordance with the instructions in steps S306-S308, the scenario generation unit 443, video generation unit 444, and audio generation unit 445 on the GPU 44 will function in sequence. By generating scene content in this way, it becomes possible to personalize the sheet content in a way that reflects the user's emotions while taking guidelines into consideration. Even when viewing the same content, users can enjoy content that changes depending on their mental state at the time.

[0084] Let's return to the explanation of Figure 7. In step S210, the CPU 41 determines whether or not text data has been received. If text data, that is, text data entered in the display area 32 located on the playback screen 30, is received, the determination in step S210 is YES, and the process proceeds to step S205, where natural language processing is performed. If no text data has been received, the determination in step S210 is NO, and the process proceeds to step S211.

[0085] In step S211, the CPU 41 determines whether the timing for emotion estimation (indicated as "estimation timing" in Figure 7) has arrived. This estimation timing is determined, for example, from the target emotion of the scene content being viewed and the emotional intonation pattern in the corresponding scene. In other words, the estimation timing is determined by identifying the timing at which the target emotion is thought to be evoked, for example, from the intonation pattern, and then determining the time at which the identified timing will arrive. If the estimation timing determined in this way has arrived, the determination in step S211 is YES and the process proceeds to step S212. If the estimation timing has not arrived, the determination in step S211 is NO and the process proceeds to step S213.

[0086] In step S212, the CPU 41 performs emotion estimation processing using features extracted from various state data. After this emotion estimation processing is completed, the process returns to step S201.

[0087] Figure 9 is a flowchart illustrating an example of emotion estimation processing. We will now refer to Figure 9 to explain the recognition process in detail. First, in step S401, the CPU 41 selects one of the features obtained in the current scene for each type of state data. When determining the estimation timing as described above, it is conceivable to select the most recently obtained feature for all state data. In the next step, S402, the CPU 41 encodes each of the selected state quantities and sends the encoding results to the GPU 44 to instruct it to perform emotion estimation. This instruction causes the emotion estimation unit 441 on the GPU 44 to function.

[0088] In the following step S403, the CPU41 analyzes the emotion recognition results returned from the GPU44 and identifies the differences from the target emotion, etc. In the subsequent step S404, the CPU 41 determines whether or not a difference exceeding the acceptable range has been identified. If a difference exceeding the acceptable range is identified, that is, if an emotion different from the target emotion has been evoked in the user, the determination in step S404 is YES and the process proceeds to step S405. If no such difference is identified, the determination in step S404 is NO and the process proceeds to step S406.

[0089] In step S405, the CPU 41 evaluates the identified difference and, if necessary, updates the sensitivity tendency characteristics by reflecting the evaluation results in the sensitivity tendency characteristics. If the sensitivity tendency characteristics for each emotion classification are represented by a predetermined range of values, for example, between 0 and 100, updating the sensitivity tendency characteristics here corresponds to manipulating the value of one of the sensitivity tendency characteristics. The manipulation of the value is not particularly limited, but for example, it may be done by multiplying the value representing the difference by a coefficient and adding the result of that multiplication to the current value. Alternatively, the value after the manipulation may be calculated by referring to the difference obtained when the user viewed other content in the past and obtaining the average value or a weighted average value, etc.

[0090] In the following step S406, the CPU 41 saves the identified difference. This difference is ultimately stored in the personal history data storage unit 461 as an emotion difference that constitutes the personal history data. After saving the difference, the emotion estimation process ends.

[0091] Let's return to the explanation of Figure 7. In step S213, the CPU 41 determines whether or not content playback has finished. The user can instruct the system to end the content being played by clicking the "End" button 34 located on the playback screen 30. This instruction is notified from the Web server 11 to the AP server 1, and the AP server 1 will stop generating scene content even if there is other scene content to be generated. The AP server 1 will also stop generating scene content after generating the last scene content and there is no more scene content to be generated. For this reason, if there is no more scene content to be generated, the determination in step S213 will be YES and the system will proceed to step S214. Otherwise, that is, if there is still a possibility of generating new scene content, the determination in step S213 will be NO and the system will return to step S201. As a result, as long as there is a possibility of generating scene content, the CPU 41 will maintain a state in which it can respond to various data, including requests sent from the user terminal 3.

[0092] In step S214, the CPU 41 generates the personal history data as data indicating the results of the user viewing the content, and stores it in the personal history data storage unit 461. After storing the personal history data, the entire process ends.

[0093] The above explanation primarily assumes a scenario where a user views content with one or more scene contents prepared for each scene, as shown in Figure 1, and content generation is performed by editing the scene contents. However, content generation can also be performed without using source content, as described above. In other words, as described above, the content content that satisfies the conditions specified by the user can be identified, the target emotion (or target emotion transition pattern) in that content can be determined, and new content can be generated.

[0094] Movies and other video content are designed to convey emotional nuances. However, some content does not anticipate such emotional fluctuations. Learning data (learning content) intended for viewing or presented as a problem typically does not anticipate emotional nuances.

[0095] For users viewing such learning content, supplementary content that acts on the visual, auditory, or olfactory senses may be generated and played to bring the estimated emotion closer to the target emotion, or to keep it from deviating from the target emotion. For example, for an excited user, calming visual effects, music, sound playback, or scent generation may be implemented. If learning content with different learning content is played consecutively, the estimated emotion may be used to select the learning content that is more desirable for the user, and playback may be performed by replacing the previous content with the selected learning content or adding the selected learning content.

[0096] This service is a membership-based service for registered members, and is intended to be provided to a large number of members. However, this service may also be provided in a form different from a membership-based service. For example, an application program for individual users may be developed, and the developed application program may be recorded on a recording medium or sold via the network 15. Such an application program may utilize the functions of other information processing devices 2, as shown in Figure 1. Furthermore, at least one of the multimodal emotion estimation unit 441, storyline generation unit 442, scenario generation unit 443, video generation unit 444, and voice generation unit 445 may be made customizable and trainable by the purchaser.

[0097] Content generation is performed using the generation AI of the story generation unit 442, scenario generation unit 443, video generation unit 444, and audio generation unit 445, but the generation AI used for content generation, the generation procedure, etc., are not particularly limited. Content may also be generated without using the generation AI. Including such modifications, the present invention can be modified in various ways. [Explanation of Symbols]

[0098] 1 Information processing device (AP server), 1A Emotion estimation unit, 1B Generation condition setting unit, 1C Content generation unit, 3 User terminal, 6 Camera, 8 Headgear, 8A Electroencephalograph, 8B Microphone, 1D Trend estimation unit, 10 Service provider, 11 Web server, 12 DB server, 13, 15 Network, 30 Playback screen.

Claims

1. A data acquisition unit that acquires status data representing the user's state, An emotion estimation unit estimates the user's emotions using the state data acquired by the data acquisition unit, When allowing the user to view content in which one or more partial content items to be generated exist, a condition generation unit generates conditions for generating the partial content based on the emotion estimation result by the emotion estimation unit and the target emotion to be targeted for the user. A content generation unit generates the partial content based on the generation conditions generated by the condition generation unit, An information processing device equipped with the following features.

2. The system further comprises a tendency estimation unit that estimates the user's sensitivity tendency from the difference between the emotion estimated by the emotion estimation unit from the user when viewing the partial content and the target emotion used to generate the generation conditions for the partial content, The condition generation unit reflects the trend estimated by the trend estimation unit in the generation of the generation conditions, and causes the content generation unit to generate the personalized partial content for the user. The information processing apparatus according to claim 1.

3. If the content includes multiple sub-contents, the condition generation unit generates, as generation conditions, an initial generation condition that specifies the target emotion for each sub-content included in the content before viewing of the content begins, and an individual generation condition that is valid for one of the sub-contents while viewing the content. The information processing apparatus according to claim 2.

4. The data acquisition unit is capable of acquiring multiple types of state data. The emotion estimation unit estimates the emotion using the various types of state data. The information processing apparatus according to claim 1.

5. In an information processing device, Obtain state data that represents the user's state, Using the acquired state data, the user's emotions are estimated. When allowing the user to view content in which one or more partial content items to be generated exist, the generation conditions for the partial content are generated based on the emotion estimation results and the target emotion to be targeted for the user. Based on the generated generation conditions, the partial content is generated. A program that executes a process.

Citation Information

Patent Citations

  • Empathic computing system and method for improved human interaction with digital content experiences

    JP2022505836A