Automatic connotation for audio and visual content using IoT sensors

The system uses IoT devices to analyze user emotions and suggest content modifications, addressing the lack of real-time emotional response capture in audiovisual content systems, enhancing user experience and engagement.

JP2025537458APending Publication Date: 2025-11-18INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025517324
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-28
Filing Date
2023-09-18
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing audiovisual content systems lack the ability to capture and respond to users' real-time emotional responses to enhance user experience by modifying content frames to better match intended emotional evocations.

Method used

A system using IoT devices to capture user emotions through sensor data, analyze them using emotion vector analysis and supervised machine learning, and generate suggestions for content creators to modify frames to align with intended emotional responses.

Benefits of technology

Enhances user experience by dynamically adjusting audiovisual content to better match the emotions intended by creators, improving user engagement and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025537458000001_ABST
    Figure 2025537458000001_ABST
Patent Text Reader

Abstract

In a method for enhancing a user's experience listening to and / or watching audiovisual content by modifying future audio and / or video frames of the audiovisual content, a processor captures a set of sensor data from an IoT device worn by a first user. The processor analyzes the set of sensor data and generates one or more connotations by transforming the emotions using emotion vector analysis techniques and supervised machine learning techniques. The processor scores the one or more connotations based on a similarity between an emotion expressed by the first user and an emotion expected to be evoked by a second user. The processor determines whether the score of the one or more connotations exceeds a preconfigured threshold level. In response to determining that the score does not exceed the preconfigured threshold level, the processor generates a suggestion to a producer of the audiovisual content.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Background of the Invention The present invention relates generally to the field of data processing, and more particularly to automatic connotation for audio and visual content using IoT sensors.

[0002] Audiovisual content consists of and / or has any combination of audiovisual components intended to inform, educate, or entertain, regardless of their duration, original intended use, or method of distribution. Examples of audiovisual content include films, television series, online streaming videos, video games, music albums, music songs, podcasts, webinars, slides, lecture notes, works of art, infographics, and photographs and images.

[0003] The Internet of Things (IoT) is the interconnection of physical devices (also called "connected devices" and "smart devices"), vehicles, buildings, and other objects that incorporate electronics, software, sensors, actuators, and network connectivity that enable them to collect and exchange data. The IoT enables objects to be remotely sensed and / or controlled across existing network infrastructures, creating opportunities for more direct integration of the physical world into computer-based systems, resulting in increased efficiency, accuracy, and economic benefits in addition to reduced human intervention. Each "thing" is uniquely identifiable through its embedded computing system, yet can interoperate within the existing Internet infrastructure. Summary of the Invention

[0004] An aspect of an embodiment of the present invention discloses a method, computer program product, and computer system for enhancing a user's experience listening to and / or watching audiovisual content by modifying future audio and / or video frames of the audiovisual content. In response to a first user expressing an emotion toward the audiovisual content, a processor captures a set of sensor data from an IoT device worn by the first user. The processor analyzes the set of sensor data and generates one or more connotations by transforming the emotion using emotion vector analysis techniques and supervised machine learning techniques. The processor uses an analysis process to score the one or more connotations based on a similarity between the emotion expressed by the first user and an emotion expected to be evoked by a creator of the audiovisual content. The processor determines whether the score of the one or more connotations exceeds a preconfigured threshold level. In response to determining that the score does not exceed the preconfigured threshold level, the processor generates a suggestion for the creator of the audiovisual content.

[0005] In some aspects of an embodiment of the present invention, the first user is a viewer of the audiovisual content, and the audiovisual content includes at least one of a film, a television series, a commercial, an online streaming video, a video game, a music album, a music song, a podcast, a webinar, slides, lecture notes, a work of art, an infographic, a photograph, and an image.

[0006] In some aspects of an embodiment of the present invention, the set of sensor data from the IoT device worn by the first user includes at least one of a heart rate of the first user, a pulse rate of the first user, a respiratory rate of the first user, changes in the nervous system of the first user, a set of neurological data of the first user, and movements performed by the first user.

[0007] In some aspects of an embodiment of the present invention, the processor identifies a first set of video frames of the audiovisual content that the first user was watching when the first user expressed an emotion.

[0008] In some aspects of one embodiment of the present invention, the first set of video frames spans a time span, the time span starting when the first user begins to express the emotion and ending when the first user stops expressing the emotion.

[0009] In some aspects of an embodiment of the present invention, a processor compares the set of sensor data with a set of historical data stored in a database to identify the emotions expressed by one or more previous users, and the processor classifies the set of sensor data as an emotion using an emotion learning model based on the results of the comparison.

[0010] In some aspects of an embodiment of the present invention, the processor compares the score assigned to the one or more connotations with a historical data model, the historical data model including a mapping of one or more emotions previously expressed by one or more previous viewers of the audiovisual content.

[0011] In some aspects of one embodiment of the present invention, the suggestions include a set of feedback regarding how the second set of video frames should be modified to more closely match the emotions expected to be evoked by the creator of the audiovisual content, the second set of video frames spanning an upcoming time span.

[0012] In some aspects of an embodiment of the present invention, after generating the suggestions to the producer of the audiovisual content, the processor enables the producer of the audiovisual content to modify the second set of video frames to more closely match the emotions expected to be evoked by the producer of the audiovisual content.

[0013] These and other features and advantages of this invention will be described, or will become apparent to those skilled in the art, in view of the following detailed description of illustrative embodiments of the invention. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a functional block diagram illustrating a distributed data processing environment, according to one embodiment of the present invention.

[0015] [Figure 2] 2 is a flowchart illustrating the operational steps of a content composition program on a server in the distributed data processing environment of FIG. 1 according to one embodiment of the present invention.

[0016] [Figure 3] 2 depicts a block diagram of components of a computing environment representative of the distributed data processing environment of FIG. 1 in accordance with one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0017] Embodiments of the present invention recognize that emotions refer to mental states and / or emotional states that arise naturally rather than through conscious effort. Emotions arise naturally due to neurophysiological changes associated with thoughts, feelings, behavioral responses, and degrees of satisfaction or dissatisfaction. Emotions often involve physical and physiological changes associated with human organs and tissues, such as the brain, heart, skin, blood flow, muscles, facial expressions, and voice.

[0018] Embodiments of the present invention recognize that in response to listening to and / or viewing audiovisual content (hereinafter "audiovisual content"), users may provide feedback through reactions, comments, shares, impressions, and clicks. Audiovisual content consists of and / or has any combination of audio and visual components intended to inform, educate, or entertain, regardless of its duration, original intended use, or method of distribution. Examples of audiovisual content include films, television series, online streaming videos, video games, music albums, music songs, podcasts, webinars, slides, lecture notes, works of art, infographics, and photographs and images. Embodiments of the present invention recognize that a user's mental state and / or emotional state may be determined from the feedback provided by the user. However, currently, the feedback provided by the user is a summary of the user's impressions of the audiovisual content.

[0019] Embodiments of the present invention recognize that audiovisual content may include multiple parts, each part evoking a different psychological and / or emotional state in a user. Embodiments of the present invention further recognize that emotions can be identified by emotion recognition methods. Emotion recognition methods can be classified into two main categories: one category uses human physical signals, such as facial expressions, speech, gestures, and posture; and the other category uses internal signals, such as physiological signs, including, but not limited to, electroencephalogram (EEG), body temperature, electrocardiogram (ECG), electromyogram (EMG), galvanic skin response (GSE), and respiratory rate.

[0020] Accordingly, embodiments of the present invention recognize a need for a system and method for capturing a user's emotional expression towards audiovisual content in real time, translating the user's emotions into suggested responses, and generating a report with one or more suggestions to producers of the audiovisual content to improve future versions of the audiovisual content.

[0021]

[0006] Embodiments of the present invention provide systems and methods for enhancing a user's experience while listening to and / or watching audiovisual content by modifying future audio and / or video frames of the audiovisual content based on one or more emotions expressed by the user. In response to a first user expressing an emotion toward the audiovisual content, embodiments of the present invention capture a set of sensor data from an IoT device worn by the first user. Embodiments of the present invention connect a feedback capture system to the IoT device worn by the first user through a cloud service. Embodiments of the present invention classify the set of captured sensor data as an emotion using an emotion learning model. Embodiments of the present invention connotate the emotion using an emotion vector analysis technique and a supervised machine learning model for frame-by-frame connotation. Embodiments of the present invention score the connotation using an analysis process based on the similarity between the emotion expressed by the first user and the emotion predicted to be evoked by a producer of the audiovisual content. Embodiments of the present invention generate a qualitative report with one or more suggestions that a producer can make to a second set of video frames to more closely match the emotion predicted to be evoked by the producer. Embodiments of the present invention may intelligently generate the second set of video frames with the producer's consent, or may allow the producer to manually generate the second set of video frames.

[0022] Implementations of embodiments of the present invention may take a variety of forms, and details of exemplary implementations are now discussed with reference to the figures.

[0023] 1 is a block diagram illustrating a distributed data processing environment, generally designated 100, according to one embodiment of the present invention. In the illustrated embodiment, distributed data processing environment 100 includes a server 120 and user computing devices 130 interconnected via a network 110. 1-N 1. Distributed data processing environment 100 may include additional servers, computers, computing devices, and other devices not shown. As used herein, the term "distributed" describes a computer system that includes multiple, physically separate devices that operate together as a single computer system. FIG. 1 is intended only to illustrate one embodiment of the invention and does not imply any limitation with regard to the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be made by one of ordinary skill in the art without departing from the scope of the invention as defined in the claims.

[0024] Network 110 operates as a computing network that may be, for example, a telecommunications network, a local area network (LAN), a wide area network (WAN) such as the Internet, or a combination of the three, and may include wired, wireless, or fiber optic connections. Network 110 may include one or more wired and / or wireless networks capable of receiving and transmitting data, voice, and / or video signals, including multimedia signals containing data, voice, and video information. Generally, network 110 may include servers 120, user computing devices 130, and other devices within distributed data processing environment 100. 1-N , and other computing devices (not shown).

[0025] The server 120 operates to execute the content composition program 122 and transmit and / or store data in the database 124. In one embodiment, the server 120 transmits data from the database 124 to the user computing device 130. 1-N In one embodiment, the server 120 may transmit data to the user computing device 130 in the database 124. 1-N In one or more embodiments, the server 120 is capable of receiving, transmitting, and processing data and transmitting the data to the user computing device 130 via the network 110. 1-N Server 120 may be a standalone computing device, an administrative server, a web server, a mobile computing device, or any other electronic device or computing system capable of communicating with user computing devices 130 via network 110. In one or more embodiments, server 120 may be a computing system utilizing clustered computers and components (e.g., database server computers, application server computers, etc.) that operate as a single pool of seamless resources when accessed within distributed data processing environment 100, such as within a cloud computing environment. In one or more embodiments, server 120 may communicate with user computing devices 130 via network 110 within distributed data processing environment 100. 1-N and other computing devices (not shown). Server 120 may be a laptop computer, tablet computer, netbook computer, personal computer, desktop computer, personal digital assistant, smartphone, or any programmable electronic device capable of communicating with other computing devices (not shown). Server 120 may include internal and external hardware components, as shown and described in more detail in FIG.

[0026] The content composition program 122 operates to enhance the user's experience while listening to and / or watching the audiovisual content by modifying future audio and / or video frames of the audiovisual content based on one or more emotions expressed by the user. In the illustrated embodiment, the content composition program 122 is a standalone program. In another embodiment, the content composition program 122 may be integrated into another software product, such as audio or video editing software. In the illustrated embodiment, the content composition program 122 resides on the server 120. In another embodiment, the content composition program 122 resides on the user computing device 130. 1-N Or it may reside on another computing device (not shown), provided that the content composition program 122 has access to the network 110 .

[0027] In one embodiment, the user computing device 130 1-N A user of the server 120 registers with the content composition program 122. For example, the user completes a registration process (e.g., user verification), provides information to create a user profile, and connects to an identified computing device (e.g., user computing device 130). 1-N) by server 120 (e.g., via content composition program 122). Relevant data includes, but is not limited to, personal information or data provided by the user or inadvertently provided by the user's devices without the user's knowledge; tagged and / or recorded user location information (e.g., to infer location or presence context (i.e., time, place, and usage)); time-stamped temporal information (e.g., to infer contextual reference points); and specifications regarding the software or hardware of the user's devices. In one embodiment, a user opts in or out of specific categories of data aggregation. For example, a user can opt in to provide all requested information, a subset of requested information, or no information. In one exemplary scenario, a user opts in to provide time-based information but opts out of providing location-based information (for all or a subset of the user's associated computing devices). In one embodiment, a user opts in or out of specific categories of data analysis. In one embodiment, users opt in or out of certain categories of data distribution. Such preferences may be stored in database 124. The operational stages of content composition program 122 are shown and described in further detail with respect to FIG.

[0028] The database 124 acts as a repository for data received, used, and / or generated by the content composition program 122. A database is an organized collection of data. The data may include information about user preferences (e.g., user computing device 130 1-Noverall user system settings, such as alert notifications for the first user; information about alert notification preferences; a set of sensor data captured from the first user, a set of historical data including emotions expressed by previous users of the content composition program 122; a historical data model; one or more generated suggestions; and any other data received, used, and / or generated by the content composition program 122.

[0029] Database 124 may be implemented in any type of device capable of storing data and configuration files that can be accessed and utilized by server 120, such as a hard disk drive, a database server, or flash memory. In one embodiment, database 124 is accessed by content composition program 122 to store and / or access data. In the illustrated embodiment, database 124 resides on server 120. In other embodiments, database 124 may reside on another computing device, server, cloud server, or may be spread across multiple devices (not shown) elsewhere in distributed data processing environment 100, provided that content composition program 122 has access to database 124.

[0030] The present invention may include various accessible data sources, such as databases 124, which may contain personal and / or sensitive business data, content, or information that a user does not want processed. Processing refers to any automatic or non-automated action or set of actions, such as collecting, recording, organizing, structuring, storing, adapting, altering, retrieving, examining, using, disclosing by transmission, disseminating, or otherwise making available, combining, restricting, erasing, or destroying personal and / or sensitive business data. The content composition program 122 enables the authorized and secure processing of personal data.

[0031] The content composition program 122 provides informed consent by notifying the user of the collection of personal and / or sensitive data, allowing the user to opt in or out of processing the personal and / or sensitive data. Consent can take several forms. Opt-in consent can require the user to take affirmative action before the personal and / or sensitive data is processed. Alternatively, opt-out consent can require the user to take affirmative action to prevent the processing of the personal and / or sensitive data before the personal and / or sensitive data is processed. The content composition program 122 provides information about the personal and / or sensitive data and the nature of the processing (e.g., type, scope, purpose, duration, etc.). The content composition program 122 provides the user with a copy of the stored personal and / or sensitive corporate data. The content composition program 122 allows for the correction or completion of incorrect or incomplete personal and / or sensitive data. The content composition program 122 allows for the immediate deletion of personal and / or sensitive data.

[0032] User Computing Device 130 1-N each of which represents a user interface 132 that allows a user to interact with the content composition program 122 on the server 120. 1-N As used herein, N represents a positive integer, and therefore the number of scenarios implemented in a given embodiment of the present invention is not limited to those shown in FIG. 1. In one embodiment, user computing device 130 1-N Each of the user computing devices 130 is a device that executes programmable instructions. 1-N are the respective user interfaces 132 1-NThe user computing device 130 may be an electronic device such as a laptop computer, tablet computer, netbook computer, personal computer, desktop computer, smartphone, or any programmable electronic device capable of executing the program and capable of communicating (i.e., sending and receiving data) with the content composition program 122 over the network 110. 1-N represents any programmable electronic device or combination of programmable electronic devices capable of executing machine-readable program instructions and capable of communicating with other computing devices (not shown) over network 110 in distributed data processing environment 100. In some embodiments, user computing device 130 1-N may include one or more user computing devices, such as a wearable computing device. Wearable computing devices, also referred to as "wearables," are a category of smart electronic devices capable of detecting, analyzing, and transmitting information about the wearer's body (e.g., vital signs and ambient data), potentially enabling instant biofeedback to the wearer. Wearable computing devices can be worn as accessories, incorporated into clothing items, implanted in the body, or even tattooed on the skin. Wearable computing devices are hands-free devices with practical uses that include a microprocessor and are enhanced with the ability to send and receive data over the Internet. For example, wearables may include, but are not limited to, smart watches, smart glasses, smart rings, and other similar wearable computing devices. In the illustrated embodiment, user computing device 130 1-N are user interfaces 132 1-N Contains each instance of

[0033] User Interface 132 1-NThe content composition program 122 on the server 120 and the user computing device 130 1-N In some embodiments, the user interface 132 acts as a local user interface between the users of the 1-N is a graphical user interface (GUI), web user interface (WUI), and / or voice user interface (VUI) that can display (i.e., visually) or present (i.e., audibly) text, documents, web browser windows, user options, application interfaces, and instructions for operation sent to the user from content composition program 122 over network 110. 1-N may also display or present alerts containing information (e.g., graphics, text, and / or sound) sent from content composition program 122 to the user over network 110. In one embodiment, user interface 132 1-N can send and receive data (i.e., to and from content composition program 122 via network 110, respectively). 1-N Through this, a user can opt in to the content composition program 122; create a user profile; set user preferences and alert notification preferences; receive one or more generated suggestions; receive feedback requests; and enter feedback.

[0034] User preferences are settings that can be customized for a particular user. A set of default user preferences is assigned to each user of the content composition program 122. A user preference editor can be used to update values ​​and change the default user preferences. Customizable user preferences include, but are not limited to, overall user system settings, specific user profile settings, alert notification settings, and machine-learned data collection / storage settings. Machine-learned data is a user's personalized corpus of data. Machine-learned data includes, but is not limited to, the results of past iterations of the content composition program 122.

[0035] Figure 2 is a flowchart, generally designated 200, illustrating operational stages of a content composition program 122 on a server 120 in the distributed data processing environment 100 of Figure 1, in accordance with one embodiment of the present invention. In one embodiment, the content composition program 122 operates to enhance a user's experience while listening to and / or viewing audiovisual content by modifying future audio and / or video frames of the audiovisual content based on one or more emotions expressed by a user. It should be understood that the process illustrated in Figure 2 illustrates one possible iteration of the process flow, and that this process flow may be repeated for each emotion expressed by a user in response to the audiovisual content.

[0036] In step 210, the content composition program 122 captures a set of sensor data in response to the first user expressing an emotion in response to the audiovisual content the first user is listening to and / or watching. In another embodiment, the content composition program 122 captures a set of sensor data in response to two or more first users expressing an emotion in response to the audiovisual content the first user is listening to and / or watching. The first users may be listening to and / or watching, but are not limited to, films (i.e., movies), television series, commercials, online streaming videos, video games, music albums, music songs, podcasts, webinars, slides, lecture notes, artwork, infographics, photos, and images. The first users wear IoT wearable devices on their bodies (i.e., user computing devices 130). 1-N ). The IoT wearable device may include, but is not limited to, a smart watch, smart glasses, smart ring, and other similar devices. In one embodiment, the content composition program 122 configures the IoT wearable device on the body of the first user (i.e., the user computing device 130). 1-N ) to capture a set of sensor data. The set of sensor data captured from the first user may include, but is not limited to, the first user's heart rate, the first user's pulse, the first user's respiratory rate, changes in the first user's nervous system, a set of neurological data for the first user, and movements made by the first user.

[0037] In one embodiment, the content composition program 122 identifies a first set of video frames of audiovisual content that the first user was listening to and / or watching when the first user expressed an emotion. The first set of video frames spans a time span that begins when the first user begins expressing an emotion and ends when the first user stops expressing an emotion, which is recorded on an IoT wearable device (i.e., user computing device 130) on the first user's body. 1-N ) is determined by a set of sensor data captured from

[0038] In step 220, the content composition program 122 analyzes the set of sensor data to generate one or more connotations. In one embodiment, the content composition program 122 compares the set of sensor data (e.g., the first user's heart rate, the first user's pulse, the first user's respiratory rate, changes in the first user's nervous system, the set of neurological data for the first user, and the first user's movements) with a set of historical data (i.e., with similar sets of sensor data) stored in a database (e.g., database 124) to identify emotions expressed by previous users of the content composition program 122. In one embodiment, the content composition program 122 uses an emotion learning model to classify the set of sensor data as an emotion (i.e., an emotion expressed by the first user) based on the results of the comparison with the historical data. In one embodiment, the content composition program 122 uses an emotion vector analysis method to convert the emotion into one or more connotations. In one embodiment, the content composition program 122 uses supervised machine learning techniques to convert the emotion into one or more connotations.

[0039] In one embodiment, when two or more first users are listening to and / or viewing the audiovisual content, the content composition program 122 performs this process (i.e., steps 210 and 220) for each user simultaneously listening to and / or viewing the audiovisual content. In one embodiment, the content composition program 122 compiles the captured set of data into a tabular format.

[0040] In step 230, the content composition program 122 applies an analytical process to score one or more connotations. In one embodiment, the content composition program 122 scores one or more connotations on a scale of 0 to 100. In one embodiment, the content composition program 122 scores one or more connotations based on the similarity (i.e., comparison) between the emotion expressed by the first user and the emotion predicted by the second user to be evoked by the audiovisual content. The second user is the creator of the audiovisual content. In one embodiment, the content composition program 122 generates a score for the first user listening to and / or watching the audiovisual content. If the emotion expressed by the first user is similar to the emotion predicted to be evoked by the second user, the generated score is on the higher end of the scale. If the emotion expressed by the first user is not similar to the emotion predicted to be evoked by the second user, the generated score is on the lower end of the scale.

[0041] In another embodiment, when two or more first users are listening to and / or watching audiovisual content, the content composition program 122 generates an aggregate score for the two or more users simultaneously listening to and / or watching the audiovisual content. In one embodiment, the content composition program 122 adjusts the feedback from two or more users simultaneously listening to and / or watching the audiovisual content using a median emotion feedback method. For example, 10 audience members are watching a movie. Five audience members like a scene in the movie and are generally happy to watch that scene in the movie. A happy emotion corresponds to a score of 10 on a scale of 1 to 10. However, five audience members do not like that scene in the movie and are generally unhappy to watch that scene in the movie. A unhappy emotion corresponds to a score of 5 on a scale of 1 to 10. The content composition program 122 uses the median emotion feedback method to determine that the median emotion score is 7.5. The content composition program 122 uses the median emotion score to determine the collective audience feedback. In another embodiment, the content composition program 122 uses a modal affective feedback method to coordinate feedback from two or more users simultaneously listening to and / or watching audiovisual content. For example, ten audience members are watching a movie. Seven audience members like a certain scene in the movie, while three audience members dislike that scene in the movie. The content composition program 122 uses a modal affective feedback method to determine that the feedback of the majority of the audience members predominates when determining collective audience feedback.

[0042] In decision step 240, content composition program 122 determines whether the score of one or more connotations exceeds a preconfigured threshold level. In one embodiment, content composition program 122 determines whether the score of one or more connotations exceeds a preconfigured threshold level by 10 points (i.e., +10) or does not exceed a preconfigured threshold level by 10 points (i.e., −10). In another embodiment, when two or more first users are listening to and / or watching the audiovisual content, content composition program 122 determines whether the total score of one or more connotations exceeds a preconfigured threshold level by 10 points (i.e., +10) or does not exceed a preconfigured threshold level by 10 points (i.e., −10). In one embodiment, content composition program 122 compares the total score to a historical data model stored in a database (e.g., database 124). The historical data model includes a mapping of emotions previously expressed by other users with similar sets of sensor data. If the content composition program 122 determines that the score of one or more connotations does not exceed a preconfigured threshold level (i.e., the emotion expressed by the first user is not similar to the emotion expected to be evoked by the second user and the generated score is on the lower side of the scale) (decision step 240, "Yes" branch), the content composition program 122 proceeds to step 250 and generates a first suggestion.If the content composition program 122 determines that the score of one or more connotations exceeds a preconfigured threshold level (i.e., the emotion expressed by the first user is similar to the emotion expected to be elicited by the second user and the generated score is on the higher side of the scale) (decision step 240, "No" branch), the content composition program 122 allows the audiovisual content to continue playing.

[0043] In step 250, the content composition program 122 generates a first suggestion. The first suggestion recommends one or more changes that the second user may make to the second set of video frames to more closely match the emotions expected to be evoked by the second user. The second set of video frames spans an upcoming time span. In one embodiment, the content composition program 122 outputs the first suggestion to the second user. In one embodiment, the content composition program 122 outputs the first suggestion to the second user on the second user computing device (i.e., user computing device 130). 1-N ) second user interface (i.e., user interface 132 1-N ) to output the first proposal to the second user.

[0044] In step 260, the content composition program 122 modifies the second set of video frames to more closely match the emotions expected to be evoked by the second user. In another embodiment, the content composition program 122 allows the second user to modify the second set of video frames to more closely match the emotions expected to be evoked by the second user. In one embodiment, the content composition program 122 allows the second user to modify the second set of video frames to remove other stored versions of frames that are similar to the first set of frames. Modifications may include, but are not limited to, adding, removing, and altering the second set of video frames.

[0045] In one embodiment, the content composition program 122 generates a second suggestion, which recommends storing multiple copies of the second set of frames in a single content pack to the second user. In one embodiment, the content composition program 122 allows the second user to select copies of the second set of frames based on the sentiment expressed by the first user for the first set of frames.

[0046] At decision step 270, the content composition program 122 allows the audiovisual content to continue playing. In one embodiment, in response to determining that the score of one or more connotations exceeds a preconfigured threshold level (i.e., the emotion expressed by the first user is similar to the emotion expected to be evoked by the second user, and the generated score is on the higher side of the scale) (decision step 240, “No” branch), the content composition program 122 allows the audiovisual content to continue playing in an unmodified state. In one embodiment, in response to modifying the second set of video frames to more closely match the emotion expected to be evoked by the second user (step 260), the content composition program 122 allows the audiovisual content to continue playing in a modified state.

[0047] In one embodiment, the content composition program 122 continues to play the audiovisual content on the user computing device 130 until the audiovisual content is finished. 1-N In one embodiment, the content composition program 122 continues to monitor the sensor data captured from the first user through the user computing device 130. 1-N determine whether a second set of data has been captured from the first user through the user computing device 130 1-NIf the content composition program 122 determines that a second set of data has been captured from the first user through the user computing device 130 (decision step 270, "Yes" branch), the content composition program 122 returns to step 220 and analyzes the set of data. 1-N If the content composition program 122 determines that a second set of data has not been captured from the first user through the audiovisual content (decision step 270, "No" branch), the content composition program 122 continues to monitor the user computing device 130 until the audiovisual content is finished. 1-N The system continues to monitor data from the first user that is captured via the system.

[0048] For example, a first user is watching a documentary about penguins, which was produced by a second user. In response to the first user expressing his or her feelings about the documentary, the content composition program 122 may transmit a message to the IoT wearable device (i.e., the user computing device 130) worn by the first user. 1-N), a set of sensor data is captured. The content composition program 122 also captures a first set of video frames from the documentary that was playing when the first user expressed emotion. The first user expressed emotion about the documentary when the first set of video frames showed the death of a penguin, which upset the first user. However, the second user did not intend to upset the first user. Instead, the second user intended to inform the first user of an event in the penguin's life. It would be a poor choice for the second user to show a clip of another penguin's death. Rather, the second user should show alternative clips, such as a clip of a baby penguin learning to swim or a clip of a group of penguins dancing. These alternative clips would uplift the first user's mood and thus hold the first user's attention. The content composition program 122 generates a suggestion recommending to the second user to modify the second set of video frames or to change the second set of video frames entirely to more closely match the emotions expected to be evoked by the second user.

[0049] 3 illustrates a block diagram of components of server 120 in distributed data processing environment 100 of FIG. 1 in accordance with one embodiment of the present invention. It should be understood that FIG. 3 is intended only as an illustration of one implementation and is not intended to imply any limitation with regard to the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be made.

[0050] Computing environment 300 includes an example environment for executing at least some of the computer code necessary to perform the methods of the present invention, such as content composition program 122 for generating mixed reality scenarios for cross-industry training. In addition to content composition program 122, computing environment 300 includes, for example, computer 301, wide area network (WAN) 302, end user device (EUD) 303, remote server 304, public cloud 305, and private cloud 306. In this embodiment, computer 301 includes processor set 310 (including processing circuitry 320 and cache 321), communications fabric 311, volatile memory 312, persistent storage 313 (including operating system 322 and the content composition program 122 identified above), peripheral device set 314 (including user interface (UI), device set 323, memory 324, and Internet of Things (IoT) sensor set 325), and network module 315. Remote server 304 includes remote database 330. The public cloud 305 includes a gateway 340, a cloud orchestration module 341, a set of host physical machines 342, a set of virtual machines 343, and a set of containers 344.

[0051] Computer 301, representing server 120 in FIG. 1, may take the form of a now-known or later-developed desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device capable of executing programs, accessing a network, or querying a database, such as remote database 330. As is well understood in the field of computer technology, and depending on the technology, execution of a computer-implemented method may be distributed among multiple computers and / or among multiple locations. However, in this description of computing environment 300, for purposes of brevity, the detailed discussion focuses on a single computer, specifically computer 301. While computer 301 is not shown in the cloud in FIG. 3, it may be located within a cloud. However, computer 301 is not required to reside within a cloud except to any extent that may be expressly indicated.

[0052] Processor set 310 includes one or more computer processors of any type now known or later developed. Processing circuitry 320 may be distributed across multiple packages, e.g., multiple coordinated integrated circuit chips. Processing circuitry 320 may implement multiple processor threads and / or multiple processor cores. Cache 321 is memory located within the processor chip package and is typically used for data or code that should be available for fast access by threads or cores executing on processor set 310. Cache memory is typically organized into multiple levels depending on relative proximity to the processing circuitry. Alternatively, some or all caches for a processor set may be located “off-chip.” In some computing environments, processor set 310 may be designed to operate with qubits and perform quantum computing.

[0053] Computer-readable program instructions are typically loaded onto the computer 301 to cause the processor set 310 of the computer 301 to perform a series of operational steps, thereby realizing a computer-implemented method, such that the instructions so executed instantiate the method specified in the flowcharts and / or descriptions of the computer-implemented method contained herein (collectively referred to as the "methods of the present invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 321 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 310 to control and direct the execution of the methods of the present invention. In the computing environment 300, at least some of the instructions for executing the methods of the present invention may be stored in the content composition program 122 in persistent storage 313.

[0054] Communications fabric 311 is the signal-conducting pathway that allows various components of computer 301 to communicate with one another. Typically, this fabric is made of switches and conductive pathways, such as the switches and conductive pathways that make up buses, bridges, physical input / output ports, etc. Other types of signal communication pathways may be used, such as fiber optic and / or wireless communication pathways.

[0055] Volatile memory 312 may be any type of volatile memory, now known or later developed. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory is characterized by random access, although this is not required unless expressly indicated. In computer 301, volatile memory 312 is located in a single package and is internal to computer 301; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located external to computer 301.

[0056] Persistent storage 313 is any form of non-volatile storage for a computer, now known or later developed. The non-volatility of this storage means that stored data remains regardless of whether power is supplied to computer 301 and / or directly to persistent storage 313. While persistent storage 313 can be read-only memory (ROM), typically at least a portion of persistent storage allows data to be written, deleted, and rewritten. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. Operating system 322 can take several forms, such as various known proprietary operating systems or open-source Portable Operating System Interface-type operating systems employing a kernel. The code included in content composition program 122 typically includes at least some of the computer code necessary to perform the methods of the present invention.

[0057] The peripheral device set 314 includes the set of peripheral devices of the computer 301. Data communication connections between the peripheral devices and other components of the computer 301 may be implemented in various ways, such as Bluetooth connections, near field communication (NFC) connections, connections formed by cables (such as universal serial bus (USB)-type cables), insertion-type connections (e.g., Secure Digital (SD) cards), connections formed through local area communication networks, and even connections formed through wide area networks such as the Internet. In various embodiments, the UI device set 323 may include components such as display screens, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. The memory unit 324 is external storage, such as an external hard drive, or insertable storage, such as an SD card. The memory unit 324 may be persistent and / or volatile. In some embodiments, the memory unit 324 may take the form of a quantum computing storage device for storing data in the form of quantum bits. In embodiments where computer 301 is required to have a large amount of storage (e.g., where computer 301 stores and manages large databases locally), this storage may be provided by a peripheral storage device designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple, geographically distributed computers. IoT sensor set 325 consists of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer and another sensor may be a motion detector.

[0058] Network module 315 is a collection of computer software, hardware, and firmware that enables computer 301 to communicate with other computers over WAN 302. Network module 315 may include hardware such as a modem or Wi-Fi signal transceiver, software for packetizing and / or depacketizing data for communications network transmission, and / or web browser software for communicating data over the Internet. In some embodiments, the network control and network forwarding functions of network module 315 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control and forwarding functions of network module 315 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing the methods of the present invention may be downloaded to computer 301 from an external computer or external storage device, typically via a network adapter card or network interface included in network module 315.

[0059] WAN 302 is any wide area network (e.g., the Internet) capable of communicating computer data over non-local distances by any technology for communicating computer data now known or later developed. In some embodiments, a WAN may be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi network. WANs and / or LANs typically include copper transmission cables, optical fiber transmissions, wireless transmissions, and computer hardware such as routers, firewalls, switches, gateway computers, and edge servers.

[0060] End-user device (EUD) 303 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 301) and may take any of the forms discussed above in connection with computer 301. EUD 303 typically receives useful and useful data from the operation of computer 301. For example, in the hypothetical case where computer 301 is designed to provide recommendations to the end user, the recommendations would typically be communicated from computer 301's network module 315 over WAN 302 to EUD 303. In this manner, EUD 303 can display or otherwise present the recommendations to the end user. In some embodiments, EUD 303 may be a client device such as a thin client, a heavy client, a mainframe computer, a desktop computer, and the like.

[0061] Remote server 304 is any computer system that provides at least some data and / or functionality to computer 301. Remote server 304 may be controlled and used by the same entity that operates computer 301. Remote server 304 represents a machine that collects and stores useful and useful data for use by other computers, such as computer 301. For example, in the hypothetical case where computer 301 is designed and programmed to provide recommendations based on historical data, then this historical data may be provided to computer 301 from remote database 330 of remote server 304.

[0062] Public cloud 305 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computer capacity, particularly data storage (cloud storage) and computing capacity, without direct active management by users. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. Direct active management of public cloud 305's computing resources is performed by computer hardware and / or software in cloud orchestration module 341. The computing resources provided by public cloud 305 are typically implemented by virtual computing environments running on various computers comprising host physical machine set 342, which is the universe of physical computers within and / or available in public cloud 305. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 343 and / or containers from container set 344. It is understood that these VCEs can be stored as images and transferred among and between various physical machine hosts, either as images or after instantiation of the VCE. Cloud orchestration module 341 manages the transfer and storage of images, deploys new instantiations of VCE, and manages active instantiations of VCE deployments. Gateway 340 is a collection of computer software, hardware, and firmware that enables public cloud 305 to communicate over WAN 302.

[0063] Some further description of virtualized computing environments (VCEs) is now provided. A VCE can be stored as an "image." A new, active instance of a VCE can be instantiated from the image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of an operating system in which the kernel allows the existence of multiple isolated user space instances, called containers. These isolated user space instances typically behave as actual computers from the perspective of programs running within them. A computer program running on a typical operating system may utilize all of the computer's resources, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and of the devices assigned to the container; this feature is known as containerization.

[0064] Private cloud 306 is similar to public cloud 305, except that the computing resources are available only for use by a single enterprise. While private cloud 306 is shown in communication with WAN 302, in other embodiments, the private cloud may be completely disconnected from the Internet and accessible only through a local / private network. A hybrid cloud is a composite of multiple clouds of different types (e.g., private, community, or public cloud types), often implemented by different vendors. While each of the multiple clouds remains a separate, discrete entity, the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the constituent clouds. In this embodiment, both public cloud 305 and private cloud 306 are part of a larger hybrid cloud.

[0065] The programs described herein are identified based on the applications for which they are implemented in specific embodiments of the invention. However, it should be understood that any particular program name herein is used merely for convenience, and thus the invention should not be limited to use in only any particular application identified and / or implied by such name.

[0066] Various aspects of the present disclosure are described by text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in computer program product (CPP) embodiments. For any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in successive flowchart diagrams may be performed in reverse order, as a single integrated step, simultaneously, or in an at least partially overlapping manner.

[0067] A computer program product embodiment ("CPP embodiment" or "CPP") is a term used in this disclosure to describe any set of one or more storage media (also referred to as "media"), collectively contained in a set of one or more storage devices, that collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. The computer-readable storage medium may be, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded device (such as pits / lands formed on the major surface of a punch card or disk), or any suitable combination of the foregoing. Computer-readable storage media, as the term is used in this disclosure, is not to be construed as storage in the form of a transitory signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses passing through fiber optic cables, electrical signals communicated through wires, and / or other transmission media. As will be understood by those skilled in the art, data is typically moved at some infrequent time during the normal operation of the storage device, such as during access, defragmentation, or garbage collection, but the above does not qualify a storage device as transitory because the data is not transitory while it is stored.

[0068] The foregoing description of various embodiments of the present invention has been presented for purposes of illustration and example, but is not intended to be exhaustive or limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the present invention. The terms used herein have been selected to best explain the principles of the embodiments, practical applications, or technical improvements over technologies found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. capturing, by one or more processors, a set of sensor data from an IoT device worn by a first user in response to the first user expressing an emotion toward the audiovisual content; analyzing, with the one or more processors, the set of sensor data to generate one or more connotations by transforming the emotions using sentiment vector analysis techniques and supervised machine learning techniques; scoring, by the one or more processors, the one or more connotations using an analytical process based on a similarity between the emotion expressed by the first user and an emotion expected to be evoked by a producer of the audiovisual content; determining, by the one or more processors, whether a score of the one or more connotations exceeds a preconfigured threshold level; generating, by the one or more processors, a suggestion to the producer of the audiovisual content in response to determining that the score does not exceed the preconfigured threshold level.

1. A computer-implemented method comprising:

2. 2. The computer-implemented method of claim 1, wherein the first user is a viewer of the audiovisual content, and the audiovisual content comprises at least one of a film, a television series, a commercial, an online streaming video, a video game, a music album, a music song, a podcast, a webinar, slides, lecture notes, a work of art, an infographic, a photograph, and an image.

3. 2. The computer-implemented method of claim 1, wherein the set of sensor data from the IoT device worn by the first user includes at least one of a heart rate of the first user, a pulse rate of the first user, a respiratory rate of the first user, changes in the nervous system of the first user, a set of neurological data of the first user, and movements performed by the first user.

4. Capturing the set of sensor data from the IoT device worn by the first user includes: identifying, by the one or more processors, a first set of video frames of the audiovisual content that the first user was watching when the first user expressed the emotion. The computer-implemented method of claim 1 , further comprising:

5. 5. The computer-implemented method of claim 4, wherein the first set of video frames spans a time span, the time span beginning when the first user begins to express the emotion and ending when the first user stops expressing the emotion.

6. generating the one or more connotations by analyzing the set of sensor data and transforming the emotions using the emotion vector analysis technique and the supervised machine learning technique, comparing, by the one or more processors, the set of sensor data with a set of historical data stored in a database to identify the emotions expressed by one or more previous users; and classifying, by the one or more processors, the set of sensor data as an emotion using an emotion learning model based on the results of the comparison. The computer-implemented method of claim 1 , further comprising:

7. Determining whether the score of the one or more connotations exceeds the preconfigured threshold level comprises: comparing the scores assigned to the one or more connotations by the one or more processors to a historical data model.

10. The computer-implemented method of claim 1, further comprising: wherein the historical data model includes a mapping of one or more emotions previously expressed by one or more previous viewers of the audiovisual content.

8. 2. The computer-implemented method of claim 1, wherein the suggestions include a set of feedback regarding how to modify a second set of video frames to more closely match the emotions expected to be evoked by the producer of the audiovisual content, the second set of video frames spanning an upcoming time span.

9. after generating the suggestions to the producer of the audiovisual content, allowing the one or more processors to modify the second set of video frames to more closely match the emotions expected to be evoked by the producer of the audiovisual content. The computer-implemented method of claim 1 further comprising:

10. one or more computer-readable storage media and program instructions stored on the one or more computer-readable storage media, the program instructions comprising: program instructions for capturing a set of sensor data from an IoT device worn by a first user in response to the first user expressing an emotion toward audiovisual content; program instructions for analyzing the set of sensor data to generate one or more connotations by transforming the emotions using sentiment vector analysis techniques and supervised machine learning techniques; program instructions for scoring the one or more connotations using an analytical process based on a similarity between the emotion expressed by the first user and an emotion expected to be evoked by a producer of the audiovisual content; program instructions for determining whether the score of the one or more connotations exceeds a preconfigured threshold level; and program instructions for generating a suggestion to the producer of the audiovisual content in response to determining that the score does not exceed the preconfigured threshold level.

1. A computer program product comprising:

11. Capturing the set of sensor data from the IoT device worn by the first user includes: program instructions for identifying a first set of video frames of the audiovisual content that the first user was watching when the first user expressed the emotion; 11. The computer program product of claim 10, further comprising:

12. generating the one or more connotations by analyzing the set of sensor data and transforming the sentiment using the sentiment vector analysis technique and the supervised machine learning technique; program instructions for comparing the set of sensor data with a set of historical data stored in a database to identify the emotions expressed by one or more previous users; and and program instructions for classifying the set of sensor data as an emotion using an emotion learning model based on a result of the comparison.

11. The computer program product of claim 10, further comprising:

13. Determining whether the score of the one or more connotations exceeds the preconfigured threshold level includes:

11. The computer program product of claim 10, further comprising program instructions for comparing the scores assigned to the one or more connotations with a historical data model, the historical data model comprising a mapping of one or more emotions previously expressed by one or more previous viewers of the audiovisual content.

14. 11. The computer program product of claim 10, wherein the suggestions include a set of feedback regarding how to modify a second set of video frames to more closely match the emotions expected to be evoked by the producer of the audiovisual content, the second set of video frames spanning an upcoming time span.

15. and program instructions for enabling the producer of the audiovisual content to modify the second set of video frames to more closely match the emotion expected to be evoked by the producer of the audiovisual content after generating the suggestion to the producer of the audiovisual content.

11. The computer program product of claim 10, further comprising:

16. one or more computer processors; one or more computer-readable storage media; program instructions collectively stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors; wherein the stored program instructions are: program instructions for capturing a set of sensor data from an IoT device worn by a first user in response to the first user expressing an emotion toward audiovisual content; program instructions for analyzing the set of sensor data to generate one or more connotations by transforming the emotions using sentiment vector analysis techniques and supervised machine learning techniques; program instructions for scoring the one or more connotations using an analytical process based on a similarity between the emotion expressed by the first user and an emotion expected to be evoked by a producer of the audiovisual content; program instructions for determining whether the score of the one or more connotations exceeds a preconfigured threshold level; and program instructions for generating a suggestion to the producer of the audiovisual content in response to determining that the score does not exceed the preconfigured threshold level. A computer system comprising:

17. Capturing the set of sensor data from the IoT device worn by the first user includes: program instructions for identifying a first set of video frames of the audiovisual content that the first user was watching when the first user expressed the emotion; 17. The computer system of claim 16, further comprising:

18. generating the one or more connotations by analyzing the set of sensor data and transforming the sentiment using the sentiment vector analysis technique and the supervised machine learning technique; program instructions for comparing the set of sensor data with a set of historical data stored in a database to identify the emotions expressed by one or more previous users; and and program instructions for classifying the set of sensor data as an emotion using an emotion learning model based on a result of the comparison.

17. The computer system of claim 16, further comprising:

19. Determining whether the score of the one or more connotations exceeds the preconfigured threshold level includes:

17. The computer system of claim 16, further comprising program instructions for comparing the scores assigned to the one or more connotations with a historical data model, the historical data model including a mapping of one or more emotions previously expressed by one or more previous viewers of the audiovisual content.

20. and program instructions for enabling the producer of the audiovisual content to modify the second set of video frames to more closely match the emotion expected to be evoked by the producer of the audiovisual content after generating the suggestion to the producer of the audiovisual content.

17. The computer system of claim 16, further comprising: