Adapt to the movie storyline
Monitoring audience situation and environmental information through neural network systems and dynamically adjusting media content, solving the problem that media content cannot be personalized in the prior art, and improving the audience experience and the effectiveness of media providers.
Patent Information
- Application Number
- CN202080059505.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-22
- Filing Date
- 2020-08-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2040-08-14
AI Technical Summary
In the prior art, media content cannot be dynamically changed according to user/viewer’s emotions, context and environment, resulting in the experience being static and unable to meet personalized needs.
The neural network system is adopted to monitor the audience through the situation context, background context and environment information modules, dynamically select and sort media story lines, and use the selection module and sorting module to realize real-time adjustment of media content.
It improves the dynamic and personalization of the media experience, enhances audience participation and satisfaction, and achieves the goals of media providers such as increasing view count and diversification of viewers.
Smart Images

Figure CN114303159B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence and artificial reality. More specifically, the present invention relates to dynamically changing and / or adapting one or more media (e.g., video and / or audio) streams based on user / audience context and / or user / audience environment to produce an enhanced audience experience and / or more effectively achieve the goals of media providers. Background Art
[0002] The fields of artificial intelligence and consumer electronics have merged with the film and entertainment industries to provide video and audio content to users in multiple venues and via multiple paths. Movies, concerts, and events are experienced not only in venues such as theaters, concert halls, and stadiums and on television, but also as media on computers and telephones, where content can be consumed on demand, continuously or intermittently, and in a variety of times, locations, and user situations.
[0003] However, generally, the media is only available for immutable content and fixed sequences of events, images, audio, and / or speech. As such, the media tends to be static, immutable, and unresponsive to changes in the environment in which the media is presented and / or the context of the user / audience consuming the media. In many cases, after the content is produced, it cannot be changed based on the mood, context, or environment of the user / audience, such as the composition (demographics), number, feelings, and emotions of the user / audience; the background of the user / audience; and / or the parameters of the environment in which the media is presented.
[0004] There is a need to be able to change media content and sequencing to adapt to the mood, context, and environment of the user / audience. Therefore, there is a need in the art to solve the above problems. Summary of the Invention
[0005] In a first aspect, the present invention provides a control system for adapting a media stream, the system comprising: a neural network having an input layer, one or more hidden layers, and an output layer, the input layer having a situation context input sub-layer and an environment input sub-layer, and the output layer having a selection / sorting output sub-layer and an environment output sub-layer, each of the layers having a plurality of neurons, each of the neurons having an activation; one or more situation context modules, each situation context module having one or more context inputs from one or more sensors for monitoring an audience and one or more emotional outputs connected to the situation context input sub-layer; one or more environment information modules having one or more environment sensor inputs and an environment output connected to the environment input sub-layer; one or more selection modules connected to the selection / sorting output sub-layer; and one or more sorting modules connected to the selection / sorting output sub-layer, wherein the selection module is operable to select one or more selected storylines, and the sorting module is operable to sort the selected storylines into a story to be played.
[0006] On the other hand, the present invention provides a method for training a neural network, comprising the steps of: respectively for a plurality of emotion activations and emotion neurons, inputting the emotion activation into the emotion neuron in a situation context input sublayer, which is a part of the input layer of the neural network, and the emotion activation forms an emotion input pattern; respectively for a plurality of environment activations and environment neurons, inputting the environment activation into the environment neuron in an environment input sublayer, which is a part of the input layer of the neural network, and the environment activation forms an environment input pattern; propagating the emotion input pattern and the environment input pattern through the neural network; determining how to change one or more weights and one or more biases by minimizing a cost function applied to the output layer of the neural network, the output layer having a selection / sorting output sublayer and an environment output sublayer, and each of the selection / sorting output sublayer and the environment output sublayer having an output activation; backpropagating to change the weights and biases; and repeating the last two steps until the output activation reaches a desired result, thereby causing the training to end.
[0007] On the other hand, the present invention provides a method for controlling the control implementation of a system according to any one of claims 1 to 14, the method comprising monitoring an audience and one or more emotion outputs connected to the situation context input sublayer; selecting one or more selected storylines; and sorting the selected storylines as the played story.
[0008] On the other hand, the present invention provides a computer program product for managing a system, the computer program product comprising a computer-readable storage medium that can be read by a processing circuit and stores instructions for executing a method for performing the steps of the present invention by the processing circuit.
[0009] On the other hand, the present invention provides a computer program stored on a computer-readable medium and loadable into the internal memory of a digital computer, the computer program comprising software code portions for performing the steps of the present invention when the program is run on the computer.
[0010] On the other hand, the present invention provides a story line control system, comprising: a neural network having an input layer, one or more hidden layers, and an output layer, the input layer having a situation context input sub-layer, a background context input sub-layer, and an environment input sub-layer, and the output layer having a selection / sorting output sub-layer and an environment output sub-layer, each layer having a plurality of neurons, and each neuron having an activation; one or more situation context modules, each situation context module having one or more context inputs from one or more sensors monitoring the audience and one or more emotional outputs connected to the situation context input sub-layer; one or more environment information modules having one or more environment sensor inputs and an environment output connected to the environment input sub-layer; one or more selection modules connected to the selection / sorting output sub-layer; and one or more sorting modules connected to the selection / sorting output sub-layer, wherein the selection module selects one or more selected story lines, and the sorting module sorts the selected story lines into the story to be played.
[0011] The present invention is a story line control system including a method of training method and an operating system.
[0012] A neural network is used in the system. In some embodiments, the neural network is a convolutional neural network (CNN). The neural network has an input layer, one or more hidden layers, and an output layer. The input layer is divided into a situation context input sub-layer, a background context input sub-layer (in some embodiments), and an environment input sub-layer. The output layer has a selection / sorting output sub-layer and an environment output sub-layer. Each layer (including sub-layers) has a plurality of neurons, and each neuron has an activation.
[0013] One or more situation context modules each have one or more context inputs from one or more sensors monitoring the audience and one or more emotional outputs connected to the situation context input sub-layer.
[0014] One or more environment information modules have one or more environment sensor inputs monitoring the audience environment and an environment output connected to the environment input sub-layer.
[0015] In some embodiments, one or more background context modules have one or more background inputs including audience characteristics and a background output connected to the environment input sub-layer.
[0016] One or more selection modules are connected to the selection / sorting output sub-layer and control the selection of one or more selected story lines.
[0017] One or more sorting modules are connected to the selection / sorting output sub-layer and sort the selected story so that the selected story modifies the story before or during the playback of the story, such as through the story line control system. Description of the Drawings
[0018] The present invention will now be described by way of example only with reference to preferred embodiments, as shown in the following drawings:
[0019] Figure 1 is a block diagram of an architecture of the present invention.
[0020] Figure 2 is a diagram of one or more users in an environment.
[0021] Figure 3 is a story having one or more main storylines and one or more alternative storylines, the alternative storylines including a start, an alternative start, a fork point, an end, and an alternative end.
[0022] Figure 4 is a system structure diagram of an embodiment of a neural network of the present invention.
[0023] Figure 5 is a flowchart of a training process for training an embodiment of the present invention.
[0024] Figure 6 is a flowchart of an operation process showing the steps of operating an embodiment of the present invention. Detailed Description of the Invention
[0025] Here, the term "audience" refers to one or more "users". The terms "user" and "audience" will be used interchangeably without loss of generality. As a non-limiting example, the audience context includes user mood, composition (e.g., user demographics - age, education, status, social and political views), profile (e.g., likes and dislikes), feelings, emotions, location, quantity, time of day, reaction to the environment, reaction to the media, and / or emotional state. In some embodiments, two user / audience contexts are used: background context and situation context.
[0026] Media refers to any sensory output of the system that causes one or more users to feel a physical sensation. Non-limiting examples of media include video, audio, voice, changes in the room / space environment (e.g., air flow, temperature, humidity, etc.), smell, lighting, special effects, etc.
[0027] The environment or user / audience environment is the sum of some or all of the sensory outputs of the system at any given time.
[0028] The experience or audience / user experience is the sensory experience felt by the user / audience at any given time caused by the environment (including the media / story being played).
[0029] One or more media include stories, each story having one or more "storylines", such as sub-components or segments. A storyline has a start and an end. Typically, the storylines are played, for example, output in sequence on one or more output devices to create one or more stories and associated environments as part of the story. One or more second storylines of the same and / or different media can start or end based on start / end and / or start or end at a queue point in the first storyline.
[0030] For example, a first video storyline plays a series of images of "Little Red Riding Hood" (Red) walking to the front door of "Grandmother's" house. When "Red" opens the door (a queue point in the first video storyline), a first audio storyline starts playing and emits the sound of the door opening. After opening the door, the first video and first audio storylines end, for example, at an end point. At the end point of the first video and first audio storylines, a second video storyline (at its start point) starts playing a series of images showing "Red" walking into the house. In the middle of the second video storyline, "Red" sees the "Big Bad Wolf". This queue point in the second video storyline queues the start of an ominous piece of music in a second audio segment (at its start point) and a first control system storyline that dims the lighting and turns up the audio volume.
[0031] The content, selection, and sequencing of these storylines drive one or more system outputs to create a dynamic user / audience environment during story playback. As the content, selection, and environment change over time, the user / audience experience changes over time.
[0032] A "story" is a selection of one or more storylines played in a sequence of storylines. Each storyline has one or more storyline contents. Each storyline is played / output in one or more media types, such as video, audio, environmental changes (temperature, airflow), smell, lighting, etc. Storyline selection, storyline content, and storyline sequencing all affect the user / audience environment and experience. Thus, changing any one of the storyline selection, content, and / or sequencing changes the user / audience environment and experience during the story. The change in the environment and experience results in a change in the user context, and the change in the user context is input into the system and, in turn, can change the storyline selection, content, and / or sequencing.
[0033] The present invention is a system and method for implementing story modification through real-time storyline selection and in-story sequencing during the conveyance and / or playback of stories to a user / audience. Story modification can be based on situational context (e.g., user / audience experience), environment, and / or background context.
[0034] In some embodiments, the system monitors / inputs user / audience context and determines how to change the selection and / or ordering of one or more storylines to change the story and / or the environment and thus change the audience experience. In some embodiments, the change is made based on some criteria, such as the goals of the media provider for experience creation, messaging, and / or measurement of audience / user response.
[0035] In some embodiments, two types of user / audience context are used: background context and situation context. The background context of the user / audience is the context that exists before the story is played. Non-limiting examples of background context include user composition (e.g., user demographics - age, education, socioeconomic status, income level, social and political views), user profiles (e.g., likes and dislikes), time of day, number of users in the audience, weather, and the location of the audience. Non-limiting examples of situation context include the emotions of one or more users in the audience, such as feelings, moods, reactions to the environment (e.g., verbal statements about the environment / story, facial expression reactions, etc.).
[0036] The system can collect background context information about one or more users in the audience. This information can be collected from social media, user input, monitoring the user's use of the system, image recognition, etc. In some embodiments, the input includes background context information that is not user information, such as weather. Background context information can be obtained from calendars, system clocks, web searches, etc. Information from user surveys can be used.
[0037] In addition, the system monitors and inputs situation context. In some embodiments, situation context is collected from a facial recognition system for identifying expressions and / or user feelings. In some embodiments, the system analyzes the word usage, meaning, expression, volume, etc. of audio input to determine the emotional state. In some embodiments, the system can measure, record, infer, and / or predict the audience's reaction to environmental input, media type, and / or changes in the story and storyline media content (e.g., through facial recognition and analysis and audio analysis) as changes in storyline selection and ordering.
[0038] When the sequence of storylines is presented in the story, the system monitors environmental input, such as sound, sound level, light level, light frequency, odor, humidity, temperature, air flow, etc.
[0039] In a preferred embodiment, a neural network is used. During the training phase, training data including situation context information, background context information, and / or environmental information is input into the neural network. The neural network is trained using known backpropagation techniques as described below.
[0040] During training, the input is fed into the input layer, and the output of the output layer is compared with the desired result. For example, this difference is used in minimizing the cost function to determine how to modify the interior of the neural network. The input is fed in again, and the new output is compared with the desired output. This process is iterated until the system is trained.
[0041] During the operation phase, the output of the neural network controls the selection and sequencing of the storylines and environmental outputs to produce one or more stories for media playback. The resulting stories are played for the audience (or a model of the audience) to create an environment.
[0042] During operation, the audience experience is monitored, and the storyline selection and / or sequencing (including environmental storyline outputs) are changed as determined by the trained neural network. To change the story, the selection controller modifies the storyline selection, and the sequence controller modifies the sequence of the storylines to produce a modified story. The modified story may have storylines that are added, deleted, and / or changed compared to the storylines in the previous story.
[0043] The trained system is more capable of selecting and sequencing storylines and environmental outputs to more predictably create one or more audience experiences during the operation phase. The operating system can create stories, monitor the audience experience, and dynamically change the selection and sequencing of one or more storylines (including environmental outputs) based on the immediate audience experience to create different environments and story outcomes for the audience with predictable audience reactions / responses. These modifications can be made for future stories or can occur "instantly" while the story is being played.
[0044] In some embodiments, again note that the operating system is trained to provide an output based on an audience experience that takes into account dynamic background context, situation context, and / or environment.
[0045] The present invention can increase audience engagement with media (such as movies) by using artificial intelligence to adjust its storylines, with the aim of making the experience more exciting, more thrilling, more impactful, etc. for the audience.
[0046] Since adapting the storylines will change the media presentation according to the background context and situation context to cause an improvement in the audience experience (e.g., as measured by achieving the goals of the media provider), and thus the viewing experience is more intense, the present invention can increase the number of viewings of media presentations (such as movies). These goals include selling more tickets, obtaining more web views, and / or reaching a more diverse audience.
[0047] The present invention can monitor and record the audience's experience of one or more storylines or stories, for example, by measuring the situation context, background context, and / or environment and storing this information, for example, on an external storage for later analysis and use.
[0048] In some embodiments, a Convolutional Neural Network (CNN) is used to modify story line selection and / or ranking and / or environmental control.
[0049] Referring now to the drawings, specifically Figure 1 , which is a block diagram of an architecture 100 of the present invention.
[0050] In one embodiment, a control system 160 is connected 110C to one or more networks, cloud environments, remote storage, and / or application servers 110. The network connection 110C can be connected to any standard network 110 using any known connection to the standard interface 110C, including the Internet, intranet, wide area network (WAN), local area network (LAN), cloud, and / or radio frequency connections such as WiFi, etc.
[0051] The network connection 110C is connected to the components of the control system 160 by any known system bus 101 through a communication interface 104. Among other things, media data can communicate with the system 160 through the communication interface 104.
[0052] One or more computer processors (such as a central processing unit (CPU) 103, a coprocessor 103, and / or a graphics processing unit (GPU)) are connected to the system bus 101.
[0053] One or more input / output (I / O) controllers 105 are connected to the system bus 101. The I / O controller 105 is connected to one or more I / O connections 115. The I / O connection / bus 115 connects I / O devices by hardwiring or wirelessly (such as radio frequency, optical connection, etc.). Examples of I / O devices include external storage devices 111 (such as portable data storage devices, such as CD / DVD devices, external hard disk drives, cloud storage devices, etc.); one or more input sensors 112; one or more displays 113, such as a graphical user interface (GUI); and one or more other controllers 114, such as an environmental controller. Media content and / or story line information can communicate with, for example, external storage devices 111, displays / GUIs 113, and other (such as environmental) outputs 114 through the network connection 110C and / or the I / O connection 115.
[0054] One or more memories 102 are also connected to the system bus 101. System software 106 (such as one or more operating systems), an operating storage device 107 (such as a cache memory), and one or more application software modules 108 are accessibly stored in the memory 102.
[0055] Other applications may reside in the memory 102 for execution on the system 160. One or more functions of these applications may also be executed external to the system 160, and the results provided to the system 160 for execution. The situation context module 120 analyzes inputs that affect the situation context during a storyline or story playback. The background context module 130 analyzes background context information. The environment module 140 analyzes the effects of the environment during a storyline or story playback. The selection module 150 selects a storyline to play. The sequencing module 165 selects the order in which the selected storyline is played. The outputs of the selection module 150 and the sequencing module 165 may be performed before the story is played, or dynamically during the story playback based on an analysis of the situation context 120, background context 130, environment information 140, and / or system configuration. The environment module 170 coordinates the environment output with the selected and sequenced storylines and programs, and controls the environment output through the I / O controller 105.
[0056] Figure 2 is a diagram of one or more users / audience members 250 in the user environment 200. The environment 200 includes a space having one or more inputs and one or more outputs. The inputs include sensors 112 such as an image capture device (e.g., a camera) 205, an audio capture device 210 (e.g., a microphone 210), and a graphical user interface (GUI) 113. The inputs also include environmental inputs such as a temperature detector 230, a humidity detector (not shown) 112, and a location detector 280. The outputs include the GUI 113, a television, a display, a home entertainment system 255, and an audio output device (e.g., a speaker) 215. Some of the environment outputs (e.g., 225) are controlled by the environment controller 105, e.g., via a bus 115, to control environmental parameters including lighting 225, temperature, humidity, and air flow.
[0057] Auxiliary devices such as a cellular phone 275, a data assistant, monitors (e.g., "fit bits" and heart rate monitors), a blood pressure monitor, a respiration detector, a motion detector, and a computer may provide inputs and / or outputs to the user, typically 250. Inputs such as the microphone 210, temperature sensor 230, cellular phone 275, and camera 205 and outputs such as the speaker 215, cellular phone 275, and lighting 225 may be located anywhere within the environment 200 including in toys 245, etc. The inputs and outputs may be hard-wired, e.g., using a bus 115, or wirelessly connected.
[0058] The audience (usually 250) includes one or more users 250. The users 250 have different characteristics. For example, one user 250 can be a parent 250P, while another user 250 can be a young person 250C. The audience 250 can change. Non-limiting examples of the audience 250 include one or more of the following: family members watching TV, a group watching a movie, a group playing a video game; a crowd at an entertainment venue, a commercial audience, students in a classroom, a group at a theme park, a cinema audience, and people attending a sports or political event.
[0059] The control system 160 can be located anywhere connected to the I / O devices (112, 113, 114, etc.). For example, the control system 160 can be located within the entertainment system 255. In some embodiments, the control system 160 is connected 110C to one or more networks / clouds 110 via the communication interface 104 (wired or wireless).
[0060] The sensor 112 collects signals (such as images and audio), especially signals from the users 250, and the signals are provided to the corresponding I / O controller 105 via the bus connection 115. The received signals are processed and the content of the signals is analyzed by one or more of the situation context module 120, the background context module 130, and the environment module 140. In some embodiments, information is also received from the network / cloud 110, user input 113 via the interface, sensors and / or inputs from auxiliary devices such as the cellular phone 275 and the storage device 111.
[0061] Using these inputs, the situation context module 120, the background context module 130, and the environment module 140 determine the situation context 260, the background context, and the environmental parameters (such as temperature, volume, lighting level, humidity, etc.) in the environment 200.
[0062] Non-limiting examples are given.
[0063] The stereo camera 205 provides a signal to the situation context module 120, which performs body pose recognition on the situation context module 120 to determine the reactions 260 of one or more users 250, such as surprise, laughter, sadness, excitement, etc.
[0064] The camera 205 provides a signal to the situation context module 120 that performs face recognition to determine the reaction 260 of the user 250 based on the facial expression.
[0065] The infrared camera 205 provides a signal to the situation context module 120 that determines the body and facial temperatures representing the reaction 260 of the user 250.
[0066] The thermometer 230 sensor 112 provides signals to the environment module 140 to measure the ambient temperature and ambient parameters of the environment 200.
[0067] The microphone 210 provides signals to the environment module 140 to determine the ambient sound level in the environment 200.
[0068] The microphone 210 provides signals to the situation context module 120 that performs speech and sound (e.g., crying, laughing, etc.) recognition. Natural language processing (NLP) can be used to determine the responses 260 of one or more users 250. In some embodiments, the situation context module 120 performs NLP to detect key phrases related to playing a story. NLP can detect the spoken phases of the user 250, which indicate the expectations of the user 250 regarding the story progress and / or end and / or the user's emotions.
[0069] Using the image data received from the sensor 112, the situation context module 120 can perform age recognition, which can give an indication of the user 250's response 260.
[0070] The situation context module 120 can use the motion data (e.g., rapid motion, no motion, gait, etc.) from the motion sensor 112 to determine the user response 260
[0071] The background context module 130 uses the information received from the network 110, the user input 113 (e.g., user surveys), and / or the stored memory 111 to form a user profile of one or more users 250. The information can be collected from statements and activities regarding social media accounts, information about user groups and friends on social media, user search activities, visited websites, requested product information, purchased items, etc. Similar information can be accessed from activities on auxiliary devices such as the cellular phone 275. Of course, appropriate permissions and access privileges are required when accessing this information.
[0072] The output of the situation context module 120 is a set of emotions and / or user responses, each with a response value / score representing the level of the emotion / response 260 of one or more users. This represents the emotional state or emotional pattern of each user 250. In some embodiments, this emotional pattern / state is input into the situation context sublayer 452 in the neural network 400 as described in Figure 4 In some embodiments, the multiple users 250 are first classified according to the emotional states of the multiple users 250. In some embodiments, an aggregator aggregates the emotions / responses 260 of the multiple users 250 to determine an aggregated set of emotions / responses 260 and corresponding values to represent the emotional states of one or more groups (250P, 250C) among the multiple audiences 250 and / or the overall audience 250. The emotional state can change while the story is being played.
[0073] The output of the background context module 130 is a user profile for one or more users 250. Each user profile has a plurality of user characteristics, which have associated characteristic values / scores representing the level of each user characteristic for a given user 250. In some embodiments, the user profiles are grouped according to similarity. In some embodiments, a profile aggregator aggregates the profiles of multiple users 250 to determine an aggregated profile with corresponding profile values / scores to represent one or more groups (250P, 250C) in the audience 250 and / or the profile of the overall audience 250. The profile may change while the story is being played, but in some embodiments, it is desirable for the change in the profile to be less dynamic. In some embodiments, as Figure 4 described, the profile of the characteristics is input as a mode of activation 410 into the background context sublayer 454.
[0074] The output of the environment module 140 is an environment profile of the environment 200. The environment profile has a plurality of environment parameters, each of which has an associated parameter value / score representing the level that each parameter has in the environment 200 at a given time. The environment profile changes while the story is being played. The control system 160 can also change the environment 200 and thus the environment profile by changing the selection and ordering of the storylines. In some embodiments, the environment profile of the environment parameters is input as a mode of activation 410 into the environment sublayer 456, as Figure 4 described therein.
[0075] Figure 3 A story 300 having one or more main storylines (304, 394) and one or more alternative storylines 383, including a start 301, an alternative start 302, a fork point 380, an ending 309, and an alternative ending 310.
[0076] Figure 3 A diagram of a story 300 having one or more storylines (303, 323, 333, 343, 363, 393, typically 383) of a main storyline or story that goes directly from a start 301 to an end 309. There are storylines 383 that define a first story and alternative storylines 383. An alternative story is developed from the first or original story by adding, inserting, and / or deleting storylines 383 in the first / original story.
[0077] The story 300 has a start (e.g., 301) and an end 309. There can be alternative starts 302 and alternative ends 310.
[0078] In addition, there are fork points. A fork point is a point in story 300 where story 300 can be changed by inserting, deleting, or adding storylines 383. In some embodiments, alternative storylines 383 begin and end at fork points (such as 340 and 345) and change the content of the original story without changing the continuity of the story.
[0079] Fork points typically coincide with the start of a storyline (such as 320, 330, 340, 360, and 370, typically 380) or with the end of a storyline (315, 325, 335, 365, typically 385). By matching the start of a storyline 380 with a fork point and the end of a storyline 385 with a fork point, story 300 can be changed by playing an inserted storyline (starting and ending at the fork point) instead of playing the storyline in the previous sequence.
[0080] For example, story 300 is initially sorted to start at 301 and play storyline 304 between point 301 and fork point 315. The system can change story 300 by starting at 302, playing storyline 303 instead of storyline 304, and ending storyline 303 at fork point 315, which is common to the end of both storylines 303 and 304. By replacing storyline 304 with storyline 303, control system 160 changes the selection (selecting 304 instead of 303) and order (playing 304 first instead of 303) of the storylines 383 in story 300.
[0081] As a further example, assume the original story 300 starts at 301 and continues by playing a single storyline directly from start 301 to end 309. This original story can be changed in multiple ways by selecting different storylines in a different order. The story can start at 301 and continue until fork point 320, where storyline 323 is selected and played until fork point 325. Alternatively, the original story can start at 302 and continue to fork point 330, where system 160 selects storyline 333 and returns to the original storyline at fork point 335. The system can also start playing a storyline at 301, and system 160 selects to start storyline 343 at fork point 340 and returns to the original storyline at fork point 345. In story 300, storylines 393 and 394 provide alternative endings (309 or 310), depending on which story ending system 160 selects and plays at fork point 370. Note that within an alternative storyline (such as 363) (such as at a queue point), there may be a fork point 367 where another storyline 305 can start or end.
[0082] By "playing" a storyline 383, one or more outputs can provide media corresponding to the selected storyline 383 sequenced and played to the viewer 250 as described above.
[0083] Figure 4 4 is a system architecture diagram of an embodiment of the neural network 400 of the present invention.
[0084] Neural network 400 includes a plurality of neurons, typically 405. Each neuron 405 may store a value called activation 410. For example, neuron 405 holds an activation 410 with a value of "3". For clarity, Figure 4 Most neurons and activations in the figure do not have reference numbers.
[0085] The neural network 400 includes a plurality of layers, such as 420, 422, 424, 425, and 426, typically 425. There is a first layer or input layer 420 and a last layer or output layer 426. Between the input layer 420 and the output layer 426, there are one or more hidden layers, such as 422, 424. Each layer 425 has a plurality of neurons 405. In some embodiments, the number of layers 425 and the number of neurons 405 in each layer of the layers 425 are determined empirically by experimentation.
[0086] In some embodiments, all neurons in the previous layer are each connected to each of the neurons in the next layer by an edge 415. For example, a typical neuron 406 in the next (hidden) layer 422 is individually connected to each neuron 405 in the input layer 420 by an edge 415. In some embodiments, one or more edges 415 have an associated weight W 418. In a similar manner 430, each neuron 406 in the next layer (e.g., 422) is individually connected to each neuron 405 in the previous layer (e.g., 420) by an edge 415. The same type of connection 415 is made between each neuron in the second hidden layer 424 to each neuron in the first hidden layer 422, and likewise between each neuron 495 in the output layer 426 and all neurons in the second hidden layer 424. For clarity, these connections 430 are not shown in FIG. Figure 4 Shown in.
[0087] In some embodiments, the activation 410 in each neuron 406 is determined by a weighted sum of the activations 410 of each connected neuron 405 in the previous layer. Each activation 410 is weighted by the weight (w, 418) of an edge 415, each edge connecting the neuron 406 to each of the corresponding neurons 405 in the previous layer (e.g., 420).
[0088] Thus, the pattern of activations 410 in the previous layer (e.g., 420) and the weights (w, 418) on each edge 415 accordingly determine the pattern of activations 406 in the next layer (e.g., 422). In a similar manner, the weighted sum of the set of activations 406 in the previous layer (e.g., 422) determines the activation (typically 405) of each neuron and, thus, the pattern of activation of the neurons in the next layer (e.g., 424). This process continues until there is a pattern of activation represented by the activations (typically 490) in each neuron 495 in the output layer 426. Thus, given the pattern of activations 405 at the input layer 420, the structure, weights (w, 418), and biases b (described below) of the neural network 400 determine the activation output pattern, which is the activation of each of the neurons 495 in the output layer 426. A change in the set of activations in the input layer 420 causes a change in the set of activations in the output layer 426 as well. The changing set of activations in the hidden layers (422, 424) is an abstract level that may or may not have physical meaning.
[0089] In some embodiments, the input layer 420 is subdivided into two or more sub - layers, such as 452, 454, and 456, typically 450.
[0090] The input sub - layer 452 is the situation context sub - layer 452 and receives the activations 410 from the output of the situation context module 120. The pattern of activations 410 at the input layer 452 represents the state of the emotion / reaction 260 of the audience 250. For example, the neurons 405 in the situation context sub - layer 452 represent reactions such as the audience being happy, excited, worried, etc.
[0091] The input sub - layer 454 is the background context sub - layer 454 and receives the activations 410 from the output of the background context module 130. The activation 410 of the neurons 405 in the background context sub - layer 454 represents the value / score of the feature in the user / audience 250 profile. In some embodiments, the pattern of activations 410 at the background context sub - layer 454 represents the background context state at a point in time of the user / audience 250 characteristic profile.
[0092] The input sub - layer 456 is the environment sub - layer 456 and receives the activations 410 from the output of the environment module 140. The activation 410 of the neurons 405 in the environment sub - layer 456 is the value / score representing the environmental parameters in the environment profile. In some embodiments, the pattern of activations 410 in the environment sub - layer 456 represents the environment profile state changing over time while the story is being played.
[0093] In some embodiments, the output layer 426 is subdivided into two or more sub - layers, such as 482 and 484, typically 480.
[0094] The output sub-layer 482 has neurons 495 that determine which storylines to select and when each of the selected storylines starts and ends, e.g., at the bifurcation points (380, 385).
[0095] The output sub-layer 484 has neurons 495 that determine how the control system 160 controls the I / O devices, particularly the environmental output 114 to change the environment 200.
[0096] A mathematical representation of the transformation from one layer to the next in the neural network 400 is as follows:
[0097]
[0098] or
[0099] a 1 = σ(Wa 0 + b)
[0100] where a n 1 is the activation 410 of the nth neuron 406 in the next level (here level 1); w n,k is the weight (w, 418) of the edge 415 between the kth neuron 405 in the current level (here level 0) and the nth neuron 406 in the next level (here level 1); and b n is the bias value of the weighted sum of the nth neuron 406 in the next level. In some embodiments, the bias value can be considered as the threshold for turning on the neuron.
[0101] The term σ is a scaling factor. For example, the scaling factor can be a sigmoid function or a rectified linear unit, e.g., ReLU(a) = max(0, a).
[0102] The neural network is trained by finding the values of all the weights (w, 418) and biases b. In some embodiments, known backpropagation methods are used to find the weight and bias values.
[0103] In some embodiments, to start the training, the weights and biases are set to random values or some initial set of values. The output (i.e., the activation pattern of the output layer 426) is compared with the desired result. The actual output is compared with the desired result through a cost function (e.g., the square root of the sum of squares of the differences), which measures how close the output is to the desired output for a given input. For example, through the gradient descent method, an iterative process is used to determine how to change the biases and weights in magnitude and direction to approach the desired output, thereby minimizing the cost function. The weights and biases are changed (i.e., propagated backward), and another iteration is performed. Multiple iterations are carried out until the output layer produces an activation pattern close to the desired result for the given activation pattern applied to the input layer.
[0104] In an alternative embodiment, the neural network is a Convolutional Neural Network (CNN). In a CNN, one or more hidden layers are convolutional layers where convolutions are performed using one or more filters to detect, emphasize, or de-emphasize patterns in the layer. There are different types of filters, for example, to detect sub-shapes in image and / or sound sub-patterns. In a preferred embodiment, the filter 470 is a matrix of values that convolves over the input of the layer to create a new pattern to the input of the layer.
[0105] Figure 5 is a flowchart of a training process 500 for training one embodiment of the present invention.
[0106] The training process 500 starts with an untrained neural network. Initially activated patterns are input 505 from the situation context module 120, the background context module 130, and the environment module 140 into the input layer 420. Alternatively, the input can be simulated.
[0107] For example, the output of the situation context module 120 is input 505 into the situation context sub-layer 452 as a pattern of activation 410 representing the state of the emotion / reaction 260 of the viewer 250. The output of the background context module 130 is input into the background context sub-layer 454 as a pattern of activation 410 representing the state of the background context state at a certain point in time (i.e., the state of the user / viewer 250 profile of features). The output of the environment module 140 is input 505 into the environment sub-layer 456 as a pattern of activation 410 representing the state of the environmental parameters in the environmental profile.
[0108] In step 510, the weights 418 and the bias b are initially set.
[0109] In step 520, the activation 410 is propagated to the output layer 426.
[0110] In step 530, the output of the sub-layer 482 is compared with the desired selection and ordering of the story line. Additionally, the output of the sub-layer 484 is compared with the desired configuration of the environmental output. The cost function is minimized to determine the magnitude and direction of the changes required for a new set of weights 418 and bias b.
[0111] In step 540, the new weights 418 and bias b are propagated backward by known methods, and a new set of outputs (482, 484) are received.
[0112] After re-entering the input activation, an inspection is made in step 550. If the outputs of the sub-layers 482 and 484 are within the tolerance of the desired selection and ordering of the story line and the environmental output, the process 500 ends. If not, control returns to step 530 to minimize the cost function, and the process iterates again.
[0113] Figure 6 It is a flowchart of an operation process 600 showing the steps of an embodiment of the present invention.
[0114] In step 605, the output of the situation context module 120 is input 605 to the situation context sublayer 452 as a pattern of activation 410 representing the state of the emotion / reaction 260 of the viewer 250. The output of the background context module 130 is input 605 to the background context sublayer 454 as a pattern of activation 410 representing the state of the background context state at a certain point in time (i.e., the state of the user / viewer 250 feature profile). The output of the environment module 140 is input 605 to the environment sublayer 456 as a pattern of activation 410 representing the state of the environmental parameters in the environmental profile.
[0115] In step 610, the output of the sublayer 484 is provided to the environment controller 170. The output of the environment controller 170 controls the environment output 114 through the I / O controller and the bus 115 to change the environment 620.
[0116] In step 610, the output of the sublayer 482 is provided to the selection module 150 and the sorting module 165. The selection module 150 determines which storylines 283 to play from the activation patterns in the sublayer 482. The sorting module 165 determines the playing order of the selected storylines.
[0117] In step 630 of one embodiment, the selection module 150 traverses the list of available storylines. If no storyline is selected, it continues through the list until it encounters and selects all the storylines specified by the activation pattern of the sublayer 482. In an alternative embodiment, the activation pattern of the sublayer 482 directly identifies which storylines are selected.
[0118] From the activation pattern of the sublayer 482, the sorting module 165 determines the order in which the selected storylines are to be played. The activation pattern of the sublayer 482 also determines the bifurcation points at which the selected storylines start and end. Step 640 monitors the playing of the story or the order in which the story is to be played. In step 640, the sorting module 165 determines whether a bifurcation point has been reached and whether the sorting module 165 needs to modify (e.g., add, delete, insert) one of the selected storylines at that bifurcation point. If not, the sorting module waits until the next bifurcation point is reached and makes the determination again.
[0119] If a storyline from the selected storylines needs to be played at a bifurcation point, then in step 650 the storyline changes. The new storyline is played and ends at the appropriate bifurcation point (e.g., as indicated by the activation pattern of the sublayer 482), and the original storyline is not played at that point in the sequence.
[0120] The description of various embodiments of the present invention has been given for illustrative purposes, but it is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be obvious to a person of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application, or the technological improvement present in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A control system for adapting media streams, the system comprising: A neural network having an input layer, one or more hidden layers, and an output layer, the input layer having a situation context input sub-layer and an environment input sub-layer, and the output layer having a selection / sorting output sub-layer and an environment output sub-layer, each of the layers having a plurality of neurons, and each neuron in the neurons having an activation; One or more situation context modules, each situation context module having one or more context inputs from one or more sensors for monitoring an audience and one or more emotion outputs connected to the situation context input sub-layer; One or more environment information modules having one or more environment sensor inputs and an environment output connected to the environment input sub-layer; One or more selection modules connected to the selection / sorting output sub-layer; And One or more sorting modules connected to the selection / sorting output sub-layer, wherein the selection module is operable to select one or more selected storylines, and the sorting module is operable to sort the selected storylines into a story to be played.
2. The system according to claim 1, wherein, The context input includes one or more of the following: facial image, infrared image, audio input, audio volume level, text, spoken word, cellular phone input, heartbeat, blood pressure, and respiratory rate.
3. The system according to claim 1 or 2, wherein The situation context module is operable to perform one or more of the following functions: Facial recognition, location recognition, natural language processing (NLP), identification of key phrases, and speech recognition.
4. The system according to claim 1, wherein The emotion output represents one or more of the following audience emotions: Feeling, mood, laughter, sadness, anticipation, fear, excitement, reaction to the environment, reaction to one or more storylines.
5. The system according to claim 1, wherein The environment sensor input includes one or more of the following: Audio acquisition device, video acquisition device, network connection, weather input, location sensor, cellular phone, thermometer, humidity detector, airflow detector, light detector, and motion sensor.
6. The system according to claim 1, wherein The environment output includes one or more of the following: volume control, lighting control, temperature control, humidity control, heating system control, cooling system control, and airflow control.
7. The system according to claim 1, wherein The selection module is operable to select one or more selected storylines based on the dynamic pattern of the environment output and the dynamic pattern of the emotion output processed by the neural network.
8. The system according to claim 1, wherein the sorting module is operable to sort one or more selected storylines based on the dynamic pattern of the environment output and the dynamic pattern of the emotion output processed by the neural network.
9. The system according to claim 1, wherein, The neural network is a convolutional neural network (CNN).
10. The system according to claim 1, wherein, One or more selected storylines start at a bifurcation point and end at a bifurcation point.
11. The system according to claim 1, wherein The neural network further includes a background context input sub-layer.
12. The system according to claim 11, wherein, The activation of one or more neurons in the background context input sub-layer includes a representation of one or more of the following: Demographics, age, education level, socioeconomic status, income level, social and political views, user profiles, likes, dislikes, time of day, number of users in the audience, weather, and location of the audience.
13. The system according to claim 11 or 12, wherein, The input to the background context input sublayer includes background context information of one or more users, the background context information being developed from one or more of the following sources: Social media posts, social media usage, cellular phone usage, audience surveys, search history, calendars, system clocks, image analysis, and location information.
14. The system according to claim 13, wherein, The selection module is operable to select one or more selected storylines, and the sorting module sorts the one or more selected storylines based on a dynamic pattern of the environmental output, background context, and sentiment output processed by the neural network.
15. The system according to claim 1, wherein the neural network is trained via the following steps: Inputting sentiment activations into sentiment neurons in a situation context input sublayer, respectively, for a plurality of sentiment activations and sentiment neurons, the situation context input sublayer being part of an input layer of the neural network, the sentiment activations forming a sentiment input pattern; Inputting environmental activations into environmental neurons in an environmental input sublayer, respectively, for a plurality of environmental activations and environmental neurons, the environmental input sublayer being part of an input layer of the neural network, the environmental activations forming an environmental input pattern; Propagating the sentiment input pattern and the environmental input pattern through the neural network; Determining how to change one or more weights and one or more biases by minimizing a cost function applied to an output layer of the neural network, the output layer having a selection / sorting output sublayer and an environmental output sublayer, each of the selection / sorting output sublayer and the environmental output sublayer having an output activation; Backpropagating to change the weights and biases; and Repeating the last two steps until the output activation reaches a desired result that causes the training to end.
16. The system according to claim 15, wherein, After the neural network training ends, the system is further configured to: Select one or more selected storylines; Insert the selected storylines into an initial story, a start point of the selected storylines starting at a start bifurcation point of the initial story, and an end point of the selected storylines ending at an end bifurcation point of the initial story.
17. The system according to claim 16, wherein The environmental output sublayer has a dynamic pattern of output activation that controls one or more environmental outputs associated with the selected storylines.
18. The system according to claim 16, wherein, The selection of the selected storylines and the sorting of the selected storylines are determined by a dynamic sentiment input pattern.
19. A method of controlling implementation for controlling a system according to any one of claims 1 to 18, the method comprising monitoring an audience and one or more sentiment outputs connected to the situation context input sublayer; Selecting one or more selected storylines; Sorting the selected storylines as stories to be played.
20. A computer program product for managing a system, the computer program product comprising: A computer-readable storage medium, readable by a processing circuit and storing instructions for execution by the processing circuit to perform the method according to claim 19.
21. A computer program stored on a computer-readable medium and loadable into the internal memory of a digital computer, the computer program comprising software code portions for performing the method according to claim 19 when the program is run on a computer.
Citation Information
Patent Citations
Method for using Widrow-Hoff learning algorithm
CN103903003A
Interactive play apparatus
WO2019076845A1