Determine lighting effects based on the level of speech in media content

By analyzing the audio part of the media content, especially the voice level, and adjusting the brightness and chromaticity of the light effect, the problem that the light effect in the prior art does not match the context of the media content is solved, and the adaptability and experience quality of the light effect are improved.

CN113261057BActive Publication Date: 2025-08-26SIGNIFY HOLDING BV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080008641.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-01-25
Filing Date
2020-01-09
Publication Date
2025-08-26
Estimated Expiration
2040-01-09

AI Technical Summary

Technical Problem

The prior art fails to fully consider the context of the media content when determining the light effect, resulting in suboptimal light effect presentation.

Method used

By analyzing the audio portion of the media content, especially the voice level, the degree and manner of light effects, including adjustments to brightness and chromaticity, to better match the semantic meaning of the media content.

Benefits of technology

A more suitable light effect presentation is achieved, the media content experience is enhanced, and suboptimal performance is avoided due to light effect mismatch.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113261057B_ABST
    Figure CN113261057B_ABST
Patent Text Reader

Abstract

A method includes obtaining (101) media content information and obtaining (103, 109) information indicating a degree of speech in an audio portion. The media content information includes media content and / or information determined by analyzing the media content, and the degree of speech is determined based on an analysis of the audio portion of the media content. The method further includes determining (107, 113) a degree to which the audio portion should be used to determine one or more light effects to be presented when the media content is being presented, and determining (117) the light effects. The degree is determined based on the degree of speech, and the light effects are determined based on an analysis (115) of the audio portion and an analysis of the video portion of the media content as a function of the degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system for determining one or more light effects to be presented when media content is being presented, the one or more light effects being determined based on an analysis of the media content.

[0002] The invention further relates to a method of determining one or more light effects to be presented when media content is being presented, the one or more light effects being determined based on an analysis of the media content.

[0003] The invention also relates to a computer program product enabling a computer system to perform such a method. Background Art

[0004] The versatility of connected light systems (such as Philips Hue) continues to grow, offering users more and more features. These new features include contextual awareness, intelligent automated behaviors, new forms of light use (such as entertainment), and more. For example, Hue Entertainment enhances the experience of watching movies, listening to music, and / or playing games by using light scripts or by creating light effects based on audio and / or video analysis. The latter is achieved using the Hue Entertainment application HueSync, which automatically creates light effects using color extraction algorithms.

[0005] An ideal lighting system for entertainment supports and enhances the experience of specific content. Currently, the focus is on low-level image statistics, such as color values ​​and image motion. However, these statistics fail to consider the semantic dimension of the scene. Two scenes that are statistically nearly identical can convey vastly different meanings.

[0006] Without context, it's impossible to determine the semantic (intended) meaning of an image of an empty bench in the grass; for example, it could be an image intended to convey a beautiful summer day or a walk in the park with family. However, when one considers that the image originated from a funeral home, the image takes on a different dimension, perhaps one of sadness or grief. Rendering lighting effects based on media content without the context of that content often results in suboptimal lighting effects.

[0007] WO 2007 / 119277 A1 discloses a device that controls a lighting device to produce a lighting effect while a video is being presented, and that takes into account the context of the video, in the form of its genre. Specifically, WO 2007 / 119277 A1 discloses a lighting control data generation unit that generates lighting control data to control a lighting device so that the lighting device emits lighting light based on the genre (e.g., a music program, a sports event, etc.) and a characteristic value of video data displayed on a display device. Regardless of the characteristic value, the lighting device continuously emits lighting light when the displayed video has a predetermined genre.

[0008] A disadvantage of WO 2007 / 119277 A1 is that by only considering the genre of the video, the rendered light effects are still suboptimal. Summary of the Invention

[0009] A first object of the present invention is to provide a system that is able to determine one or more light effects while taking the context of the media content into account in a better way in order to create more suitable light effects.

[0010] A second object of the present invention is to provide a method which is able to determine one or more light effects while taking the context of the media content into account in a better way in order to create more suitable light effects.

[0011] In a first aspect of the present invention, a system for determining one or more light effects to be presented when media content is being presented, the one or more light effects being determined based on an analysis of the media content, the system comprising at least one input interface, at least one output interface, and at least one processor, the at least one processor being configured to: obtain media content information using the at least one input interface, the media content information comprising the media content and / or information determined by analyzing the media content; and obtain information indicating a degree of speech in the audio portion, the degree of speech being determined based on an analysis of the audio portion of the media content.

[0012] At least one processor is further configured to: determine the extent to which the audio portion should be used to determine one or more light effects, the extent being determined based on the determined degree of speech; determine one or more light effects to be presented on one or more light sources when media content is being presented, the one or more light effects being determined based on an analysis of the audio portion according to the extent and at least based on an analysis of the video portion of the media content; and use the at least one output interface to control the one or more light sources to present the one or more light effects and / or output a light script specifying the one or more light effects.

[0013] By using speechiness as an indicator of the semantic meaning of a scene, the context of the media content can be better considered in order to create more appropriate lighting effects. Even when only the spectral composition of speech is considered, this can still be highly informative about the semantic meaning of a scene (e.g., whispering versus screaming or laughing versus crying). Scenes containing a lot of dialogue will generally benefit more from subtle lighting effects than scenes that are visually similar (in terms of overall scene dynamics, saturation, and color) but do not include a lot of dialogue.

[0014] For example, the voice level may include the amount of voice and / or one or more voice categories.For example, the system may be part of a lighting system including one or more devices, or may be used in a lighting system including one or more lighting devices.

[0015] The degree can indicate whether the brightness and / or chromaticity of the one or more light effects should be determined based on the intensity and / or loudness of the audio portion. Changing the brightness and / or chromaticity of light effects based on the intensity and / or loudness of the audio portion of a media content item is particularly beneficial for music video clips and scenes with sound effects (such as explosions), but is not appropriate for scenes with a large amount of dialogue. The intensity of audio is typically the power carried by a sound wave per unit area in a direction perpendicular to the area. The loudness of audio is typically a subjective perception of sound pressure.

[0016] As a first example, a light effect with high brightness may be presented in response to an audio section with high intensity and / or loudness, and a light effect with low brightness may be presented in response to an audio section with low intensity and / or loudness. As a second example, a light effect with saturated colors may be presented in response to a segment of an audio section with high intensity and / or loudness, and a light effect with desaturated colors may be presented in response to a segment of an audio section with low intensity and / or loudness.

[0017] Alternatively or additionally, the degree may indicate whether the brightness and / or color of the one or more light effects should be determined based on one or more different characteristics of the audio portion. Speechiness is typically determined based on characteristics other than audio intensity and / or loudness. The brightness and / or color of the light effects may also be varied based on these other characteristics: for example, based on a perceived emotion determined from the narration and / or singing. Perceived emotion may be determined, for example, as described in the Proceedings of the ISCA Symposium on Speech and Emotion (https: / / www.isca-speech.org / archive_open / speech_emotion / spem.pdf).

[0018] The degree of speech in the audio portion can be determined by determining the amount of speech in the audio portion and classifying the audio portion as primarily speech or primarily non-speech based on the amount of speech. This classification can be used as described in the next two paragraphs.

[0019] The at least one processor may be configured to determine a first degree as the degree based on the audio portion being classified as primarily speech, and to determine a second degree as the degree based on the audio portion being classified as primarily non-speech, the second degree indicating that the brightness and / or chromaticity of the one or more light effects should be determined based on the intensity and / or loudness of the audio portion, and the first degree indicating that the brightness and / or chromaticity of the one or more light effects should not be determined based on the intensity and / or loudness of the audio portion. Changing the brightness and / or chromaticity of light effects based on the intensity and / or loudness of the audio portion of the media content item is particularly beneficial for music video clips and scenes with sound effects (such as explosions), but is not appropriate for scenes with a large amount of dialogue.

[0020] The at least one processor can be configured to determine the one or more light effects using a first luminance and / or chromaticity range based on the audio portion being classified as primarily speech, and using a second luminance and / or chromaticity range based on the audio portion being classified as primarily non-speech, the first luminance and / or chromaticity range having a lower average luminance and / or chromaticity than the second luminance and / or chromaticity range. Typically, scenes classified as primarily speech focus on dialogue, and these scenes preferably use lower intensity light than scenes classified as primarily non-speech (which typically focus on visual aspects) so as not to distract from the dialogue.

[0021] The degree of speech in the audio portion can be determined by classifying the audio portion as diegetic or non-diegetic. Non-diegetic sound is typically defined as sound that comes from a source outside the story space, such as a narrator's comments, sound effects added for dramatic effect, or mood music. Diegetic sound is typically defined as sound whose source is visible on screen or whose source is implied by the action of the film, such as a character's voice, sounds made by objects in the story, or music from instruments in the story. This classification is often difficult to detect from the audio and may therefore be manually included in the content metadata. It may sometimes be possible to detect whether the source of speech / sound in an audio portion is on screen or off screen and influence the lighting effects accordingly.

[0022] When the speech in the audio portion is classified as diegetic or non-diegetic, this can be used to determine the lighting effects based on audio analysis (and optionally video analysis) if the speech is classified as non-diegetic, and can be used to determine the lighting effects based solely on video analysis if the speech is classified as diegetic. The diegetic / non-diegetic classification can also be used, for example, to distinguish between theme songs played for atmosphere (non-diegetic) and songs that are part of a movie, such as songs listened to by characters in a club (diegetic). In the former case, for example, the lighting effects can be determined based solely on video analysis. In the latter case, for example, the lighting effects can be determined based on audio analysis (e.g., to help create the feeling of being in a club).

[0023] The degree of speech in the audio portion can be determined by classifying the audio portion into one of a plurality of categories, the plurality of categories comprising at least two of the following: conversation, whispering, screaming, narration, and singing. This classification can be used as described in the next two paragraphs.

[0024] The at least one processor may be configured to determine a first degree as the degree based on the audio portion being classified as conversation, and to determine a second degree as the degree based on the audio portion being classified as singing, the second degree indicating that the brightness and / or chromaticity of the one or more light effects should be determined based on the intensity and / or loudness of the audio portion, and the first degree indicating that the brightness and / or chromaticity of the one or more light effects should not be determined based on the intensity and / or loudness of the audio portion. In the event that the audio portion is classified as singing (rather than being classified as conversation), normal light effects may be presented, i.e., light effects determined based on an analysis of the audio portion. This is beneficial, for example, if a music video clip is classified as primarily speech due to the presence of singing or if the audio portion is not classified as primarily speech or primarily non-speech.

[0025] The one or more light effects may include a plurality of light effects, and the at least one processor may be configured to determine a transition speed between the plurality of light effects based on the classification. For example, if the audio portion is classified as screaming, the dynamics of the light effect may be adjusted to high, if the audio portion is classified as conversation, to medium, and if the audio portion is classified as whispering, to low. The same transition speed may be used for transitions between different chroma settings and for transitions between different brightness settings, but different transition speeds may alternatively be used.

[0026] The audio portions can be classified by analyzing the spectral composition of the audio portions. For example, by considering the spectrum and intensity differences between casual speech and shouted speech, it is possible to determine whether people are talking at a conversational level or screaming.

[0027] The one or more light effects may include a plurality of light effects, and the at least one processor may be configured to determine whether the amount of speech in the audio portion exceeds a threshold and determine a transition speed between the plurality of light effects based on the amount of speech exceeding the threshold. For example, a scene that includes a lot of conversation may be presented using low dynamics, while the same scene that includes a lot of screaming—even though the audio portion of this scene may have the same intensity and / or loudness—may be presented with a higher dynamics. The same transition speed may be used for transitions between different chroma settings and for transitions between different brightness settings, but different transition speeds may be used instead.

[0028] The at least one processor can be configured to determine the words spoken in the audio portion by identifying the spoken words in the audio portion and / or obtaining the spoken words from subtitles associated with the media content. The words spoken in the audio portion can be used to more accurately determine the atmosphere of the scene. As a first example, highly dynamic light effects can be presented for scenes that are emotionally charged, and slightly dynamic light effects can be presented for scenes that are not emotionally charged. As a second example, it may be inappropriate to present light effects in a cheerful green color during a funeral scene. Instead, a softer, desaturated green may be more suitable.

[0029] The at least one processor may be configured to determine the amount of speech by using subtitles associated with the media content and / or by focusing on a center channel in or obtained from the audio portion. Since the center channel in a surround setup typically includes dialogue, this is the best channel to focus on to determine the amount of speech and / or identify spoken words. Although a stereo audio portion may not include a center channel, such a center channel may then be obtained from the audio portion by determining common components in the two stereo channels. The size of the subtitle file or the number of words in the subtitle file may be a good indicator of the amount of speech in the media content.

[0030] In a second aspect of the present invention, a method for determining one or more light effects to be presented when media content is being presented, the one or more light effects being determined based on an analysis of the media content, the method comprising: obtaining media content information, the media content information including the media content and / or information determined by analyzing the media content; and obtaining information indicating a degree of speech in the audio portion, the degree of speech being determined based on an analysis of the audio portion of the media content.

[0031] The method further includes: determining the extent to which the audio portion should be used to determine one or more light effects, the extent being determined based on the determined degree of speech; determining one or more light effects to be presented on one or more light sources when media content is being presented, the one or more light effects being determined based on an analysis of the audio portion and at least based on an analysis of the video portion of the media content as a function of the extent; and controlling the one or more light sources to present the one or more light effects and / or outputting a light script specifying the one or more light effects. The method can be performed by software running on a programmable device. This software can be provided as a computer program product.

[0032] Furthermore, a computer program for carrying out the methods described herein and a non-transitory computer-readable storage medium storing the computer program are provided. For example, the computer program can be downloaded from or uploaded to an existing device, or stored when these systems are manufactured.

[0033] A non-transitory computer-readable storage medium stores software code portions that, when executed or processed by a computer, are configured to perform executable operations for determining one or more light effects to be presented when media content is being presented, the one or more light effects being determined based on an analysis of the media content. The executable operations include obtaining media content information, the media content information including the media content and / or information determined by analyzing the media content; and obtaining information indicating a degree of speech in an audio portion, the degree of speech being determined based on the analysis of the audio portion of the media content.

[0034] The executable operations further include: determining the extent to which the audio portion should be used to determine one or more light effects, the extent being determined based on the determined degree of speech; determining one or more light effects to be presented on one or more light sources when the media content is being presented, the one or more light effects being determined based on an analysis of the audio portion according to the extent and at least based on an analysis of the video portion of the media content; and controlling the one or more light sources to present the one or more light effects and / or output a light script specifying the one or more light effects.

[0035] As will be appreciated by those skilled in the art, aspects of the present invention may be embodied as devices, methods or computer program products. Thus, aspects of the present invention may take the form of complete hardware embodiments, complete software embodiments (including firmware, resident software, microcode, etc.) or embodiments of combined software and hardware aspects, which may all be generally referred to herein as "circuits," "modules," or "systems." The functions described in this disclosure may be implemented as algorithms executed by a processor / microprocessor of a computer. Additionally, aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable media having a computer-readable program code embodied (e.g., stored) thereon.

[0036] Any combination of one or more computer-readable media can be utilized. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or apparatus, or any suitable combination of the foregoing. More specific examples of computer-readable storage media can include, but are not limited to, the following: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of the present invention, a computer-readable storage medium can be any tangible medium that can contain or store a program used by or in conjunction with an instruction execution system, device or device.

[0037] A computer-readable signal medium may include a propagated data signal having computer-readable program code embodied therein (e.g., in baseband or as part of a carrier wave). Such a propagated signal may take any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0038] The program code embodied on computer-readable media can use any suitable medium to transmit, including but not limited to wireless, wired, optical fiber, cable, RF etc., or any suitable combination as mentioned above.The computer program code for realizing the operation of aspects of the present invention can be written with any combination of one or more programming languages, and these one or more programming languages ​​comprise object-oriented programming languages ​​(such as Java (TM), Smalltalk, C++ etc.), traditional procedural programming languages ​​(such as " C " programming languages ​​or similar programming languages) and functional programming languages ​​(such as Scala, Haskel etc.).Program code can be used as an independent software package and is executed completely on the user's computer, partly on the user's computer, partly on the user's computer and partly on a remote computer, or is executed completely on a remote computer or server. Under the latter scenario, the remote computer can be connected to the user's computer by any type of network (including local area network (LAN) or wide area network (WAN)), or can be connected (for example, by using the internet of an internet service provider) with an external computer.

[0039] Aspects of the present invention are described below with reference to the flowchart illustrations and / or block diagrams of the methods, devices (systems) and computer program products according to embodiments of the present invention. It will be understood that each frame of the flowchart and / or block diagram and the combination of frames in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device, particularly a microprocessor or central processing unit (CPU), to produce a machine so that instructions executed by the processor of the computer, other programmable data processing device or other equipment create a device for implementing the function / action specified in the flowchart and / or one or more block diagram frames.

[0040] These computer program instructions may also be stored in a computer-readable medium, which may direct a computer, other programmable data processing apparatus, or other device to operate in a specific manner so that the instructions stored in the computer-readable medium produce an article of manufacture that includes instructions for implementing the functions / actions specified in the flowchart and / or one or more block diagram blocks.

[0041] The computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide a process for implementing the functions / actions specified in the flowchart and / or block diagram blocks.

[0042] The flow chart and block diagram in the figure illustrate the possible implementation architecture, function and operation of the device, method and computer program product according to various embodiments of the present invention.In this respect, each frame in the flow chart or block diagram can represent a module, segment or part of a code, which includes one or more executable instructions for implementing the specified (multiple) logical functions.It should also be noted that in some alternative embodiments, the functions described in the frame may not occur in the order described in the figure.For example, the two frames shown in succession can in fact be performed substantially simultaneously, or sometimes these frames can be performed in reverse order depending on the functions involved.It will also be noted that each frame of the block diagram and / or flow chart and the combination of the frames in the block diagram and / or flow chart can be implemented by a system based on dedicated hardware or a combination of dedicated hardware and computer instructions that performs a specified function or action. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] These and other aspects of the invention are apparent from and will be further elucidated, by way of example, with reference to the accompanying drawings, in which:

[0044] Figure 1 is a block diagram of an embodiment of the system;

[0045] Figure 2 is a flow chart of a first embodiment of the method;

[0046] Figure 3 is a flow chart of a second embodiment of the method;

[0047] Figure 4 is a flowchart of a third embodiment of the method;

[0048] Figure 5 is a flowchart of a fourth embodiment of the method;

[0049] Figure 6 is a flowchart of a fifth embodiment of the method;

[0050] Figure 7 is a flowchart of a sixth embodiment of the method;

[0051] Figure 8 An example of audio classification of a first media item is shown;

[0052] Figure 9 An example of audio classification of a second media item is shown; and

[0053] Figure 10 is a block diagram of an exemplary data processing system for executing the method of the present invention.

[0054] Corresponding elements in the drawings are denoted by the same reference numerals. DETAILED DESCRIPTION

[0055] Figure 1 An embodiment of a system for determining one or more light effects to be presented when media content is being presented is shown: a mobile device 1. The one or more light effects are determined based on an analysis of the media content. This analysis can be performed by the mobile device 1 or by another device. The mobile device 1 is connected to a wireless LAN access point 23. A bridge 11 is also connected to the wireless LAN access point 23, for example via Ethernet. The light sources 13-17 communicate wirelessly with the bridge 11, for example using the Zigbee protocol, and can be controlled, for example, by the mobile device 1 via the bridge 11. For example, the bridge 11 can be a Philips Hue bridge and the light sources 13-17 can be Philips Hue lamps. In an alternative embodiment, the light sources are controlled without a bridge.

[0056] The TV 27 is also connected to the wireless LAN access point 23. For example, media content can be presented by the mobile device 1 or by the TV 27. The wireless LAN access point 23 is connected to the Internet 24. An Internet server 25 is also connected to the Internet 24. For example, the mobile device 1 can be a mobile phone or a tablet. For example, the mobile device 1 can run the Philips Hue Sync application. The mobile device 1 includes a processor 5, a receiver 3, a transmitter 4, a memory 7, and a display 9. Figure 1 In the embodiment of FIG. 1 , the display 9 comprises a touch screen. The mobile device 1 , the bridge 11 and the light sources 13 - 17 are part of a lighting system 21 .

[0057] exist Figure 1 In an embodiment, the processor 5 is configured to use the receiver 3 to obtain media content information. The media content information includes media content and / or information determined by analyzing the media content. For example, the media content information can be obtained from an internet server 25. The processor 5 is further configured to obtain information indicating the degree of voice in the audio portion. For example, this information can be obtained from the media content information. The degree of voice is determined based on the analysis of the audio portion of the media content. The processor 5 is further configured to determine the degree to which the audio portion should be used to determine one or more light effects. The degree is determined based on the determined degree of voice.

[0058] The processor 5 is further configured to determine one or more light effects to be presented on one or more light sources, such as one or more of the light sources 13-17 or the light sources that have not yet been identified, when the media content is being presented. The one or more light effects are determined based on an analysis of the audio portion and at least based on an analysis of the video portion of the media content according to the extent. The processor 5 is further configured to use the transmitter 4 to control one or more of the light sources 13-17 to present the one or more light effects and / or to use an internal interface (not shown) to output a light script specifying the one or more light effects to the memory 7.

[0059] For example, the degree may indicate whether the brightness and / or chromaticity of one or more light effects should be determined based on the intensity and / or loudness of the audio portion. Depending on the algorithm used for light effect creation, different ways of applying speech classification can be envisaged:

[0060] Transition Speed. If the colors used for light effect creation are extracted from a predefined analysis area within the on-screen content (e.g., as done in HueSync), voice classification can then be used to influence the transition speed between light effects that present the extracted colors.

[0061] When transitioning to a light effect, the colors extracted from the screen may be desaturated to lighter colors, or saturated to more vivid colors.

[0062] Brightness. Similar to above, you can adapt brightness but not saturation.

[0063] Extraction algorithm. Instead of modifying the colors extracted from the screen, speech classification can control what algorithm is used to select the color, what color is selected, and from which analysis area.

[0064] Audio Input: Often, the primary way to select the intensity and chroma of the light is based on the intensity and chroma of the video signal. However, in addition to that, some additional intensity (i.e., brightness) modulation is often added based on the audio intensity and / or loudness. This can make certain effects (such as explosions) extra dramatic by intensifying the effect or providing any effect (as they may be detectable in audio but not in video). However, for speech, it is clear that such intensity variations based on the audio signal are highly undesirable. Therefore, this audio input will then be enabled / disabled depending on whether speech is detected.

[0065] exist Figure 1In the embodiment of the mobile device 1 shown in FIG, the mobile device 1 includes a processor 5. In alternative embodiments, the mobile device 1 includes multiple processors. The processor 5 of the mobile device 1 can be a general-purpose processor (e.g., from Qualcomm or based on ARM) or a dedicated processor. For example, the processor 5 of the mobile device 1 can run an Android or iOS operating system. The memory 7 can include one or more memory units. For example, the memory 7 can include solid-state memory. For example, the memory 7 can be used to store an operating system, applications, and application data.

[0066] For example, the receiver 3 and transmitter 4 may use one or more wireless communication technologies, such as Wi-Fi (IEEE 802.11), to communicate with the wireless LAN access point 23. In alternative embodiments, multiple receivers and / or multiple transmitters are used instead of a single receiver and a single transmitter. Figure 1 In the embodiment shown in , a separate receiver and a separate transmitter are used. In an alternative embodiment, the receiver 3 and the transmitter 4 are combined into a transceiver. For example, the display 9 may include an LCD or OLED panel. The mobile device 1 may include other typical components of a mobile device, such as a battery and a power connector. The present invention may be implemented using a computer program running on one or more processors.

[0067] exist Figure 1 In an embodiment, the system of the present invention is a mobile device. In alternative embodiments, the system of the present invention is a different device (e.g., a PC or a video module), or includes multiple devices. For example, the video module can be a dedicated HDMI module that can be placed between a TV and a device that provides an HDMI input so that it can analyze the HDMI input.

[0068] exist Figure 1 In the embodiments described above, the system of the present invention is used within a lighting system to illustrate that the system can be used both to create light scripts and to render light effects in real time. However, the system need not be part of a lighting system. For example, the system could be a PC used solely for creating light scripts. In this case, light effects are typically not created for specific light sources. Light effects can be created for one or more light sources in a particular part of a room (e.g., to the left of a TV), or for any light source.

[0069] exist Figure 1In an embodiment, the light sources in the lighting system can be used to render light effects in real time during normal use of the lighting system, or can be used to test light scripts. If the system of the present invention is not used in a lighting system, light scripts can also be tested. In this case, one or more light sources can be virtual / simulated. Bridging and communication between devices can also be simulated. In addition, the presentation of media content does not require a TV. For example, the media content can be rendered on a PC used to create a light script (for example for testing purposes). For example, the PC might be running software like Adobe Premier, and the user might be given: an additional window showing a virtual environment with lights; or even a simpler representation to show how the effect will look if the parameters are adjusted in a certain way.

[0070] Figure 2 A first embodiment of the method is shown in FIG. The method is for determining one or more light effects to be presented when media content is being presented. The one or more light effects are determined based on an analysis of the media content. Figure 2 In an embodiment, the one or more light effects include a plurality of light effects. Step 101 includes obtaining media content information. The media content information includes media content and / or information determined by analyzing the media content.

[0071] Steps 103 and 109 include obtaining information indicating the degree of speech in the audio portion. The degree of speech is determined based on an analysis of the audio portion of the media content. Steps 107 and 113 include determining the degree to which the audio portion should be used to determine one or more light effects. The degree is determined based on the degree of speech determined in steps 103 and 109.

[0072] exist Figure 2 In the embodiment of , step 103 includes sub-steps 141 and 143. Step 141 includes determining the amount of speech in the audio portion. Figure 2 In an embodiment of the present invention, this is achieved by spectrally analyzing the audio portion, focusing on the typical frequency region of human speech (i.e., from about 300 to 3400 Hz). Speech detection can be further enhanced by, for example, detecting subtitles in the content, or by focusing on the center channel in or obtained from the audio portion. The audio portion including the center channel is typically presented in a surround sound setting. Additionally, online subtitle libraries can contain timestamps of scenes containing speech, and this information can be used to further optimize speech detection.

[0073] Step 143 includes classifying the audio portion as primarily speech or primarily non-speech based on the amount of speech by determining whether speech is present in more than 50% of the audio portion. Next, step 105 is performed. Step 105 includes determining whether the audio portion has been classified as primarily speech or primarily non-speech. If the audio portion has been classified as primarily speech, step 151 is performed. If the audio portion has been classified as primarily non-speech, step 153 is performed. Steps 151 and 153 are substeps of step 107.

[0074] Step 151 includes determining a first degree. The first degree indicates that the brightness and / or chromaticity of one or more light effects should not be determined based on the intensity and / or loudness of the audio portion, and that the one or more light effects should use a first brightness and / or chromaticity range. Step 109 is performed after step 151. Step 153 includes determining a second degree. The second degree indicates that the brightness and / or chromaticity of one or more light effects should be determined based on the intensity and / or loudness of the audio portion, and that the one or more light effects should use a second brightness and / or chromaticity range. The first brightness and / or chromaticity range has a lower average brightness and / or chromaticity than the second brightness and / or chromaticity range. Step 115 is performed after step 153.

[0075] Step 109 includes classifying the audio portion into one of a plurality of categories. The plurality of categories includes at least two of the following: talking, whispering, screaming, narrating, and singing. Figure 2 In one embodiment, audio segments are classified by analyzing their spectral composition. Differences in spectral composition are then used to determine appropriate behavior for the dynamic lighting system. By considering the spectral and intensity differences between casual speech and shouted speech, it is possible to determine whether people are talking at a conversational level or screaming. This results in a lighting system that can support and enhance content in a manner consistent with its meaning and semantics.

[0076] Next, step 111 includes determining in which category the audio portion has been classified, and steps 161 and 162 include determining the speed of transitions between the plurality of light effects based on this category. If the audio portion is classified as conversation or whispering (Group 1), step 161 is performed. If the audio portion is classified as screaming (Group 2), step 163 is performed. If the audio portion is classified differently (Group 3), the degree determined in step 151 is not modified. In this case, step 115 is performed after step 111. As indicated by the degree determined in step 161, a scene that includes a lot of conversation or a mother whispering to her baby is rendered with low dynamics, while the same scene that includes a lot of screaming or a couple having a shouting argument - even though the audio portions of this scene may have the same intensity and / or loudness - is rendered with higher dynamics, as indicated by the degree determined in step 163.

[0077] After the degree has been determined—that is, one of steps 151 and 153 has been performed, and one of steps 161 and 163 has been conditionally performed—step 115 is performed. Step 115 includes analyzing the video portion of the media content, for example by performing color extraction, and analyzing the audio portion of the media content if step 153 has been performed.

[0078] Therefore, the result of step 143 is that either 1) the audio is primarily speech or 2) the audio is primarily non-speech. Based on this classification, a first level of dynamic adjustment of the light effects is performed in steps 151 and 153. Generally speaking, scenes focused on dialogue should result in lower intensity light effects than scenes focused on visual aspects (otherwise the light effects may actually distract from the dialogue). In addition, the dynamics of the audio signal for speech should not be considered as an input for modulating the intensity of the light effects, which is likely to be more appropriate for non-speech. If it is determined in step 105 that the audio portion has been classified as speech, the spectral content is further analyzed and classified into multiple categories (e.g., conversation, whispering, and screaming) in step 109. Based on this classification, the dynamics of the system are further adjusted in steps 161 and 163.

[0079] Step 117 includes determining one or more light effects to be presented on one or more light sources while the media content is being presented. If step 153 has been performed, the one or more light effects are determined based on the analysis of the audio portion performed in step 115, but they are at least determined based on the analysis of the video portion performed in step 115. Step 119 includes controlling the one or more light sources to present the one or more light effects. Step 121 includes outputting a light script that specifies the one or more light effects.

[0080] In this way, the method optimizes the behavior of a dynamic lighting system based on spectral analysis of the audio content. Low-level spectral analysis allows the identification of speech characteristics such as "regular" conversation, whispers, screams, etc. The system then uses and applies this information to adaptively change the dynamics of the light to correspond to the scene content. Thus, the system enhances media content by adjusting the light in a meaningful way, corresponding to the semantics of the content.

[0081] Figure 3 A second embodiment of the method is shown in FIG. Figure 3 In the embodiment, Figure 2 Step 101 has been replaced with step 201, Figure 2 Step 103 has been replaced by step 203, and Figure 2 Step 109 has been replaced by step 209. Step 201 differs from step 101 in that not only the media content itself is obtained, but also metadata associated with the media content. Similar to steps 103 and 109, steps 203 and 209 include obtaining information indicating the degree of speech in the audio portion. However, in steps 203 and 209, this information is not obtained by analyzing the media content, but from metadata. The metadata may include one or more classifications and / or speech volume and / or spectral analysis information for each time interval of the media content.

[0082] exist Figure 3 In an embodiment, step 203 includes determining from the metadata whether the (current) audio portion is primarily speech or primarily non-speech. Step 209 includes determining from the metadata whether the (current) audio portion belongs to one or more of a plurality of categories, the plurality of categories including at least two of the following: conversation, whispering, screaming, narration, and singing. The audio portion may also be classified as a non-speech category, such as music or natural sounds.

[0083] Figure 4 A third embodiment of the method is shown in FIG. Figure 4 In the embodiment, Figure 3 Step 201 has been replaced by step 301, Figure 3 Step 217 has been replaced by step 317, and Figure 3 Step 115 has been omitted. Step 301 is different from step 201 in that the media content itself is no longer obtained, but only metadata related to the media content is obtained. Figure 3 In addition to the information described, the metadata further includes information extracted from the video and audio portions of the media content that allows the light effect to be determined, such as color extracted from the frames of the video portion or loudness / intensity information extracted from the audio portion. Since it is no longer necessary to analyze the media content to obtain this information, step 115 is omitted. Step 317 is similar to Figure 3 In step 217 , the information obtained in step 301 is used to determine one or more light effects and one or more further light effects.

[0084] Figure 5 A fourth embodiment of the method is shown in FIG. Figure 5 In the embodiment, Figure 2 Steps 103, 105, 107, 109, 111 and 113 have been replaced with steps 401, 403 and 405. Figure 2 Step 103, Figure 5 Step 401 includes step 141, but step 401 does not include Figure 2 Thus, step 401 does not include classifying the speech as being primarily speech or primarily non-speech. Step 141 includes determining the amount of speech in the audio portion, for example using spectral analysis.

[0085] Step 403 involves determining whether the amount of speech determined in step 141 exceeds a threshold. For example, this threshold can be a percentage. If this threshold is set to 50%, this results in a determination of whether the audio portion comprises primarily speech or primarily non-speech. However, the threshold can advantageously be set to a percentage lower or higher than 50%.

[0086] Step 405 is performed after step 403. Step 405 includes sub-steps 407 and 409. If it is determined in step 403 that the threshold has been exceeded, step 407 is performed. If it is determined in step 403 that the threshold has not been exceeded, step 409 is performed. Step 407 includes determining a first degree. Step 409 includes determining a second degree.

[0087] The first degree indicates a first transition speed (i.e., a first dynamics) between the plurality of light effects. The second degree indicates a second transition speed between the plurality of light effects. The second transition speed is higher than the first transition speed. Thus, a light effect accompanying a scene containing more than a certain amount of speech is rendered with low dynamics, while a light effect accompanying the same scene containing less than the certain amount of speech—even though the audio portion of the scene may have the same intensity and / or loudness—is rendered with higher dynamics.

[0088] Figure 6 A fifth embodiment of the method is shown in FIG. Figure 6 In the embodiment, Figure 2Steps 109, 111, and 113 have been replaced with steps 421, 427, 429, and 431. In this fifth embodiment, not only is the spectral content considered, but a semantic analysis of the speech is also performed. Step 421 is performed after step 151, which is performed if the audio portion is classified as primarily speech. In step 421, spoken words are obtained. Step 423 includes determining the spoken words in the audio portion by identifying the spoken words in the audio portion. Step 425 includes obtaining the spoken words from subtitles associated with the media content. In an alternative embodiment, only one of steps 423 and 425 is performed.

[0089] In step 427, the mood of the scene is determined based on the spoken word determined in step 421. In step 429, it is determined whether the mood of the scene is emotional. If the mood of the scene is emotional, a higher transition speed between the plurality of light effects is selected as the degree in step 433. If the mood of the scene is not emotional, a lower transition speed between the plurality of light effects is selected as the degree in step 435. Steps 433 and 435 are sub-steps of step 431.

[0090] Figure 7 A sixth embodiment of the method is shown in FIG. Figure 7 In the embodiment, Figure 2 Step 113 has been replaced with step 451. Step 111 comprises determining whether the audio portion has already been classified as narration or singing or has been classified differently. If the audio portion has already been classified as narration or singing (group 5), step 451 is performed. Step 153 is performed as a sub-step of step 451. Thus, the degree is determined as if the audio portion were classified as primarily non-speech and normal light effects were applied. If the audio portion has already been classified differently—for example, as conversation or screaming (group 4)—then the degree is not modified and step 115 is performed next.

[0091] Figure 8 An example of audio classification of a first media content item, which is an episode of a TV series, is shown in the form of a graph. Time is plotted along the x-axis of the graph. Four possible categories are shown along the y-axis of the graph. Figure 8 In the audio classification depicted in , an audio portion with a duration of one second is classified. The chart shows the categories detected within a 30-second period. From one to six seconds, the music category 53 is detected. From seven to fourteen seconds, the conversation category 57 is detected. From fifteen to twenty seconds, the screaming category 55 is detected. From twenty-one to thirty seconds, the conversation category 57 is again detected. The singing category 51 is not detected in this audio portion. Based on these classifications, the time interval from zero to thirty seconds can be classified as primarily speech, since screaming and conversation are speech categories.

[0092] Although Figure 8 In the example, only one category is detected per second, but in Figure 9 In this example, we detect multiple categories at the same time. Figure 9 An example of audio classification for a second media content item, which is a music video clip, is shown in the form of a graph. From 0 to 30 seconds, the music category 53 is detected. From 4 to 10 seconds, 12 to 18 seconds, and 23 to 30 seconds, the singing category 51 is detected. Based on these classifications, the time interval from 0 to 30 seconds can be classified as primarily non-speech because the music category 53 is detected for 30 seconds and the singing category 51 is detected for 22 seconds.

[0093] Figure 10 Depicted instructions can be executed as reference Figures 2 to 7 A block diagram of an exemplary data processing system for the described methods.

[0094] like Figure 10 As shown in , data processing system 500 may include at least one processor 502 coupled to a memory element 504 via a system bus 506. In this manner, the data processing system may store program code within the memory element 504. Further, processor 502 may execute program code accessed from the memory element 504 via the system bus 506. In one aspect, the data processing system may be implemented as a computer suitable for storing and / or executing program code. However, it should be appreciated that the data processing system 500 may be implemented in the form of any system including a processor and a memory capable of performing the functions described in this specification.

[0095] Memory element 504 may include one or more physical memory devices, such as, for example, local memory 508 and one or more mass storage devices 510. Local memory may refer to random access memory or (a plurality of) other non-persistent memory devices typically used during the actual execution of program code. Mass storage devices may be implemented as hard drives or other persistent data storage devices. Processing system 500 may also include one or more cache memories (not shown), which provide temporary storage of at least some program code to reduce the number of times program code must be retrieved from mass storage devices 510 during execution. For example, if processing system 500 is part of a cloud computing platform, processing system 500 may also be able to use the memory element of another processing system.

[0096] Optionally, input / output (I / O) devices, depicted as input device 512 and output device 514, may be coupled to the data processing system. Examples of input devices may include, but are not limited to, a keyboard, a pointing device such as a mouse, a microphone (e.g., for voice and / or speech recognition), etc. Examples of output devices may include, but are not limited to, a monitor or display, speakers, etc. Input and / or output devices may be coupled to the data processing system directly or through intervening I / O controllers.

[0097] In an embodiment, the input and output devices may be implemented as a combined input / output device (in Figure 10 512 and output device 514). An example of such a combined device is a touch-sensitive display, sometimes also referred to as a "touch screen display" or simply a "touch screen." In such an embodiment, input to the device can be provided by movement of a physical object (such as, for example, a user's finger or a stylus) on or near the touch screen display.

[0098] Network adapter 516 may also be coupled to the data processing system to enable it to couple to other systems, computer systems, remote network devices, and / or remote storage devices through intervening private or public networks. The network adapter may include a data receiver for receiving data transmitted to the data processing system 500 by the system, device, and / or network, as well as a data transmitter for transmitting data from the data processing system 500 to the system, device, and / or network. Modems, cable modems, and Ethernet cards are examples of different types of network adapters that may be used with data processing system 500.

[0099] like Figure 10 As shown in FIG, memory element 504 can store application programs 518. In various embodiments, application programs 518 can be stored in local memory 508, one or more mass storage devices 510, or separate from local memory and mass storage devices. It should be appreciated that data processing system 500 can further execute an operating system (OS) that can facilitate the execution of application programs 518. Figure 10 ). Application 518, implemented in the form of executable program code, may be executed by data processing system 500 (e.g., by processor 502). In response to executing the application, data processing system 500 may be configured to perform one or more operations or method steps described herein.

[0100] Various embodiments of the present invention may be implemented as a program product for use with a computer system, wherein the program(s) of the program product define the functionality of the embodiments (including the methods described herein). In one embodiment, the program(s) may be contained on various non-transitory computer-readable storage media, wherein, as used herein, the expression "non-transitory computer-readable storage media" includes all computer-readable media with the sole exception of temporary propagation signals. In other embodiments, the program(s) may be contained on various temporary computer-readable storage media. Illustrative computer-readable storage media include, but are not limited to: (i) non-writable storage media on which information is permanently stored (e.g., a read-only memory device within a computer, such as a CD-ROM disk readable by a CD-ROM drive, a ROM chip, or any type of solid-state non-volatile semiconductor memory); and (ii) writable storage media on which changeable information is stored (e.g., a flash memory, a floppy disk within a floppy disk drive or hard drive, or any type of solid-state random access semiconductor memory). The computer program may be executed on the processor 502 described herein.

[0101] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the present invention. As used herein, the singular forms "a" or "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "comprise" and / or "comprising" specify the presence of the features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0102] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements in the claims below are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The descriptions of the embodiments of the present invention have been presented for illustrative purposes and are not intended to be exhaustive or limited to the embodiments in the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention. The embodiments are chosen and described in order to best explain the principles of the invention and some practical applications, and to enable others of ordinary skill in the art to understand the invention for various embodiments with various modifications suitable for the particular use under consideration.

Claims

1. A system (1) for determining one or more light effects to be presented when media content is being presented, the one or more light effects being determined based on an analysis of the media content, the system (1) comprising: - at least one input interface (3); - at least one output interface (4); and - at least one processor (5) configured to: - using said at least one input interface (3) to obtain media content, - determining one or more light effects to be presented on one or more light sources (13-17) while the media content is being presented, the one or more light effects being determined based on: - analysis of the audio portion of said media content, and - analysis of the video portion of said media content, and - using the at least one output interface (4) to control the one or more light sources (13-17) to present the one or more light effects, wherein the processor (5) is further configured to: - obtaining information indicative of a degree of speech in the audio portion, the degree of speech being determined based on the analysis of the audio portion; - determining that the audio portion should be used to determine an extent to which the one or more light effects are to be used, the extent being determined based on the determined degree of speech; as well as - determining said determined extent of said one or more light effects in dependence on said audio portion which should be used, the brightness and / or chromaticity of said one or more light effects being determined based on the intensity and / or loudness of said audio portion.

2. The system (1) of claim 1, wherein the degree of speech in the audio portion is determined by determining an amount of speech in the audio portion and classifying the audio portion as primarily speech or primarily non-speech based on the amount of speech.

3. A system (1) as claimed in claim 2, wherein the at least one processor (5) is configured to determine a first degree as the degree based on the audio portion being classified as primarily speech and to determine a second degree as the degree based on the audio portion being classified as primarily non-speech, the second degree indicating that the brightness and / or chromaticity of the one or more light effects should be determined based on the intensity and / or loudness of the audio portion, and the first degree indicating that the brightness and / or chromaticity of the one or more light effects should not be determined based on the intensity and / or loudness of the audio portion.

4. The system (1) of claim 2, wherein the at least one processor (5) is configured to determine the one or more light effects using a first luminance and / or chrominance range based on the audio portion being classified as primarily speech and using a second luminance and / or chrominance range based on the audio portion being classified as primarily non-speech, the first luminance and / or chrominance range having a lower average luminance and / or chrominance than the second luminance and / or chrominance range.

5. The system (1) of claim 1, wherein the degree of speech in the audio portion is determined by classifying the audio portion into one of a plurality of categories (51, 53, 55, 57), the plurality of categories (51, 53, 55, 57) comprising at least two of the following: conversation (57), whispering, screaming (55), narration, singing (51).

6. A system (1) as claimed in claim 5, wherein the at least one processor (5) is configured to determine a first degree as the degree based on the audio portion being classified as conversation and to determine a second degree as the degree based on the audio portion being classified as singing, the second degree indicating that the brightness and / or chromaticity of the one or more light effects should be determined based on the intensity and / or loudness of the audio portion, and the first degree indicating that the brightness and / or chromaticity of the one or more light effects should not be determined based on the intensity and / or loudness of the audio portion.

7. The system (1) of claim 5, wherein the one or more light effects comprises a plurality of light effects, and the at least one processor (5) is configured to determine a transition speed between the plurality of light effects according to the category.

8. The system (1) of claim 5, wherein the audio portions are classified by analyzing the spectral composition of the audio portions.

9. The system (1) of claim 1, wherein the degree of speech in the audio portion is determined by classifying the audio portion as diegetic speech or non-diegetic speech.

10. The system (1) of claim 1, wherein the one or more light effects include a plurality of light effects, and the at least one processor (5) is configured to determine whether the amount of speech in the audio portion exceeds a threshold and determine a transition speed between the plurality of light effects based on the amount of speech exceeding the threshold.

11. The system (1) of claim 1, wherein the at least one processor (5) is configured to determine the words spoken in the audio portion by identifying the spoken words in the audio portion and / or obtaining the spoken words from subtitles associated with the media content.

12. The system (1) of claim 1, wherein the at least one processor (5) is configured to determine the degree of speech by using subtitles associated with the media content and / or by focusing on a center channel in or obtained from the audio portion.

13. A lighting system (21) comprising the system (1) according to any one of claims 1 to 12 and one or more light sources (13-17).

14. A method of determining one or more light effects to be presented while media content is being presented, the one or more light effects being determined based on an analysis of the media content, the method comprising: -Get (101, 201, 301) media content; - determining (117, 317) one or more light effects to be presented on one or more light sources while the media content is being presented, the one or more light effects being determined based on an analysis of an audio portion of the media content and an analysis of a video portion of the media content; as well as - controlling (119) said one or more light sources to render said one or more light effects, wherein the method further comprises: - obtaining (103, 109, 203, 209, 401, 421) information indicative of a degree of speech in the audio portion, the degree of speech being determined based on an analysis of the audio portion; - determining (107, 113, 405, 431, 451) that the audio portion should be used to determine a degree of the one or more light effects, the degree being determined based on the determined degree of speech; as well as Wherein the determined extent to which the one or more light effects should be used is determined according to the audio portion, the brightness and / or chromaticity of the one or more light effects being determined based on the intensity and / or loudness of the audio portion.

15. A computer program product comprising at least one software code portion configured to enable the method of claim 14 to be performed when run on a computer system.

Citation Information

Patent Citations

  • Audiovisual environment control device, audiovisual environment control system, and audiovisual environment control method

    WO2007119277A1

  • Multimedia playing system with lighting display device

    CN101388234A

  • Light adjusting method and device, intelligent illumination equipment and storage medium

    CN107509287A