Predicting physiological responses of people to media content

By correlating prosodic properties with physiological data through Interpretation Maps, the method predicts and tailors media content for desired physiological responses, addressing the limitations of qualitative music therapies and medication side effects, enabling precise and targeted media selection.

WO2026008967A1PCT designated stage Publication Date: 2026-01-08KINGS COLLEGE LONDON
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2025/051432
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-05
Filing Date
2025-06-30
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing music-based therapies for cardiovascular health are qualitative and difficult to model quantitatively, limiting their application at scale, and medication for hypertension often has negative side effects with poor compliance, necessitating a non-pharmacological, individualized, and pleasurable alternative to manage physiological responses.

Method used

Measuring physiological responses to media content, generating predictive models using Interpretation Maps that correlate prosodic properties with physiological data, and creating an index to predict responses to other media content, allowing for tailored selection based on desired physiological outcomes.

Benefits of technology

Enables precise analysis and prediction of physiological responses to media content, facilitating therapeutic and entertainment applications by providing a library of physiologically classified media content for targeted physiological effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025051432_08012026_PF_FP_ABST
    Figure GB2025051432_08012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to measuring physiological responses of consumers of media content, and then generating models of the consumers' responses. Such models can subsequently be used to make predictions of a viewer / listener / musician response to other media content. Some embodiments receive physiological measurements of subjects exposed to media content and generate predictive models by processing the received measurements together with labels describing the structure of the media content. In this way, the impact of the structure on the consumer's physiological state can be determined. These results can be used to build an index that allows physiological responses to other items of media content to be predicted.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PREDICTING PHYSIOLOGICAL RESPONSES OF PEOPLE TO MEDIA CONTENT Field The present disclosure concerns methods, computing systems, computer programs and computer readable media for predicting physiological responsesof people consuming media content.BackgroundIt is known that music can impact the autonomic nervous system (ANS), whichregulates unconscious bodily processes like heart rate, heart rate variability,breathing, thermoregulation, digestion and others. As a result, music-basedtherapies have been successfully used in many aspects of healthcare, including cardiac healthcare (e.g., Hanser SB, Mandel SE (2005). The effects of music therapy in cardiac healthcare. Cardiol Rev. 13(1):18-23; and Hanser SB (2014). Music therapy in cardiac health care: current issues in research. Cardiol Rev. 22(1): 37-42). However, the predominant approach has been largely qualitative, and relies strongly on one-to-one sessions with trained therapists, making the effects difficult to model quantitatively and apply at scale. Much of therapy focuses on mental and psychological effects of music, but a growing body of research shows that music has the potential to affect physiology, and in particular the heart, in tangible ways (Chew E, P Loui, G Leslie, C Palmer, J Berger, E Large, N Bernardi, S Hanser, J Thayer, M Casey,P Lambiase (2021). How Music Can Literally Heal the Heart. Scientific American(Opinion), 18 Sep 2021 online, Dec 2021 / Jan 2022 6(3):28-30 print, www.scientificamerican.com / article / how-music-can-literally-heal-the-heart). Most studies of music-induced physiological effects consider music as a single class of sounds (comparing physiology with and without music) and some go a step further to consider different genres of music (Bernardi L, C Porta, P Sleight. (2005) Cardiovascular, cerebrovascular, and respiratory changes induced by different types of music in musicians and non-musicians: the importance of silence. Heart 92, 445–452; and Hilz MJ, S Peter, G Thomas, N Juliane, L Habib-Romstoeck, B Stemper, S Buechner, S Wong, J Koehn (2014). Music induces different cardiac autonomic arousal effects in young and older persons. Auton. Neurosci. 183, 83–93); a few consider general features of individual compositions like the complexity of the rhythm (Bernardi L, C Porta, P Sleight. (2005) Cardiovascular, cerebrovascular, and respiratory changesinduced by different types of music in musicians and non-musicians: theimportance of silence. Heart 92, 445–452), difficulty of the music (Wright SE, C Palmer (2020). Physiological and behavioral factors in musicians’performance tempo. Frontiers in Human Neuroscience. 2020 Aug 25;14:311;and Mulcahy D, J Keegan, A Fingret, C Wright, A Park, J Sparrow, D Curcher,KM Fox (1990). Circadian variation of heart rate is affected by environment: astudy of continuous electrocardiographic monitoring in members of a symphony orchestra. Heart 64, 388–392), and tempo (Bernardi L, C Porta, P Sleight. (2005) Cardiovascular, cerebrovascular, and respiratory changes induced by different types of music in musicians and non-musicians: the importance of silence. Heart 92, 445–452; and Nomura S, K Yoshimura, Y Kurosawa (2013). A pilot study on the effect of music-heart beat feedback system on human heart activity. J. Med. Informatics Technol. 22), vocal and orchestral crescendos, phrases and emphasis (Bernardi L, C Porta, G Casucci, R Balsamo, NF Bernardi, R Fogari, P Sleight (2009). Dynamic Interactions Between Musical, Cardiovascular, and Cerebral Rhythms in Humans. Circulation 119, 3171–3180). Some studies and meta-analyses suggest the benefits of using music as a non- pharmacological approach to support the therapy of cardiovascular diseasesdriven by the dysfunction of the ANS, such as hypertension, which afflicts 1 in3 adults worldwide (do Amaral MA, Neto MG, de Queiroz JG, Martins-Filho PR,Saquetto MB, Carvalho VO. Effect of music therapy on blood pressure of individuals with hypertension: A systematic review and Meta-analysis. International journal of cardiology. 2016 Jul 1;214:461-4 and Teng XF, Wong MY, Zhang YT. The effect of music on hypertensive patients. In2007 29th Annual International Conference of the IEEE Engineering in Medicine andBiology Society 2007 Aug 22 (pp. 4649-4651). IEEE. and Cao M, Zhang Z.Adjuvant music therapy for patients with hypertension: a meta-analysis and systematic review. BMC complementary medicine and therapies. 2023 Apr 6;23(1):110). Medication for hypertension, while effective, can have negative side effects. Drug compliance is usually poor, leading the therapy to be ineffective, and some patients’ bodies are resistant to the medication.Although the general effect of music on blood pressure levels is known andreported in the literature, drug compliance may be poor. Therefore, improving knowledge of how media may affect people's physiological responses may facilitate and optimise a non-pharmacological, non-invasive, individualised, and pleasurable alternative to hypertension medication to lower blood pressureor improve cardiovascular conditions with few (if any) side effects. Moreover,an improved understanding of how people respond physiologically to mediacontent can be used to tailor media content to achieve other physiological responses in non-medical contexts. Therefore, it is an object of the present disclosure to improve understanding of physiological responses to media content. SummaryThe present disclosure relates to measuring physiological responses ofconsumers of media content, and then generating models of the consumers’responses. Such models can subsequently be used to make predictions of aviewer / listener / musician / performer response to other media content. Whilethe present disclosure primarily focuses on physiological responses to music, it will be recognised that various forms of media (e.g., film, television, etc.)include visual content or audible content that can change over time in waysthat affect the consumer physiologically, and so the methods of the presentdisclosure can be applied to any type of media content. Some embodimentsreceive physiological measurements of subjects exposed to media content and generate predictive models by processing the received measurements together with labels describing the structure of the media content. In this way, the impact of the structure on the consumer’s physiological state can be determined. These results can be used to build an index that allows physiological responses to other items of media content to be predicted. Some embodiments evaluate musicians’ performance decisions (i.e., their individual interpretation) as mathematical inputs into computational models for predicting any listeners’ (including the musicians’ own) autonomic response. For example, in some embodiments, physiological responses are mapped to items of media content by defining features of the media contentin an “Interpretation Map”. An Interpretation Map may capture expressivestructures and interpretive decisions made by musicians when shaping musical communication. These may be described as prosodic properties of media content and may be associated with respective prosodic labels (which could identify where in an item of media content a respective prosodic property is located). These decisions can include moments of climax and release, melodic interest, interaction or dialogue, significant accompaniment, and moments of concern. The Interpretation Map can be extended by annotations related to the structure of the piece, including the labels associated with, for example, introducing a novel melody, returns to the previous melodies, significant pauses, and others.Interpretation Maps can be marked by musician(s), a viewer, a listener, orautomatically extracted (e.g., by calculating change points for loudness,timbre-related signals like Mel-Frequency Cepstral Coefficients (MFCC), ortempo). Interpretation Maps may be mathematical representations that can beintegrated into computational models of consumers’ physiological responses.For example, an Interpretation Map may be stored in a digital structure that relates specific features of media content to specific points in time. The Interpretation Map may be used alongside other video and / or musical features,such as tempo, loudness, and note density, to improve descriptive andpredictive models of physiological reactions to media content. Once a physiological predictive model has been generated, it can then be used to predict the physiological response of a consumer (e.g.,musician / listener / viewer / player) to other media content.Some embodiments of the present disclosure make a quantitative prediction of the physiological response that will be generated in a musician or other listener by a given piece of music. This has advantages in that it allows for thecreation of transparent, explanatory models that link specific music structures resulting from interpretive decisions to particular physiological responses and can identify moments in a piece of music that are most influential in eliciting physiological reactions, offering a more precise analysis than general continuous features like loudness or tempo. Moreover, once it is possible to predict the physiological response of aconsumer to an item of media content, or to an individual property within anitem of media content, it can then become possible to predict physiological responses to other media content (e.g., to a plurality of other, known piecesof music, or to other videos or films) to select one of the items of media contentthat will give an intended physiological response. For example, this can allowfor quantitative selection, based on predicted physiological response, of a piece of music to match the intended “mood” that a director (for example) of a film is attempting to invoke in a watcher of a film through the choice of music.Intended physiological responses could be used for therapeutic purposes (e.g.,for music-related interventions), for example, to lower blood pressure, or for entertainment purposes (for example to increase the tension in a movie), or for other purposes (e.g. to elicit the same response as a third person formatchmaking, for example). Such applications can be facilitated by being ableto systematically evaluate media content and thereby quantify / predict physiological responses to such media content. Some embodiments of the present disclosure use generated predictive modelsto make predictions of physiological responses of consumers to other pieces ofmedia content, which allows a library of physiologically classified items ofmedia content to be generated, where the physiological response has been predicted. Once such a library has been built up, the library can be accessed to select items of media content that give desired physiological responses. Thiscould be for any one of a number of purposes, including therapeutic (e.g., tolower blood pressure), entertainment purposes (e.g., to select exciting musicor visual effects for a movie), or for any other purposes (e.g., to identify groupsof people who have a similar physiological responses to media content).Accordingly, against this background and in accordance with an aspect of thepresent disclosure, there is provided a computer-implemented method according to claim 1. In further aspects, a computing system according to claim 23 is provided, a computer program product according to claim 24 is provided, and a computer readable medium according to claim 25 is provided.Specifically, the present disclosure provides a computer-implemented methodfor predicting physiological responses of people consuming media content, anda computing system comprising a processor configured to perform the method,the method comprising: recording one or more physiological responses of aperson as the person consumes an item of media content, the item of media content being associated with a plurality of prosodic labels defining prosodic properties of the item of media content; processing the recorded one or morephysiological responses in conjunction with the prosodic labels to generate apredictive model predicting physiological responses of people to media content; applying the predictive model to a plurality of further items of mediacontent to generate further predicted physiological responses to the respectivefurther items of media content; and storing an index comprising the further predicted physiological responses indexed to the respective further items of media content. Advantageously, a predictive model may be used to create one or more indexes of physiological responses caused by various items of media content. Thiscould be used for specific sub-groups of population (differentiated by, forexample, demographics, health status, music sophistication, music preferences, and the like). Hence, a library of media content is provided, which is physiological-response-classified. Then, when a specific indexedphysiological response is desired (e.g., as indicated by a selection provided bya user input), pieces of media content can be selected depending on the intended use for the media content. Accordingly, some embodiments of the present disclosure provide a method further comprising: receiving data (e.g., based on a user selection) indicative of a desired physiological response (e.g., a physiological response that the userwishes to elicit from a person); and selecting (e.g., by searching within theindex for the desired physiological response) one or more of the further items of media content using the index, the selected one or more further items ofmedia content having predicted physiological responses that correspond (e.g., are the same as or are similar to) to the desired physiological response as indicated by the received data. For example, using this approach, media content that can provide a desired physiological response may be provided in response to a query. In some embodiments, the method may further comprise providing, in response to receiving the data indicative of the desired physiological response, the selected one or more further items of media content. The actual item ofmedia content may be provided in response to the data indicative of thedesired physiological response. In some embodiments, an indication of the relevant item of media content may be provided instead of the actual mediacontent per se. For example, where bandwidth is limited, the method mayreturn an indication (e.g., a title of, or an identifier of) of the selected one or more further items of media content. In some embodiments, in the methods described herein, the desired physiological response and the predicted physiological responses of each of the selected one or more further items of media content may have a similarity metric that exceeds a threshold value. As a simplified example, if a desiredphysiological response is a 1% reduction in blood pressure, and the indexindicates that an item of media content is likely to cause a 0.9% reduction in blood pressure, then the method may determine that 0.9% is close enough to 1% to return the relevant item of media content. Therefore, some similarity metric (e.g., a percentage difference threshold, or a percentage point difference threshold, or some other appropriate metric for quantifying similarity) may be determined. Where a threshold is used, the threshold could be arbitrary and / or user-defined and / or adjustable. In some embodiments, at least one item of media content, and optionally each item of media content, may comprise any one or more of: an audio file; a pieceof music; a video file; a television programme; a video game; and / or a film.The above-mentioned media content could, in some embodiments, be generated automatically by artificial intelligence (AI). Any media content with audible sound can be used with and can benefit from the predictive models described herein. The person, or people, used in embodiments of the disclosure may be any oneof: an audience of the media content (e.g., listener, watcher, viewer); aparticipant of the media content; and a player of music in the media content(e.g., a player who is currently playing the music in the media content). Forexample, musicians’ physiological responses when playing music can be correlated with either or both of: the physiological responses of consumers of the music; and the physiological responses of other performers when they play the music. Accordingly, so the predictive models described herein can be used to predict physiological responses of various different people. Moreover, the physiological response of a person who listens to a musical song might be the same as, or at least similar to, the physiological response of a person who hears the same song while watching a movie. Therefore, various different people’s physiological responses can be predicted using the methods described herein. In some embodiments, the one or more physiological responses may comprise any one or more of: a cardiovascular measurement; heart rate; beat-to-beatheart intervals (RR intervals); heart rate variability parameters (HRV); bloodpressure; and respiration rate. Other physiological measurements can be takenand predicted using the methods of the present disclosure. The advantages noted above and other advantages of the present disclosure will be apparent from the following detailed description. Listing of Figures Embodiments of the present disclosure will now be described by way of example and with reference to the following figures, in which: Figure 1 shows a method 100 according to an embodiment of thepresent disclosure; Figure 2 shows a method 200 according to an embodiment of thepresent disclosure; Figure 3 shows an index generated according to an embodiment of thepresent disclosure; Figure 4 shows an example computing system 400 an embodiment of the present disclosure, for implementing the methods described herein; Figure 5 shows changes in systolic and diastolic blood pressure (BP) andPulse Pressure (PP) in response to annotated expressive music structures andprosody parameters; Figure 6 shows mean difference values of physiological parameterscalculated for each category and physiological signal; p-value annotations: *p<0.05, ** p<0.01, ***p<0.001; the significant differences remaining after using Bonferroni correction were marked by a black thick frame; Figure 7 shows mean values and 95%CI (Confidence Interval) foranalysed parameters in each second before and after the onsets of transitions; Figure 8 shows data collection in two settings: 1) performers setting,with Polar straps (a) to collect ECG (electrocardiogram) data and a Zoomrecorder gathering audio data from each performance (b); 2) listener settingwith Polar strap and other sensors to collect ECG, RR intervals, respiration signals, BP, BP waveform; Figure 9 shows an Interpretation Map; Figure 10 shows the original and predicted RR interval series from anexample performance, together with the musicians’ annotations and loudness and tempo signals in the score-time domain for all musicians; and Figure 11 shows an example of a binary and Gaussian signal,demonstrating how sequential annotated sections of moment of concern are summed. Detailed Description Figure 1 shows a method 100 according to an embodiment of the present disclosure. The method 100 comprises recording 101 one or more (e.g., oneor a plurality of) physiological responses (e.g., as measured using any one ormore of the following sensors: heart rate sensors, ECGs, blood pressuresensors, respiratory sensors, etc.) of a person as the person consumes an itemof media content (e.g., while the person is exposed to the content, measurements may be taken on that person), the item of media content being associated with a plurality of prosodic labels (e.g., as defined in an Interpretation Map) defining prosodic properties (e.g., items of melodic interest) of the item of media content.The method 100 further comprises processing 102 the recorded one or morephysiological responses in conjunction with the prosodic labels to generate apredictive model predicting physiological responses of people to media content. For example, the one or more physiological responses and the prosodic labels may be received as inputs, and associations between them may be made. For instance, temporal relationships may be inferred between times at which physiological responses occur and where in the music certain prosodic properties can be found. Hence, logic (e.g., statistical methods) can be applied to determine a predictive model based on the recorded physiological response(s) and the prosodic labels. After the step of processing 102, the method 100 further comprises applying 103 the predictive model to a plurality of further items of media content (which could be, for example, retrieved from a database, or a training data set) to generate further predicted physiological responses to the further items of media content. In this way, the physiological responses caused by multiple different items of media content can be ascertained.Then, the method 100 comprises storing 104 one or more indexes comprisingthe further predicted physiological responses indexed to the respective furtheritems of media content. Hence, an index can be built that provides a convenientway of identifying items of media content that are likely to elicit certainphysiological responses for a whole population (e.g., if the population isrelatively homogenous) or for a certain sub-group of a population. The index may associate one or a plurality of physiological responses with a particular item of media content. For example, where an item of media content is likely to cause a reduction in blood pressure as well as heart rate, then two physiological responses may be stored in the index that is built. The method 100 may also comprise storing, in the index, the recorded one or more physiological responses indexed to the item of media content (i.e., the initial media content used to measure physiological responses may be stored in the index in addition to the further items of media content stored in the index). The steps of recording one or more physiological responses and processing one or more physiological responses may be performed for each of a plurality of different items of media content to generate the predictive model. Additionally or alternatively, the steps of recording one or more physiological responses and processing one or more physiological responses may be performed for a plurality of different people to generate the predictive model. The physiological responses may be indicated by physiological measurements (e.g., usingphysiological sensors, such as heart rate strap(s) and / or blood pressuremonitor(s) and / or respiratory sensor(s)) taken on one or more people. Thus,various different items of media content and people can be taken into account when determining the relationships between prosodic properties and physiological responses. The step of recording 101 the one or more physiological responses of the person may comprise: exposing the person to the item of media content (e.g., playing the content to them); recording one or more times at which the one ormore physiological responses occurred (e.g., using various physiologicalsensors); optionally collecting additional information (e.g., user data) aboutthe person (demographics, health status, music preferences, musicsophistication, etc.); and associating one or more of the prosodic labels withthe one or more times at which the one or more physiological responses occurred (e.g., by correlating the times at which particular physiological responses are sensed with points of time in the media content, which could beexpressed in score-time or in other arbitrary units of time). Processing therecorded one or more physiological responses in conjunction with the prosodic labels to generate the predictive model may be performed in dependence on the association between the one or more prosodic labels and the one or more times. For example, values measured in standard units of time can be represented using musical notation in score-time, since a decrease in blood pressure occurring at the start of the second bar of a specific song may be more meaningful (and hence more useful for the predictive models described herein) than data indicating that the change in blood pressure occurred at 11:29am on 17 May 2023. The predictive models described herein can predict different physiological responses to media content depending on collected additional information about the person (including demographics, health status, music preferences,music sophistication, etc.). For instance, there may not be a single universalreaction to media content. Different for different sub-groups of a population might exhibit different responses. For example, the methods of some embodiments may comprise recording user data of the person (e.g., before or after the person consumes the item of media content) and determining the predictive model based on the user data. Thus, predictive models can be generated that allow specific user data to be taken into account whenpredicting physiological responses. The user data may comprise any one ormore of: demographic data of the user (e.g., any social and / or socioeconomicfactors, such as age, ethnicity, gender, marital status, education, and / oremployment); an indication of a degree of media content sophistication of theuser (e.g., musical sophistication, indicating the user’s level of knowledge about musical compositions, structures, etc.); and / or one or more media content preferences of the user.Figure 2 shows a method 200 that builds on the method of Figure 1. Themethod 200 comprises: recording 201 one or more (e.g., one or a plurality of)physiological responses of a person as the person consumes an item of mediacontent, the item of media content being associated with a plurality of prosodic labels defining prosodic properties of the item of media content; processing 202 the recorded one or more physiological responses in conjunction with the prosodic labels to generate a predictive model predicting physiological responses of people to media content; applying 203 the predictive model to a plurality of further items of media content to generate further predicted physiological responses to the further items of media content; and storing 204 an index comprising the further predicted physiological responses indexed to the respective further items of media content. In this regard, steps 201-204 are substantially similar to steps 101-104 of the method 100 of Figure 1. Themethod 200 of Figure 2 further comprises steps of: receiving 205 dataindicative of a desired physiological response; and selecting 206 one or more of the further items of media content using the index, the selected one or more further items of media content having predicted physiological responses that correspond to the desired (optionally for a given sub-group of population) physiological response as indicated by the received data. Using the methods of Figures 1 and steps 201-204 of Figure 2, predictive models can be built that index physiological responses of people to specific items of media content. Using steps 205-206 of Figure 2, a particular item of media content can be identified (and optionally provided to a person) when it is desired to elicit a specific physiological response in a person. It will be appreciated that the steps of Figure 2 could be performed on a pre- existing index. For example, some aspects of the present disclosure could provide a method comprising: receiving data indicative of a desired physiological response; and selecting one or more items of media content using an index that comprises physiological responses indexed to respective items of media content, wherein the selected one or more items of media content have predicted physiological responses that correspond to the desired (optionally for a given sub-group of population) physiological response as indicated by the received data. Figure 3 shows an index that can be generated and used in accordance with embodiments of the present disclosure. The index includes identifiers of numerous items of media content (Song1, Film1, TV1, Song2, Song3, Film2),numbers of sub-items in each media content (1, 2, 3, etc.) and associatedphysiological response(s) (A, B, C, D, E, F, G, H) of those items of mediacontent. When a user wishes to identify media content item that will elicitphysiological response “E”, the third column of the index in Figure 3 can be searched for “E” and the identifiers all entries that contain “E” can be located.In this instance, Song1 (items 1 & 4), Film1 (item 1), Song2 (item 1) andSong3 (item 1) would be identified, since each of these items of media contentis indexed to physiological response “E”. The number of sub-items in an item of media content can be any positive integer. For example, an entire item of media content can be analysed as a whole and indexed using the methods described herein, in which case, the number of sub-items is 1. In some cases, an item of music might include a relatively energetic portion and a relatively low-energy portion, in which case the two portions could be designated as sub-items and the number of sub- items would be 2, with the number of two sub-items being determined based on the number of different physiological responses induced by that item of content. In other cases, a video could be considered to have 2 different sub- items, with one sub-item being a stream of visual images that elicit a certain physiological response and another sub-item being an accompanying soundtrack of the video that might elicit the same or a different physiologicalresponse. The visual and / or audio content could be analysed and indexed asseparate sub-items of the same item of media content. If the audio content ofa video is fairly homogenous but the visual content changes regularly, then asingle video file might be associated with multiple visual sub-items but only one audio sub-item. A sub-item of content could be described as an excerpt,an extract or an individual structure within an item of content. The number ofsub-items of an item of content could be determined automatically (e.g., using AI) or could be input by a user (e.g., a user with knowledge of how media is structured). Figure 4 shows an example computing system 400, according to another embodiment of this disclosure, for implementing the methods described in this disclosure. Specifically, Figure 4 is a block diagram illustrating an arrangement of a system 400 for implementing the present disclosure. Some embodiments of the present disclosure are designed to run on general-purpose desktop or laptop computers. Therefore, according to an embodiment, a computing system 400 is provided having a central processing unit (CPU) 402, and random access memory (RAM) 404 into which data, program instructions, and the like can be stored and accessed by the CPU 402. The computing system 400 is provided with a display screen 406, and input peripherals in the form of a keyboard 408, and a mouse 410. The computing system 400 can include an audio output, such as a speaker, headphones, or the computing system 400 can include or be connected to a musical instrument operating based on electronic instructions (e.g., in a MIDI format). The display screen 406 could include an integrated speaker system for outputting audio content. The keyboard 408 and the mouse 410 communicate with the system 400 via a peripheral input interface 412. Similarly, a display controller 414 is provided to control display 416, so as to cause it to display images under the control of CPU 402. Data 418, for example audio files in any audio format (e.g., mp3, .wav, etc.)or any video format (e.g., .mp4, .mov, etc.), can be input into the system 400and stored via the data input 420. In this respect, the system 400 comprises a computer readable storage medium 422, such as a hard disk drive, writable CD or DVD drive, zip drive, solid state drive, USB drive or the like, upon which data 418 can be stored. Alternatively, the data 418 could be stored on a web-based platform, for example, a database, and accessed via an appropriate network. A computer readable storage medium 422 also stores various programs, which when executed by the CPU 402 cause the system 400 to operate in accordance with some embodiments of the present disclosure. In particular, a control interface program 424 is provided, which when executed by the CPU 402 provides overall control of the computing apparatus, and inparticular provides a graphical interface on the display 416, and accepts userinputs using the keyboard 408 and the mouse 410 by the peripheral interface 412.Such a control interface 424 can be used when setting up the predictive modelsdescribed herein, or when providing an input indicating a desired physiologicalresponse when a user wishes to make use of the predictive models describedherein. The control interface program 424 may also call, when appropriate, other programs to perform specific processing actions. The user launches the control interface program 424. The control interface program 424 is loaded into RAM 404 and is executed by the CPU 402. The user then launches a program 426, which acts on the input data 418 as described above. The program instructions 426executed by the CPU 402 relate to the method as described in any of the otherembodiments of this disclosure.Embodiments of the present disclosure, therefore, provide a way of consideringhow the music is communicated to consumers (i.e., through the prosodiclabels) and integrating quantified ways of representing these prosodicstructures. The content that impacts the body can be found in the inherentstructures of media content and in the ways in which media content is communicated, the prosodic structures of the communicated media content, and so the predictive models described herein can help to identify the effects of such structures.Being able to predict cardiac response to media content at a large scale (i.e.,rather than having a clinician manually evaluate individuals on a case-by-case basis) requires having a way to represent performers’ prosodic choices and integrate them in descriptive and predictive models of autonomic activity; autonomic activity may be, as measured, for example, as features derived from electrocardiographic signals such as beat-to-beat heart intervals (RR intervals), pulse wave obtained from the photoplethysmography signals, and other markers of heart rate variability (HRV), and related cardiovascularparameters such as blood pressure and respiration. Hence, an advantageousconcept introduced in the present disclosure is the use of information onmusicians’ interpretations (or other people’s interpretations, such as musicalexperts’ interpretations) of music pieces (or other media content, such as videothat does not include audio), represented in Interpretation Maps, in the modelling of physiological signals. In general terms, the methods may comprise receiving one or more of theprosodic labels as user input (i.e., Interpretation Maps may be input by a user).The prosodic labels may comprise one or more manual annotations of the media content provided by a user (e.g., annotations provided by a performer, or by an expert). Such annotations could be based on a musical expert’ssubjective opinion on the prosodic properties or on a film expert’s subjectiveopinion on visual prosodic properties (e.g., change in contrast, flashing lights, bright colours, warm colours, etc.). The prosodic properties may comprise one or more automatically-detected properties of the media content (e.g., using logic to automatically determine an Interpretation Map). In some instances, the prosodic properties may comprise one or more measured properties of the media content (e.g., the prosodic properties may be objective, measurableproperties, such as pitch, tempo, note density, etc. as applied to music, orbrightness, contrast, etc. as applied to visual content).The present disclosure demonstrates that incorporating information from anInterpretation Map into computational models can account for over 50% of thevariation in the musicians’ heart rate response when they play musical mediacontent, a marked increase from models that do not incorporate any form ofInterpretation Map. Features of the Interpretation Maps described herein canalso be used for explaining non-player or non-performer (e.g., listener orviewer) responses as well, for music and for other types of media content.Interpretation Maps can be used as advantageous representations ofmusicians’ choices when shaping the musical communication. Theseinterpretive decisions can include moments of climax and release, melodicinterest, melodic interaction or dialogue, significant accompaniment, andmoments of concern. An example of an Interpretation Map can be seen inFigure 10, which is discussed in more detail subsequently. These expressive structures may be marked by the musician(s), by a listener, or automaticallyextracted (e.g., using machine learning techniques). Hence, in general terms,the methods described herein may comprise automatically detecting one or more of the prosodic labels by analysing the media content (e.g., using learning techniques) and / or receiving one or more of the prosodic labels as user input (e.g., as data 418).Interpretation Maps can be represented mathematically and integrated asindependent variables into computational models of consumers’ physiologicalresponses. They can, therefore, act as an important input used together withother musical features, such as tempo, loudness, MFCC (Mel-FrequencyCepstral Coefficients), timbre spectral parameters (e.g.: spectral centroid,spectral spread, spectral flux, spectral entropy), harmonic tension parameters(diameter, distance from the key) and note density, to improve descriptive andpredictive models of physiological reaction to music (which could be live orrecorded). Examples 1 and 2 of the present disclosure show how the Interpretation Map can be used to significantly improve the accuracy of predicting players’ RR intervals. The results of Examples 1 and 2 aresummarised below in generalised terms.Interpretation Maps have been successfully used in linear mixed models forpredicting RR interval (interval between heartbeats, i.e. the inverse of heart rate) time series collected from a trio of professional musicians (pianist, violinist, and cellist). The trio performed Schubert’s Trio No. 2, Op. 100,Andante con moto in nine rehearsals over five days. In generalised terms,therefore, in the methods described herein, the predictive model may be a linear mixed model. The models in some of the examples described herein used the following data: ECG signals measured with a Polar H10 (Polar Electro Oy, Kempele, Finland)HR monitor (to provide physiological responses); Audio signals recorded witha Zoom H5 (Zoom, Tokyo, Japan) handheld recorder for extracting musicfeatures such as tempo and loudness (and from which various prosodicproperties can be determined and used for associating physiological responseswith prosodic labels); Annotations of musicians’ interpretation choices in anInterpretation Map (for defining prosodic labels). The musicians noted theirinterpretation choices, capturing their intention, immersion and reflection,components of the performances that are likely to affect heart rate.In the examples described herein, the Interpretation Map comprised 7 salientcategories of musical action: 1) melodic interest, 2) melodic interaction (e.g., melody & counter melody together), 3) dialogue (e.g., call and answer, asynchronous), 4) significant accompaniment (as opposed to accompaniment that is background), 5) climax (usually preceded by a build-up into the climax, like a crescendo), 6) return / repose (however, it was not used in the model due to too small occurrence), and 7) moments of concern (during which musicians need to manage greater risk). Other prosodic properties of media content can be incorporated. Accordingly, in generalised terms, in the methodsdescribed herein, the prosodic properties may comprise any one or more ofand / or may comprise a change (e.g., an indication of a point at which a change in a prosodic property occurs, which may be described as a “change point”) in any one or more of:^ tempo;^ loudness;^ articulation (contrastive durations or accent patterns to make notesstand out);^ note density (number of notes in a unit of real-time (such as a second)or score time (such as a beat));^ Mel-Frequency Cepstral Coefficients (MFCC; a specific representation ofthe short-term power spectrum of a sound, related to, among other things, timbre and loudness);^ timbre (a quality of sound that differentiates it from others of the samepitch or volume) and its parameters (parameters: spectral centroid, spectral spread, spectral crest, spectral flux, spectral kurtosis, spectral decrease, spectral entropy, spectral flatness);^ pitch (the frequency of sounds) and their combination (harmony);^ melody (a sequence of pitches) and shorter structures (e.g.: phrases);^ tonality (hierarchical organisation of pitches) such as chord and keys,and the onset of chord / key change;^ harmonic tension (tension caused by combinations of pitches) such asdissonance (e.g. diameter of a pitch cluster in the spiral array), distance from the key (tensile strain), rate of chord change (momentum);^ rhythm (a sequence of durations which may or may not repeat);^ meter (periodic grouping of musical beats);^ climax (an intense, exciting and / or important point in the music);^ melodic interest (a salient melody that captures the listener's attention);^ melodic interaction (the interrelation of two or more melodies performedby two parts (voices) through one or more instruments; in Western classicalmusic, this is related to the counterpoint);^ dialogue (synchronous interaction between two melodies like a call-and-answer);^ accompaniment (a musical part that supports the salient melody);^ return (a return to the melody or theme previously played, usually arelease preceded by anticipation);^ repose / relief (a state of calm or ease after effort, tension, and / orstrain);^ moment of concern (a state when musicians need to manage greaterrisk);^ crescendo / decrescendo (gradual increase / decrease of sound intensity);^ swell (a gentle increase followed by a decrease of music intensity);^ emphasis (highlighting a particular note or groups of notes, or chord,usually through a dynamic accent or could also be time);^ resolution / resolve / release (the movement of a note or chord frominstability (such as dissonance) to stability (such as consonance));^ tension buildup (a distinctive increase of music intensity in terms ofvolume, harmonic tension, emphasis, engagement and / or tempo);^ close (an emphasised end of a phrase, passage, or whole composition(a cadence), linked with harmonic resolution);^ imitation (a pattern, say melodic fragment, repeated in a different voiceor register in a polyphonic texture within a short time span);^ novel melody (the introduction of a novel melody, usually accompaniedby other distinctive musical material, for the first time in the piece);^ fast sequence, such as tickly, rise, and / or climb (a rapid ascending ordescending note sequence or pattern);^ significant silence (a salient pause during the piece, after which the nextpart of the piece continues or commences);standout articulation such as accent, jagged, and / or ping (a distinctivearticulation of notes, such as sharply detached notes); and / or^ instability (a state which conveys an impression of imbalance anduncertainty). It should be noted that the above examples of prosodic properties are provided by way of example and that various other prosodic properties can be incorporated into embodiments of the present disclosure.In the examples described herein, an example of mathematical representationwas developed. A binary vector was created for each annotated category; a profile was then generated using the sum of Gaussian kernel density estimationfunctions. Hence, in general terms, at least some of the prosodic labels, andoptionally each of the prosodic labels, may be stored as an annotated musicalscore (e.g., an Interpretation Map) indicating the distribution of the prosodiclabels throughout the item of media content. Since the annotated musical score could apply to any other type of media content, at least some of the prosodic labels, and optionally each of the prosodic labels, may be stored as a set ofannotations (e.g., an Interpretation Map could represent an annotated filmscript, or television show script, or annotated stage directions, etc.) indicating the distribution of the prosodic labels throughout the item of media content. The annotated musical score may be stored in an array. For example, the array may be a one-dimensional vector. Each element (i.e., each entry) in the array may comprise an indication (which could be a binary indication of the presence or absence) of whether a prosodic property of the media content is present at a point in time in the media content. The indication could comprise an indication of whether a prosodic property is present at a particular time in a musical score (as measured using score-time, i.e., by calculating the point in time in the media content mathematically). In some embodiments, processing the recorded one or more physiological responses in conjunction with the prosodic labels to generate the predictive model may comprise generating a profile (see Figures 11 and 12, for instance) of the item of media content by summing contributions (e.g., by summing Gaussian functions where the presence of properties is indicated by binary indications) from the prosodic labels in the array to thereby generate the predictive model in dependence on (i.e., based on) the profile.An advantage of using Interpretation Maps is the step improvement in themodelling of the physiological reaction as compared to models without Interpretation Maps. Interpretation Maps can help to explain the music component in the overall dynamic changes in physiological signals in response to music. As observed in the proof of concepts described herein, more than half of the variability in the prediction results can be associated with the structures captured by Interpretation Maps.Interpretation Maps can, therefore, facilitate transparent, explanatory modelsto be created that associate precise music structures resulting from interpretive decisions (communicated through expressive musical sound structures) with specific physiological responses. Furthermore, InterpretationMaps are able to point to moments in items of media content that are mostinfluential for eliciting physiological reactions, whereas general continuous features like the loudness envelope may give only weak indicators of possibleinfluence. With Interpretation Maps of a musical playlist, models of individuallistener response can be created in real-time to provide therapeuticinterventions with feedback in digital music therapeutics. In some cases, movie productions could benefit from testing or validating listener responses to movie music and peak responses based on elements in Interpretation Maps. These expressive structures may be marked by the musician(s), by a listener, or automatically extracted. The Interpretation Map can be represented mathematically and integrated as independent variables into computationalmodels of listeners’ (including the players’) physiological responses. AnInterpretation Map can be provided as an input used together with other musical features such as tempo, loudness, and note density to improve descriptive and predictive models of physiological reaction to music. Forinstance, Interpretation Maps can be used to significantly improve the accuracyof predicting players’ RR intervals; while the physiological response may be the physiological response while a musician plays the music, the physiological responses of listeners can also be predicted. In some embodiments, the predictive models described herein may be based on a machine learning algorithm that learns the relationship between prosodic labels and physiological responses. The machine learning algorithm could beor could comprise any one or more of: a neural network; a support vectormachine; a decision tree; a random decision forest; a convolutional neuralnetwork; and / or a regression model. In some cases, embodiments may play aretrieved one or more pieces of media content to the person; and record theperson’s actual physiological response to the played media content. Themethods may further comprise comparing the person’s actual physiological response to the played media content with the predicted physiological response to the same media content; and updating the predictive model based on the comparison. A difference between the predicted and actual physiologicalresponse could be used as a cost function to train the predictive model. Thus,the predictive models described herein could be trained iteratively. As shown in Figure 5, using embodiments of the present disclosure, substantialchanges in systolic and diastolic blood pressure (BP) and Pulse Pressure (PP) areobserved in response to the annotated expressive music structures and prosodyparameters (for example, melody interaction, an increase of MFCC, novel melody).For example, in the ‘melodic interaction’ category, this factor is significantly related to a change in Pulse Pressure (p<0.001); a significant increase of MFCC with a change in diastolic (p<0.001) and pulse pressure (p<0.001). The results of such observations are shown in Figure 5. To further demonstrate the efficacy of the methods described herein, twoexamples are provided. The examples relate to musical performances but canbe extended to other types of media content. For example, while the examples describe methods for analysing predicting physiological responses of people asthey listen to music, other types of media content (e.g., any audio files, anypieces of music, any video file whether it includes sound or not, televisionprogrammes, and / or film) can elicit physiological responses when consumed.Such other types of media content can also be used with the predictive modelsdescribed herein. Example 1 In a first example, beat-by-beat analysis of heart rate variability (HRV) and respiration in the vicinity of musical structures (e.g., contrast, novelty, return, change points in loudness, MFCC, tempo, etc.) was investigated. Music changes elicit physiological effects, specifically in the autonomic nervous system (ANS). However, there is a lack of analyses on ANS response to specific music structures in a short time scale. Some embodiments of this disclosure evaluate ANS response to music structures (or other prosodic properties of media content) at an unprecedented scale. In this example, designed to test embodiments of the present disclosure, 83 participants listened to nine classical pieces (original and altered versions)rendered on a reproducing piano. Experts manually annotated the start ofmusic structures, such as novel melodic material, melodic returns, arrival,buildup, climax, closing, drop / gentle, emphasis, fast sequence release / sparkle / relief, imitation, melodic interaction, resolve / release, run / fast sequence (tickly, rise, climb), significant melody, significant silence, standout articulation (accent, jagged, ping), swell, unstable / transition, which areaccompanied by heterogeneous changes. Apart from manual annotations, automatic annotations of change points in loudness, tempo, MFCC signal, diameter (harmonic tension), and spectral centroid (timbre) were detected using the change point detection algorithm (Rebecca Killick, Paul Fearnhead,and IA Eckley. Optimal detection of changepoints with a linear computationalcost. Journal of the American Statistical Association, 107(500):1590{1598,2012.). Differences in heart rate variability (HRV) parameters, blood pressureparameters (systolic, diastolic, pulse pressure), and respiratory parametersbefore and after these transitions were calculated. Statistical analysis of themean values of physiological parameters (systolic BP, diastolic BP, pulse pressure (PP), respiratory intervals (resp. Intervals), instantaneous peak frequency in low and high frequency band FLF, FHF, power in low and high frequency spectrum PLF, PHF) before and after the onset of different categories were calculated using statistical tests.Several significant effects of music prosody parameters on physiologicalparameters were observed, including significant decrease of respiratoryintervals with the onset of novel melody (p<0.001), change point related to an increase of spectral centroid, MFCC, and tempo (p<0.001), significant increase of FHF with the onset of novel melody (p<0.001), and increasing of MFCC (p<0.001), a decrease of FLF with the onset of novel melody (p<0.001), and an increase of diastolic BP with an increase of MFCC (p<0.001).Therefore, music transitions, especially novel melodic material and MFCCchanges, can induce increasing sympathetic activity and respiratory rate, which can inform the design of musical stimuli. This can be used for therapeutic and other applications. For example, physiological responses can be used to build an index relating predicted physiological responses to items of media content, so that items of media content that are likely to elicit a specific physiological response can be identified. This procedure is described in further detail below. Music (and other items of media content) constantly changes its structure within a piece; for example, contrast, novelty, returns, and various other prosodic properties can be identified within music. The uniqueness of musicpieces, the musical expressions introduced during its communication, the individual perception of listeners due to the effects of familiarity, musical sophistication, transitory memory and attention and even health status can make the study of physiological response to music a complex problem. Existing studies focus primarily on the overall effects of music (or simpleauditory stimuli) - classified by genre, happy / sad, pleasant / unpleasant -indicating changes in physiological signals, in comparison to the baseline (Bernardi L. Cardiovascular, Cerebrovascular, and Respiratory Changes Induced by Different Types of Music in Musicians and Non-Musicians: the Importance of Silence. Heart 2005;92(4):445–452; and Tschacher W, Greenwood S, Ramakrishnan S, Tröndle M, Wald-Fuhrmann M, Seibert C, Weining C, Meier D. Audience synchronies in live concerts illustrate the embodiment of music experience. Scientific Reports October 2023;13(1). ISSN 2045-2322). Music is often treated as a whole without consideration for the impact of its structures or parts. There is a lack of studies reporting the instantaneous response of ANS to elements of music in a short time scale. One motivation behind the present disclosure is to investigate physiological responses during music listening (or other media content consumption) under ecological conditions. The aim is to expand the knowledge base on music- induced changes in ANS activity for potential use in various fields, such as indexing / classifying content, cardiovascular therapeutics. Music pieces can be decomposed into events that elicit physiological changes in heart rate, HRV parameters, respiration patterns, and blood pressure (BP). Relatively strong physiological response may be associated with the sensing of the onsets of musical prosodic structures, with the perception of transitions from one distinguishable part of music to another. Some embodiments of this disclosure analyse ANS response during music listening in the vicinity of musical structures, for example based on instantaneous changes in RR intervals, respiratory rate, and spectral HRV parameters. A detailed beat-by-beat analysis of physiological signals before and after the onsets of thesestructures provides explanations for subtle music-induced ANS changes on an unprecedented scale. Participants were invited to listen to nine pieces of Western classical music in their original or altered (fast / slow, loud / soft) versions, rendered on areproducing piano - eight unique pieces in randomly selected versions and therepeat of the first piece in a version different from the first time - followed by5-minute baseline. The pieces were designed to elicit diverse physiological responses due to a broad variation of music structures, transitions, and compositional styles. The following physiological signals were collected: ECG traces and RR intervals acquired by a heart rate monitor chest strap (a Polar H10 by Polar Electro Oy, Kempele, Finland), respiratory signals via a respiratory sensor band (BIOPAC by BIOPAC Systems, Goleta, USA), and continuous BP via a BP sensor (CNAP sensor by CNSystems, Graz, Austria). All music pieces were recorded on a handheld audio recorder (Zoom H5 by Zoom, Tokyo, JP) for calculating musical features, such as loudness, tempo, MFCC, harmonic tension, and timbre parameters (spectral centroid, spectral spread, spectral crest, and spectral flux). Participants were also asked to report their impressions after each piece, including whether the music was familiar to them.Figure 8(2) (Listener setting) depicts the process of data collection andpossible analysis. In some embodiments, the participant can be exposed to music and media content using different methods (audio signals, reproducing piano, live music, speakers, headphones). All collected signals were resampled to 5 Hz. RR intervals were filtered using a high-pass Butterworth filter with a cut-off frequency of 0.03 Hz before calculating the HRV spectral parameters, such as instantaneous peak frequencies FLF, FHF and the instantaneous powers PLF, PHF of LF and HF bands in a spectrogram obtained using the short-time Fourier transform (STFT). While the LF component was defined by a standard range of frequencies (0.04-0.15 Hz), the range for the HF component was made time-varying and dependent on the respiratory frequency Fr(t) (based on the respiratoryintervals extracted from the respiratory signal), Fr(t) ± [−0.125; 0.125] Hz.This methodology based on the method of Orini M et al. (Orini M, Bailón R, EnkR, Koelsch S, Mainardi L, Laguna P. A method for continuously assessing the autonomic response to music-induced emotions through hrv analysis. Medical amp Biological Engineering amp Computing March 2010; 48(5):423–433. ISSN 1741-0444) improves the estimating of respiratory sinus arrhythmia (RSA) in continuous time. All physiological signals were analysed by calculating their mean values in 10-second windows before and after each onset of the music transition category. Music structures were annotated by two experts who listened to and annotated music pieces exposed to the participants in the same laboratory environment. Experts manually annotated the start of music structures such as novel melodic material, melodic returns, arrival, buildup, climax, closing, drop / gentle, emphasis, fast sequence release (also sparkle, relief), imitation, melodic interaction, resolve / release, run / fast sequence (tickly, rise, climb), significant melody, significant silence, standout articulation (accent, jagged, ping), swell, unstable / transition, which are accompanied by heterogeneous changes. Apart from manual annotations, automatic annotations of change points in loudness, tempo, MFCC signal, diameter (harmonic tension), and spectral centroid (timbre) were detected using the change point detection algorithm (RebeccaKillick, Paul Fearnhead, and IA Eckley. Optimal detection of changepoints witha linear computational cost. Journal of the American Statistical Association,107(500):1590{1598, 2012.). Altering the tempo and / or loudness of arecorded music piece changes its perception, thus the annotations could vary between different versions of the piece. The differences in mean values of HRV spectral parameters, RR and respiratory intervals obtained from 10-second windows before and after the onsets of the music transitions were analysed. The significance of these differences was calculated using paired t-Student or Wilcoxon tests (whenever the distribution was not normal), separately for each annotation category. Surrogate data analysis was performed to determine if random effects could be the cause of changes observed in the original data. The surrogates analysis was 50iterations of the original analysis - based on comparing means before and after- applied to random timestamps in the piece. Moreover, the mean values ofphysiological parameters and 95% confidence interval in each one-second timewindow, [t-0.5 s; t+0.5 s], where t is the time before or after the transition,were calculated and tested to determine whether each mean significantly differs from zero. Moreover, Bonferroni correction was used to adjust the p- value considering multiple comparisons and control the false discovery rate (FDR) or the family-wise error rate (FWER). The results were obtained from 83 participants (52 females and 31 males) with a mean age of 42.2±13.3 years (range 21-78 years); 26 participants were found to have high baseline BP (>140 / 90 mmHg). The total number of music structures over all pieces, annotated by experts, was 827. During all sessions, participants were exposed to 17,341 onsets. Among all categories of prosodic parameters and physiological signals, the strongest effects were observed in a significant decrease of respiratory intervals with the onset of novel melody (p<0.001), change point related to an increase of spectral centroid (p<0.001), MFCC (p<0.001), significant increase of FHF with the onset of novel melody (p<0.001), and increasing of MFCC (p<0.001), a decrease of FLF with the onset of novel melody (p<0.001), and an increase of diastolic BP with an increase of MFCC (p<0.001). These effects remained significant after using Bonferroni correction to adjust p-value. The mean values, 95% confidence intervals, and annotations of all significanteffects for all parameters are shown in Figure 6. Analysis of surrogate datashowed no significant differences for any parameter.Figure 7 shows examples of parameter changes in each of ±10 seconds fromthe onset of novel melody, change points related to increasing of MFCC andspectral centroid (a ⋆ marks the one-second segment with values significantlydifferent from zero). It can be observed that the maximum change of theparameter values from the event onset, in some cases, does not occur instantlybut with a delay (approximately 4 seconds). The second-by-second analysis also shows that in some cases (e.g. FHF for novel melodies), the mean values significantly differ from zero before the transition. A significant effect of music structures on ANS can be observed. Among alldefined categories, the strongest mean responses were observed for novelmelodies, an increase of MFCC, and an increase of spectral centroid parameter,especially for parameters respiratory intervals, FLF, FHF. A significant increaseof FHF and a decrease in respiration intervals in both groups suggest a tendency for the respiratory rhythm to accelerate during onsets of these parameters. Figure 7 shows that the respiratory intervals start decreasing even before the transition, which might anticipate the autonomic changes. It can also be observed that analysed parameters reach their minimum or maximum values after a delay (for example, the minimum value of PHF in the ’novel melody’ group occurs at the 4th second and the maximum FHF in the ’return’ group occurs at the 7th second after transition onset). These delays could be associated with the limited time resolution of the spectrogram, but they can also be due to the natural reaction time to auditory stimuli or the effect of music expectations (Juslin PN, Västfjäll D. Emotional responses to music: The need to consider underlying mechanisms. Behavioral and Brain Sciences October 2008; 31(5):559–575). In some embodiments, a method for tracking participants’ attention while consuming media content (apart from questioning them about their impressions after each piece) may be provided. This may improve the reliabilitywith which physiological response data can be attained.Accordingly, acute changes in physiological signals can be observed in response to musical structures and prosody parameters to novel melodic material, changes in MFCC, and musical timbre, or in response to other prosodic properties of media content. The aggregate results suggest a tendency to sympathetic activation with increasing respiratory rate due to reactions to these events in music. The outcome emphasizes the non- stationary nature of physiological signals while listening to music. Example 2 In a second example, performers’ beat-to-beat heart intervals were modelled using music features and Interpretation Maps. As mentioned previously, musicstrongly modulates the ANS. This modulation is evident in musicians’ beat-to-beat heart (RR) intervals, a marker of heart rate variability (HRV), and can berelated to music features and structures. This second example uses a novelapproach to modelling musicians’ RR interval variations, analysing detailedcomponents within a music piece to extract continuous music features andannotations of musicians’ performance decisions.A professional ensemble (violinist, cellist, pianist) performed Schubert’s TrioNo. 2, Op. 100, Andante con moto nine times during rehearsals. RR intervalseries were collected from each musician using wireless electrocardiography(ECG) sensors. Linear mixed models were used to predict their RR intervalsbased on music features (tempo, loudness, note density), interpretive choices(as indicated in an Interpretation Map), and a starting factor. The modelsexplain approximately half of the variability of the RR interval series for allmusicians, with R-squared = 0.606 (violinist), 0.494 (cellist), and 0.540(pianist), demonstrating robust performance. The features with the strongestpredictive values were loudness, climax, moment of concern, and startingfactor. Accordingly, the method revealed the relative effects of different musicfeatures on autonomic response and shows, for the first time, a strong linkbetween an Interpretation Map and RR interval changes. Modelling autonomicresponse to music stimuli can be used to develop medical and non-medicalinterventions. The models described herein can serve as a framework forestimating performers’ physiological reactions using only music information (orany sound content in any media content) and these predictive models can alsoapply to listeners (i.e., to any person listening to any type of media content,whether or not the person is aware that they are listening to the media content, e.g., the content may be played in the background).Live music performance provides a unique context in which to investigatecardiac and other physiological responses in ecological and engagingsettings. Musicians engage in physical activity and mental coordinationwhile performing, which stimulates the body through the autonomic nervoussystem (ANS). One of the most relevant measures of physiological responseto music stimuli is heart rate variability (HRV), a marker of ANS.Studies of musicians’ autonomic response have mainly focused onperformance stress and the effects of performance context rather than theimpact of the music itself. Prior work has considered changes in physiologicalsignals with factors such as the intensity of the physical effort on specificinstruments (Iñesta, C., Terrados, N., García, D., and Pérez, J. A. (2008).Heart rate in professional musicians. Journal of Occupational Medicine andToxicology 3. doi:10.1186 / 1745-6673-3-16), ecological setting (Mulcahy,D., Keegan, J., Fingret, A., Wright, C., Park, A., Sparrow, J., et al. (1990). Circadian variation of heart rate is affected by environment: a study of continuous electrocardiographic monitoring in members of a symphonyorchestra. Heart 64, 388–392. doi:10.1136 / hrt.64.6.388; and Williamon, A.,Aufegger, L., Wasley, D., Looney, D., and Mandic, D. P. (2013).Complexity of physiological responses decreases in high-stress musicalperformance. Journal of The Royal Society Interface 10, 20130719.doi:10.1098 / rsif.2013.0719), and difficulty of particular pieces (Mulcahy etal.). One such example is the analysis of Williamon et al., which examinedmusicians’ beat-to-beat heart rates on a continuous scale. In this study, aprofessional musician performed Bach’s English Suite in A minor (BWV 807)for a large audience and in a laboratory setting while measuring hisphysiological response. The analysis also combined the performer’s basicannotations of playing challenge–the first and third movements were markedas the most challenging–along with changes in stress levels as measuredphysiologically. The results suggested autonomic responses are more powerfulin live, ecological settings.However, there is a lack of work that considers detailed analyses of themusicians’ autonomic responses on a continuous scale and their performancedecisions, intentions, and moments of concern while playing a music piece. Theeffort expended in performance involves planning and action to achieve musicalintentions (Hatten, R. S. (2017). A Theory of Musical Gesture and itsApplication to Beethoven and Schubert. In Music and Gesture (Routledge). 1– 18. doi:10.4324 / 9781315091006-2). Performing music also affects musiciansin the present (current actions), past (recovery from previous actions), andfuture (anticipating next actions). Embodiments of the present disclosure, as exemplified in this example, address this gap by examining musicians (violin, cello, and piano) playing Schubert’s Trio No. 2, Op. 100, Andante con moto and predicting their RR intervals based on continuously measured music features and an “ Interpretation Map containing annotations of the musicians’ interpretive decisions. Some embodiments provide some explanation of the physiological response to music performance using regression modelling that considers these music markers so as to inform music training and potential cardiovascular therapeutic applications (or for non-therapeutic, e.g., for recommending content intended to elicit a certain response). Embodiments of the present disclosure are unusual in that they represent for the first time a music-only method being used to predict autonomic response in musicians.The participants in this example were a trio of professional musicians with over20 years’ performance experience, and more than 10 years playing together.The musicians were instructed to get at least 7 hours of sleep and avoid anycaffeinated and alcoholic beverages for 8 hours and food for 1 hour prior to theagreed rehearsal time.The trio performed Schubert’s Trio No. 2, Op. 100, Andante con moto(henceforth, “the Schubert”). The piece offered a balance between varied musicfeatures and clear musical structures. The alternating sections marked by themusical themes enabled comparisons between similar and repeated musicfeatures. The tempo of the piece resembles a typical human HR, which allowsmusic features to be aligned to physiological data without oversampling thesignals. The musicians had not practised the Schubert together before the firstrecording. To measure ECG in each performance, the Polar H10 (Polar Electro Oy,Kempele, Finland) heart rate sensor was used. The Polar unit registers a one-channel ECG signal with a sampling frequency of 130 Hz and automatically detects QRS complexes to create RR interval series. Data samples from the Polar straps were collected via Bluetooth on three iPhones (Apple, Cupertino,CA, USA) running the HeartFM mobile app on iOS 16.1.2 or higher. A Zoom H5(Zoom, Tokyo, Japan) handheld recorder was used for audio recording. The Polar straps (Figure 8(1)(a)) were moistened and worn across each musician’s sternum. The musicians were seated in a typical trio configuration. The Zoom recorder (Figure 8(1)(b)) was placed approximately 2 metres from the musicians. The three iOS devices and Zoom were synchronised with a clapper for later alignment of the signals. The Schubert was recorded nine times over five different days.Accordingly, Figure 8(1) shows data collection: the trio perform the Schubert9 times while wearing Polar straps (a) to collect ECG data. A Zoom recorder gathers audio data from each performance (b). Figure 9 shows how players annotated categories of performed structures inthe music score. On a separate date following the recordings, the musicianscollaborated to annotate the music score, developing a novel set of categoriesreflecting their negotiated roles and collective actions. This representation is referred to herein as an Interpretation Map. The Interpretation Map captures the musicians’ experienced cognitive and physical load, expressive choices, and how they negotiated a path through the music. The trace of the musicians’ intention, immersion and reflection embedded in this Interpretation Map represent components of the performances that likely affected their heart rates. The Interpretation Map comprises one or more manual annotations of the Schubert indicating the distribution of prosodic labels throughout the item of media content.The categories in the Interpretation Map in this example were: (1) melodicinterest (main melody); (2) melodic interaction (e.g., melody and counter-melody together); (3) dialogue (e.g., asynchronous call and answer); (4)significant accompaniment (as opposed to accompaniment that isbackground); (5) climax (usually preceded by a build-up into the climax, likea crescendo); (6) return / repose (however, it was not used in the model dueto low occurrence); and (7) moment of concern (during which musicians needto manage greater risk). Figure 10 shows the distribution of tempo, loudness, Interpretation Map categories (melody, dialogue, significant accompaniment, accompaniment, climax, moments of concern), kernel Gaussian functions based on theaforementioned annotations (in dashed lines), a time series introducing thestarting factor (black dashed line) and RR intervals in score-time for each of the three musicians. The RR intervals are predicted using Model Sets 1 (tempo, loudness), 2 (+ Interpretation Map), 3 (+ starting factor), as explained in more detail below. Elements of the data preparation are discussed below: Score-time DomainTo focus on musically salient features and their effect on physiological dataacross multiple performances, the RR interval time series and performanceaudio were synchronised to score-time (Chew, E. and Callender, C. (2013).Conceptual and Experiential Representations of Tempo: Effects on Expressive Performance Comparisons. In Mathematics and Computation in Music. MCM 2013 (Berlin, Heidelberg: Springer), Lecture Notes in Computer Science. 76—-87. doi:10.1007 / 978-3-642-39357-0 6) – i.e., with musical beats instead ofseconds as time axis – to match them to the performers’ score annotations.Other continuous time signals were adjusted to the annotated beats usinglinear interpolation.The conversion to score-time also matched the physiological signals to musicfeatures extracted from the score and makes the signals from all performances comparable. The half-beat and the eighth note pulse in the 2 / 4 meter were used as a unit for the 848 eighth note beats in Schubert. Physiological DataThe RR interval series generated by Polar’s automatic QRS complex detectionwere reviewed for inaccurate beat detection and premature beats; prematureatrial and ventricular beats were removed. The series of normal beats werenormalised. The normalisation of the RR intervals allows a comparison of theeffect of the music structures annotated in the Interpretation Map between the musicians. Since the musical parameters extracted from the score are commonfor each performance, it is useful to assess their effects independent of thebaseline RR interval value. Music Audio FeaturesThe eighth note pulse in the recorded performance audio was manuallyannotated using Sonic Visualizer (Cannam, C., Landone, C., and Sandler, M.(2010). Sonic Visualiser: An Open Source Application for Viewing, Analysing, and Annotating Music Audio Files. In Proceedings of the ACM Multimedia 2010International Conference (Firenze, Italy), 1467–1468); the annotations werecorrected to note onsets using the TapSnap algorithm of the CHARM MazurkaProject. Three music features related to the musicians’ physical effort whileplaying were computed: note density; loudness; and tempo. Note density, thenumber of note events per beat, was calculated using the Matlab MIDI Toolbox(Eerola, T. and Toiviainen, P. (2004). MIDI Toolbox: MATLAB Tools for MusicResearch (Jyväskylä, Finland: University of Jyväskylä)) as a measure oftechnical complexity. Perceptual loudness in songs was derived using the MusicAnalysis Matlab Toolbox (E., P. (2004). A Matlab Toolbox to Compute MusicSimilarity from Audio. In ISMIR International Conference on Music InformationRetrieval). The tempo was computed from the beat annotations in beats perminute (BPM). Tempo and loudness from each performance were normalised,and loudness was filtered with a low-pass Butterworth filter with order N = 3and Wn = 0.125 (cutoff frequency parameter in butter function in Python).Score AnnotationsThe musicians’ score annotations were transformed into signals used as aninput to the models. A length L binary vector (where L is the number of samplesin the piece in score-time domain, here L = 848) was created for eachannotated category, with 1 indicating the occurrence of an annotation (0otherwise). These binary vectors are the reference for the input signals, whichare the sum of Gaussian kernel density estimation functions (Silverman, B.(2018). Density Estimation for Statistics and Data Analysis (The city:Routledge)). When the i-th sample is 1 in the binary vector, a Gaussian function(standardised to sum = 1) centred at the i + 16th (4 bars) index and SD =16 / 2 (two bars) was added to the final vector. These parameters relate to the2 / 4 meter of the Schubert. The generated Gaussian functions are summed ifthere is more than one 1 sequentially in the binary vector. In these cases, the number of these values are related to the amplitude (thus the strength of theeffect of a given category) of the function. Figure 11 illustrates an example ofa binary and Gaussian signal, demonstrating how sequential annotated sections of moment of concern (cellist) are summed. Starting factorThe musicians’ physiological responses were observed to be different at thebeginning of playing due to individual activation of autonomic mechanismsunderlying cardiovascular reactivity rather than any specific physical or musicfeatures. The first part of the Schubert is not physically challenging (lowloudness, simple piano accompaniment, calm introduction of the cello theme),and the decline in RR intervals is relatively high through the first 20 bars ofmusic. The initial stress reaction to mental tasks caused by the process ofswitching from rest to stimulation has been associated with the disruption ofbaseline homeostasis (see: Hughes, B. M., Lü, W., and Howard, S. (2018).Cardiovascular stress-response adaptation: Conceptual basis, empirical findings, and implications for disease processes. International Journal ofPsychophysiology 131, 4–12); and Widjaja, D., Orini, M., Vlemincx, E., VanHuffel, S., et al. (2013). Cardiorespiratory dynamic response to mental stress: a multivariate time-frequency analysis. Computational and mathematicalmethods in medicine 2013; and Kelsey, R. M., Blascovich, J., Tomaka, J.,Leitten, C. L., Schneider, T. R., and Wiens, S. (1999). Cardiovascular reactivity and adaptation to recurrent psychological stress: Effects of prior taskexposure. Psychophysiology 36, 818–831; and Hughes, B. M., Howard, S.,James, J. E., and Higgins, N. M. (2011). Individual differences in adaptation of cardiovascular responses to stress. Biological Psychology 86, 129–136).This initial physiological reaction was modelled by introducing a starting factor(tailored to the Schubert), expressed as a time series based on the followingformula:where t is the score-time, m(t) and a(t) terms are binary indicators of whetherthe musician plays a melody or accompaniment (or significant accompaniment)at score-time t, and b1 = −0.2364 and b2 = −0.0871, mean values of theexponents obtained by fitting tb to each musician’s first 80 RR intervals (inscore-time) from the cellist (b1) and pianist (b2) over all performances. Thus,for the cellist, who plays the opening melody, the factor equals t−0.2364, andfor the pianist t−0.0871, where t ∈ [ 0, 80 ], and 0 otherwise. It should be notedthat the violin does not play – she has only rests – during the first 20 bars,so her starting factor has a rectangular shape and equals tb1*0+b2*0 = t0 = 1.The example of the result obtained from the usage of that function (for cellist) is presented in Figure 11.Once the data is prepared, statistical analysis was performed. Linear mixedmodels (LME) were used as predictive models to predict the musicians’ RRintervals. Three models with differing complexity were created for eachmusician:^ Model Set 1: Considered only loudness and tempo (music-relatedfeatures) extracted from the audio files.^ Model Set 2: Included tempo, loudness, note density, and theInterpretation Map (annotations of melodic interest, dialogue, accompaniment,significant accompaniment, climax, moments of concern).^ Model Set 3: Included tempo, loudness, note density, theInterpretation Map, and the starting factor time series.The feature in Model Sets 1-3 were considered as fixed effects. The performance number, 1-9, was added to the models as a random effect (random intercept). In this way, the model’s performance was progressively observed as its complexity increased.The model coefficients of the LME models are presented in Table 1 with 95%confidence intervals and p-values calculated using the bootstrapping method –sampling with replacement from the original dataset 1000 times (Haukoos, J.S. (2005). Advanced Statistics: Bootstrapping Confidence Intervals for Statistics with ”Difficult” Distributions. Academic Emergency Medicine 12, 360– 365. doi:10.1197 / j.aem.2004.11.018). In the bootstrapping method, the p- values are estimated by generating distributions for the coefficients with mean 0 using coefficient and mean values extracted from the original data, i.e. distributions for the null hypothesis that the feature has no effect in the model. Then, the probability is calculated, considering a two-tailed hypothesis, that the mean is significantly different from zero. If the probability is less than 0.05, the null hypothesis (that the feature has no effect in the model) is rejected.The backward, stepwise method based on the Akaike Information Criterion(AIC) was used to eliminate insignificant variables in all models. The R2(coefficient of determination) was used to estimate the variation in thedependent variable explained by the independent variables.The results of the LMEs considering all three model sets for each musician are shown in Table 1:

[0002] Table 1. Parameters used to model RR intervals during performing the Schubert for three model sets. Coefficients (and confidence intervals) that are significant are reported; p < 0.01 was obtained for significant accompaniment in the violinist model and note density and significant accompaniment in the pianist model; otherwise, p < 0.001. It can be observed that the models with music audio features, the Interpretation Map, and the starting factor (Model Set 3) explained more than half of the variability of the musicians’ RR interval time series for the violinist and the pianist, and almost half for the cellist. Based on the R2 values, the greatest improvement in explaining the variability in the RR intervals is observed after adding the Interpretation Map (Model Set 2); adding the starting factor (Model Set 3) further boosted the results.R2 increased with model complexity for all musicians and all groups of models(Model Sets 1-3). Model Set 1 (only audio parameters loudness and tempo)explained 29.3%, 7.2% and 23.7% of the RR interval variability for theviolinist, cellist and pianist, respectively. Adding the Interpretation Map inModel Set 2 increased the R2 values to 54.0%, 48.5%, and 44.6%,respectively. The most complex Model Set 3, adding the starting factor,obtained the highest R2: 60.6%, 49.4%, and 54.0%, respectively. Figure 10shows the original and predicted RR interval series from an exampleperformance, together with the musicians’ annotations and loudness andtempo signals in the score-time domain for all musicians. Table 1 shows that most of the coefficients are negative, meaning that the musicians’ RR intervals tend to decrease with the chosen features, but they vary in value. The results in the final Model Set 3 show that the componentswith the strongest effect on the dependent value are loudness (–0.495), climax(–0.330), and initialisation factor (0.300) for the violinist; climax (–0.712), moment of concern (–0.506), and loudness (–0.246) for the cellist; climax (– 0.495), loudness (–0.431), and initialisation factor (0.327) for the pianist. Climax was the strongest common factor amongst all features for all musicians.It is observed that the largest simultaneous decrease in RR intervals at thebeginning of climatic parts: from bar 67 (268 eighth notes in score time), from bar 115 (460 eighth notes in score time; the musicians’ response in that part is a combined effect of climax, melodic interaction, and moments of concern) and from bar 158 (632 eighth notes in score time). A moment of concern is a type of stress. It has the lowest coefficient value for the cellist (–0.506). This is probably associated with a solo passage of 16th notes in bars 119-121 (476- 484 eighth notes in the piece). The coefficient of the moment of concern was slightly negative for the pianist (–0.151) but slightly positive for the violinist (0.166). However, the violinist labelled only one short segment of the piece (only 3 bars) as a moment of concern, which may be insufficient information to obtain a representative coefficient for this category. The rest of the annotations, including melody, dialogue, significant accompaniment, and accompaniment, have larger coefficient values (> −0.2), the singular exception being the pianist’s annotation of the dialogue (−0.228) which for that instrument occurs largely with the climax. Loudness and tempo have limited value for explaining the musicians’ physiological response. Between the two, loudness has a stronger effect than tempo. Its coefficient has a larger absolute value for all musicians (based on the final Model Set 3): –0.495 (loudness) vs. –0.214 (tempo) for the violinist, –0.246 vs. –0.143 for the cellist, and –0.431 vs. 0.094 for the pianist. Thissuggests that playing louder tends to decrease RR intervals more than playingfaster. One exception is the positive tempo coefficient for the pianist, meaningthe RR intervals increase with tempo, but the effect is relatively small. Thepositive effect of tempo on heart rate was previously reported as a significantfactor in music playing (Iñesta, C., Terrados, N., García, D., and Pérez, J. A. (2008). Heart rate in professional musicians. Journal of Occupational Medicine and Toxicology 3. doi:10.1186 / 1745-6673-3-16).Introducing the starting factor further improved model performance. A sharpdecrease in RR intervals at the start of playing was not as well-captured by themusic-based features in the previous models. The starting factor capturesthe initial disruption of homeostasis toward activation of the sympathetic nervous system. The music-based features in Model Sets 1 and 2 insufficiently accountedfor the sharp decrease in RR intervals at the start of playing – compare the resultsfor the different models in Figure 2 during the first 80 samples. This is likely due to the music features not encoding information that could predict such changes: there were no large changes in loudness and tempo, and no climax or momentof concern. The starting factor also depends on the musician’s role during theinitial bars. The largest reaction was observed for the cellist, who plays the main melody with piano accompaniment. Although the opening motif in the Schubert is relatively simple and easy to play, a disproportionate drop in RR intervals atthe start for the cellist is observed, potentially boosted by the activation of thesympathetic nervous system. A slower decrease in RR intervals is observed for the pianist, who plays an accompanying part. The initialisation factor for the violinist is constant, due to the silence during the first 80 eighth notes.The starting factor can be tailored to each piece according to specific musicstructures at its beginning and over time. Another option is to forego thestarting factor and begin RR interval prediction after the initial autonomicreaction.A strength of the models demonstrated herein is the ability to use only music-based information to model RR intervals, and its explanatory value incomparing the impact of individual factors on RR. Extensions to these modelscan include other physiological measures, such as respiration and physicalmovement. Estimating cardiac response to music playing using onlyinformation from the performance audio – how musicians modulate thedynamics of the piece – and interpretation – the musicians’ decisions andactions – provides a strong link between the autonomic nervous system andplaying music, with the potential for using active music-making incardiovascular therapies.Thus, presented herein is a novel way of modelling the heart rate variability(RR intervals) of musicians based on music features and the musicians’interpretation of the piece. This information can be used for variousadvantageous purposes. Notably, by using only music information extractedfrom audio recordings or the music score, the models were able to explain morethan half of the variability of the RR interval series for all musicians (R2 = 0.540for the pianist, 0.606 for the violinist and 0.494 for the cellist). Loudness,climaxes, and moments of concern were found to be useful features formodelling the players’ RR intervals. These features may be related to physicaleffort or mental challenges while performing the most demanding and engagingparts of a music piece. Another useful feature is the starting factor, indicatingthe importance of separately modelling the initial physiological reaction toplaying music.Hence, it has been shown how instantaneous changes in RR intervals can depend on time-varying expressive music properties and decisions, and a framework is presented for estimating performers’ physiological reactionsusing music-based information alone. Active engagement in music making isan application of cardiovascular variability. Because listeners receive theresults of their musical actions, the approaches described herein can alsoapply to modelling listeners’ physiological responses to music (or other mediacontent) and can be used to develop non-pharmacological therapies. Forinstance, an index of media content can be built, the index allowing a desired physiological response to be attained using a query. Furtherextensions of the methods described herein can compare the autonomicresponse to music stimuli between healthy players and those withcardiovascular diseases. The analysis of differences between these groups ofplayers would help in monitoring autonomic balance during music playing.It will be understood that many variations may be made to the above apparatus, systems and methods whilst retaining the advantages noted previously. For example, where specific components have been described, alternative components can be provided that provide the same or similar functionality. Each feature disclosed in this specification, unless stated otherwise, may be replaced by alternative features serving the same, equivalent or similar purpose. Thus, unless stated otherwise, each feature disclosed is one example only of a generic series of equivalent or similar features. As used herein, including in the claims, unless the context indicates otherwise, singular forms of the terms herein are to be construed as including the plural form and, where the context allows, vice versa. For instance, unless the context indicates otherwise, a singular reference herein including in the claims, such as "a" or "an" (such as a physiological response or a person) means "one or more" (for instance, one or more physiological responses, or one or more people). Throughout the description and claims of this disclosure, the words "comprise", "including", "having" and "contain" and variations of the words, for example "comprising" and "comprises" or similar, mean that the describedfeature includes the additional features that follow, and are not intended to(and do not) exclude the presence of other components. The use of any and all examples, or exemplary language ("for instance", "such as", "for example" and like language) provided herein, is intended merely to better illustrate the disclosure and does not indicate a limitation on the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure. Any steps described in this specification may be performed in any order or simultaneously unless stated or the context requires otherwise. Moreover, where a step is described as being performed after a step, this does not preclude intervening steps being performed. All of the aspects and / or features disclosed in this specification may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. In particular, the preferred features of the disclosure are applicable to all aspects and embodiments of the disclosure and may be used in any combination. Likewise, features described in non-essential combinations may be used separately (not in combination).

Claims

Claims:

1. A computer-implemented method for predicting physiological responsesof people consuming media content, the method comprising:recording one or more physiological responses of a person as the personconsumes an item of media content, the item of media content beingassociated with a plurality of prosodic labels defining prosodic properties of the item of media content; processing the recorded one or more physiological responses in conjunction with the prosodic labels to generate a predictive model predicting physiological responses of people to media content; applying the predictive model to a plurality of further items of media content to generate further predicted physiological responses to the respective further items of media content; and storing an index comprising the further predicted physiological responses indexed to the respective further items of media content.

2. The method of claim 1, further comprising:receiving data indicative of a desired physiological response; and selecting one or more of the further items of media content using the index, the selected one or more further items of media content having predicted physiological responses that correspond to the desired physiologicalresponse as indicated by the received data.

3. The method of claim 2, further comprising providing, in response toreceiving the data indicative of the desired physiological response, the selected one or more further items of media content.

4. The method of claim 2 or claim 3, wherein the desired physiologicalresponse and the predicted physiological responses of each of the selected one or more further items of media content have a similarity metric that exceeds a threshold value.

5. The method of any preceding claim, further comprising performing thesteps of recording one or more physiological responses and processing one ormore physiological responses for each of a plurality of different items of media content to generate the predictive model.

6. The method of any preceding claim, further comprising performing thesteps of recording one or more physiological responses and processing one or more physiological responses for a plurality of different people to generate the predictive model, and optionally for a plurality of different sub-groups of people.

7. The method of any preceding claim, further comprising:automatically detecting one or more of the prosodic labels by analysing the media content; and / or receiving one or more of the prosodic labels as a user input.

8. The method of any preceding claim, wherein the prosodic labelscomprise any one or more of:one or more manual annotations of the media content provided by a user; one or more automatically-detected properties of the media content; and / or one or more measured properties of the media content.

9. The method of any preceding claim, wherein:the prosodic properties comprise any one or more of: tempo; loudness; note density; Mel-Frequency Cepstral Coefficients (MFCC); timbre; pitch; melody; key; rhythm; climax; melodic interest; melodic interaction; dialogue; accompaniment; return; repose; moment of concern; crescendo; swell;emphasis; resolution; tension buildup; close; relief; imitation; novel melody;fast sequence, such as tickly, rise, and / or climb; significant silence; standoutarticulation, such as accent, jagged, and / or ping; instability; and / or release;and / or the prosodic properties comprise a change in any one or more of:tempo; loudness; note density; Mel-Frequency Cepstral Coefficients (MFCC); timbre; pitch; melody; key; rhythm; climax; melodic interest; melodic interaction; dialogue; accompaniment; return; repose; moment of concern;crescendo; swell; emphasis; resolution; tension buildup; close; relief; imitation; novel melody; fast sequence, such as tickly, rise, and / or climb; significant silence; standout articulation, such as accent, jagged, and / or ping; instability; and / or release.

10. The method of any preceding claim, wherein recording the one or morephysiological responses of the person comprises: exposing the person to the item of media content; recording one or more times at which the one or more physiological responses occurred; and associating one or more of the prosodic labels with the one or more times at which the one or more physiological responses occurred.

11. The method of claim 10, wherein processing the recorded one or morephysiological responses in conjunction with the prosodic labels to generate the predictive model is performed in dependence on the association between theone or more prosodic labels and the one or more times at which the one ormore physiological responses occurred.

12. The method of any preceding claim, wherein the physiological responsesare indicated by physiological measurements taken on one or more people.

13. The method of any preceding claim, wherein at least some of theprosodic labels, and optionally each of the prosodic labels, are stored as anannotated musical score indicating a distribution of the prosodic labelsthroughout the item of media content.

14. The method of claim 13, further comprising storing the annotatedmusical score in an array.

15. The method of claim 14, wherein each element in the array comprisesan indication of whether a prosodic property of the media content is present at a point in time in the media content.

16. The method of claim 14 or claim 15, wherein the array is a one-dimensional vector.

17. The method of any of claims 14 to 16, wherein processing the recordedone or more physiological responses in conjunction with the prosodic labels to generate the predictive model comprises generating a profile of the item of media content by summing contributions from the prosodic labels in the array to thereby generate the predictive model in dependence on the profile.

18. The method of any preceding claim, wherein the predictive model is alinear mixed model.

19. The method of any preceding claim, wherein at least one item of mediacontent, and optionally each item of media content, comprises any one or more of: an audio file; a piece of music; a video file; a television programme; a video game; and / or a film.

20. The method of any preceding claim, wherein the person is any one of:an audience of the media content; a participant of the media content; and a player of music in the media content.

21. The method of any preceding claim, wherein the one or morephysiological responses comprise any one or more of: a cardiovascular measurement; heart rate; beat-to-beat heart intervals (RR intervals); heartrate variability (HRV) parameters; blood pressure; and respiration rate.

22. The method of any preceding claim, further comprising recording userdata of the person and generating the predictive model based on the user data, optionally wherein the user data comprises any one or more of: demographic data of the user; an indication of a degree of media content sophistication of the user; and / or one or more media content preferences of the user.

23. A computing system for predicting physiological responses of peopleconsuming media content, the computing system comprising a processor adapted to:record one or more physiological responses of a person as the personconsumes an item of media content, the item of media content beingassociated with a plurality of prosodic labels defining prosodic properties of the item of media content; process the recorded one or more physiological responses in conjunction with the prosodic labels to generate a predictive model predicting physiological responses of people to media content; apply the predictive model to a plurality of further items of media content to generate further predicted physiological responses to the respective further items of media content; and store an index comprising the further predicted physiological responses indexed to the respective further items of media content.

24. A computer program product comprising instructions that, when theprogram is executed by a computing system, cause the computing system to carry out the method of any of claims 1 to 22.

25. A computer readable medium having stored thereon the computerprogram product of claim 24.

Citation Information

Patent Citations

  • Method and system for analysing sound

    EP2729931B1

  • Initialising of a system for automatically selecting content based on a user's physiological response

    US20110179054A1