Identifying media items for target groups
By employing AI models to analyze media content and map descriptors to personality traits, the method addresses the challenge of personalized media recommendations and targeted advertising, achieving enhanced user engagement and relevance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-03-25
AI Technical Summary
Existing systems lack the ability to effectively analyze media content to determine media content descriptors, emotion profiles, and personality profiles, which are crucial for personalized media recommendations and targeted advertising.
A method is developed to determine a set of media content descriptors that best match a target profile by using artificial intelligence models, such as deep learning neural networks, to analyze media items and generate semantic and emotional descriptors, which are then mapped to personality traits using mapping rules learned through machine learning techniques.
This approach enables accurate identification of media items and users that match a target profile, allowing for personalized recommendations and targeted advertising based on user emotions and personalities, enhancing user engagement and relevance.
Smart Images

Figure 2026053716000001_ABST
Abstract
Description
Background Art
[0001] This application relates to the analysis of media content to determine a media content descriptor, a media profile, an emotion profile, and a personality profile from the generated semantic descriptors of media items. The media content descriptor, the media profile, the emotion profile, and the personality profile can be used in many use cases, for example, for determining media items that match a target profile or for determining media users that have a personality profile that matches a target profile. Use cases can include media recommendation engines, virtual reality, smart assistants, advertising (target marketing), and computer games.
Summary of the Invention
[0002] In a broad aspect, the present disclosure relates to determining a set of media content descriptors that best matches a target set of content descriptors from a plurality of sets of media content descriptors associated with a media item. The media item can be any type of media content, particularly an audio or video clip. The audio media item preferably includes a piece of music or a portion of a piece of music, preferably a piece of music. Photos, a series of photos, videos, slides, and graphic representations are further examples of media items. The set of media content descriptors includes semantic descriptors of each media item or group of media items.
[0003] A method for determining the best-matching media content descriptor set involves obtaining a target profile with multiple profile scores for its profile elements. Typically, the target profile is based on a personality scheme (value pairs representing personality traits) that defines many profile elements, including attributes. These profile element values are also referred to as profile scores. Examples of personality schemes include the Myers-Briggs Personality Indicator (MBTI), Ego Equilibrium, the Big Five personality traits (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism (OCEAN)), or the Enneagram. Other schemes are possible for defining personality profile elements.
[0004] A target profile is mapped to a set of target content descriptors having multiple characteristics. This mapping may be performed by applying at least one mapping rule that defines how to calculate the characteristics of the target content descriptor set from the profile score.
[0005] Next, multiple sets of media content descriptors are obtained for one or more media items. Each set of media content descriptors is associated with a media item or group of media items. A set of media content descriptors for a media item (also called the media profile of a media item, or the music profile in the case of a music media item) or group of media items contains many media content descriptors (also called characteristics) that characterize the media item in terms of different aspects. A set of media content descriptors includes, optionally, a semantic descriptor for the media item (group), among other descriptors. A semantic descriptor describes the content of a media item at a high level, such as the genre to which the media item belongs. In this sense, it is possible to classify a media item into one of many semantic classes and indicate with high probability which semantic class the media item belongs to. For example, a semantic descriptor can be represented as a binary value (0 or 1) indicating the class affiliation of a media item, or as a real number indicating the probability that the media belongs to a semantic class. A semantic descriptor may also be an emotional descriptor that indicates that a media item corresponds to an emotional aspect such as atmosphere. An emotional descriptor can classify a media item into one or more emotional classes and indicate with high probability which emotional class the media item belongs to. An emotional descriptor can be represented as a binary value (0 or 1) indicating the class the media item belongs to, or as a real number indicating the probability that the media belongs to an emotional class.
[0006] Media content descriptors may be calculated from identified media items, or they may be read from a database containing media content descriptors that have been pre-analyzed for multiple media items. Thus, the step of obtaining a set of media content descriptors for each of the one or more identified media items may include reading a set of media content descriptors for the media items from the database. The step of obtaining a set of media content descriptors may be performed before the target profile is obtained and mapped to the target content descriptors. Some media content descriptors have a numerical value that quantifies the degree of each semantic descriptor and / or sentiment descriptor present for a media item. For example, a numerical media content descriptor may be normalized to have a value between 0 and 1 or between 0% and 100%.
[0007] This method further includes searching for at least one media content descriptor set that is the best-matching media content descriptor set with respect to the target profile and its target content descriptor. Searching for the best-matching content descriptor set may include comparing the target content descriptor set with multiple media content descriptor sets of media items.
[0008] For further activities related to the target profile, such as receiving information associated with the target profile, at least one identified best-matching media content descriptor set is selected. If the target profile corresponds to a brand or product, it becomes possible to select users or user groups that best match the product or brand in terms of their personality or emotions. If the target profile corresponds to individual users or user groups, media items corresponding to the best-matching media content descriptor set may be selected for playback or recommendation to the user or user group.
[0009] The results of the decision step may be displayed on a computer device or transmitted to a database server. For example, identification information for at least one determined set of media content descriptors is transmitted to the database server. Identification information for at least one determined set of media content descriptors may be used in many use cases, such as determining media users with personality profiles that match a target profile (e.g., media recommendation engines, smart assistants, smart homes, advertising, product targeting, marketing, virtual reality, and games).
[0010] A set of media content descriptors and / or target content descriptors for a media item may further include one or more acoustic descriptors for the media item. The acoustic descriptors (also referred to as acoustic attributes) of a media item may be determined based on an acoustic digital audio analysis of the media item's content. For example, the acoustic analysis may be based on a spectrogram derived for the audio content of the media item. Various techniques may be employed to obtain acoustic descriptors from an audio signal. Examples of acoustic descriptors include tempo (beats per minute), duration, key, mode, presence or absence of rhythm, and energy (of the spectrum).
[0011] A set of media content descriptors for a media item may be determined at least in part on one or more artificial intelligence models that determine one or more emotional descriptors and / or one or more semantic descriptors for the media item. One or more semantic descriptors may include at least one of the following: genre, or vocal attributes such as the presence or absence of vocals, and the gender of the vocals (low-pitched or high-pitched, respectively). Examples of emotional descriptors include the mood of the music and the mood of the rhythm. The artificial intelligence models may be based on machine learning techniques such as deep learning (deep neural networks). For example, an artificial neural network may be used to determine the emotional and semantic descriptors for a media item. The neural network may be trained on a broad set of data provided by music experts and data science experts. It is also possible to use an artificial intelligence model or machine learning technique (e.g., a neural network) to determine the acoustic descriptor (such as bpm or key) of a media item.
[0012] The media item may be segmented, and based on the results of the analysis of individual segments, a set of media content descriptors for the media item may be determined. For example, a media item may be segmented into media item parts, acoustic analysis and / or artificial intelligence techniques may be applied to the individual parts, and the acoustic descriptors and / or semantic descriptors generated for these parts may then be aggregated to form an acoustic descriptor and / or semantic descriptor for the complete media item, similar to how the media content descriptors for a media item are aggregated for the entire group of media items.
[0013] The profile score of a target profile (i.e., attribute values (value pairs of profile elements)) may be used by mapping rules that define how to calculate the properties of a target content descriptor from the profile score. The mapping rules may define which profile scores contribute to the properties of a set of target content descriptors and how they contribute. For example, the properties of a target content descriptor are determined based on a weighted profile score. Based on the weighting, different scores may contribute to the target content descriptor to different degrees. Furthermore, the properties of a target content descriptor may be determined based on the presence or absence of a target profile score. In other words, if a score exists for a given profile element, a contribution to the properties of the content descriptor set may be made. Alternatively, if no profile score is expected to exist, the contribution to the properties of a content descriptor may be represented by weighting the difference between 1 and the score value (which has a value between 0 and 1).
[0014] Mapping rules may be learned using machine learning techniques. For example, the weights that a target profile score contributes to the characteristics of a target content descriptor may be determined by machine learning using a number of target profiles (real-world user profiles) and a suitable machine learning technique capable of determining rules and / or weights for mapping content descriptors to profile scores and vice versa. Such machine learning techniques may also determine which content descriptors can contribute to the profile score and select each content descriptor, and vice versa.
[0015] Searching for the best-matching media content descriptor set may be based on comparing a target content descriptor set with multiple media content descriptor sets. For example, comparing a target content descriptor set with media content descriptor sets may be based on matching elements of content descriptors and selecting content descriptor sets of media items that have identical or similar elements to the target content descriptors. Furthermore, comparing content descriptor sets may be based on similarity searches in which corresponding elements of the target content descriptor set and media content descriptor sets are compared and a matching score value indicating the similarity of each pair of content descriptor sets is calculated. The matching score for a pair of target content descriptor sets and media content descriptor sets may be based on the individual matching scores of the corresponding elements of the content descriptor sets. For example, the difference between corresponding values (scores) of content descriptor sets can be calculated (e.g., Euclidean distance, Manhattan distance, cosine distance, etc.), and from this, matching scores can be calculated for multiple pairs of target content descriptor sets and media content descriptor sets being compared. Furthermore, multiple best-matching media content descriptor sets may be determined and ranked according to their matching scores. This makes it possible to determine the best-matching media content descriptor set, the second-best match, and so on.
[0016] Comparing content descriptor sets may further depend on the context or environment of the user or user group corresponding to the target profile or media content descriptor set. Examples of context or environment include the user's location, date and time, weather, and other people nearby. Similar contexts or examples may be applied to user groups.
[0017] In embodiments addressing the selection of suitable media items for a user or user group, the target profile is based on a personality scheme (value pairs representing personality traits) that corresponds to individual users or user groups and defines many profile elements, including attributes. The values of the profile elements are also referred to as profile scores. Examples of personality schemes include the Myers-Briggs Scale of Inventory (MBTI), Ego Equilibrium, the Big Five personality traits (openness, conscientiousness, extraversion, agreeableness, neuroticism (OCEAN)), or the Enneagram.
[0018] The target profile may be determined based on the user's or user group's playlist or streaming history, which identifies a group of media items associated with the user or user group. Media items may be identified in a playlist by referencing their storage location (e.g., by URL), by listing their name or title (e.g., artist, album, song), or by a unique identifier (e.g., ISRC, MD5 sum, audio identification fingerprint, etc.). The storage location of the corresponding audio / video file may be determined by a table lookup or search procedure.
[0019] Media items identified in a playlist enable the determination of a user's personality profile. If the media items relate to a user's short-term media consumption history (e.g., recently listened-to songs), the generated personality profile characterizes the user's current or recent mood. If the media items relate to a playlist that identifies a user's long-term media item usage history, the generated personality profile characterizes the user's long-term personality profile. In some embodiments, particularly in advertising and branding use cases, a combination of the long-term personality profile and the short-term personality profile (based on the mood of recently listened-to songs) can also be considered as the user's relevant personality profile. In other embodiments, the target profile may be determined based on the short-term media consumption history of the user or user group, characterizing the user's or user group's current mood.
[0020] Further media items corresponding to the best-matching media content descriptor set may be selected for presentation to the user or user group. These media items may be audio or video clips selected for playback or recommendation to the user or user group. Information associated with the media items corresponding to the best-matching media content descriptor set may be provided to the user or a user device associated with the user.
[0021] In other embodiments addressing the identification of suitable users that match a target profile, the target profile corresponds to a product or brand, and each set of media content descriptors is associated with a user or user group. This method may further include determining the user or user group associated with the media content descriptor set that best matches the target content descriptor set. In other words, a user or user group with an emotional profile that best matches the target content descriptor set for a product or brand can be identified and selected for further activity. The target content descriptors may be determined from elements of the product or brand profile by applying mapping rules as described above.
[0022] A set of media content descriptors may include collective characteristics that characterize a group of media items presented to a user or user group, particularly media items identified in a playlist associated with the user or user group. In this regard, a set of collective media content descriptors may be determined for the entirety of one or more media items in a group, based on each media content descriptor of individual media items. The collective media content descriptor characterizes the semantic descriptors and / or emotional descriptors (and optionally, acoustic descriptors) of the media items in the group. A set of collective media content descriptors, including atmosphere and associated with a user or user group, may also be referred to as the emotional profile of the user or user group. A collective media content descriptor can be calculated, in particular for media content descriptors with numerical values, by averaging the values of the individual media content descriptors of the media items. Note that methods other than a simple average of the values of individual media content descriptors are possible. For example, root mean square (RMS) or other techniques (e.g., "log-mean exponential averaging") may be applied to emphasize larger values within a set. For this reason, the step of obtaining a set of content descriptors for multiple media items may include calculating a collective numerical content descriptor from each of the numerical content descriptors for the media items in that group.
[0023] Furthermore, user personality profiles may be provided to users or user groups associated with a media content descriptor set (for example, associated with the sentiment profile of the user or user group). For example, a user personality profile may be provided to a user whose sentiment profile best matches the target content descriptor set. The user personality profile includes multiple personality scores for profile elements that represent the personality traits of the user or user group based on a personality scheme. The personality scheme may be one of the schemes described above. User personality profiles corresponding to the best-matching content descriptor set may be selected, stored, and / or transferred to other computer equipment for further activities.
[0024] A user personality profile can be generated by mapping the characteristics of an associated (collective) set of media content descriptors (e.g., an emotional profile) to a personality score. Depending on the media items corresponding to the set of media content descriptors, the user personality profile may characterize the user's long-term personality or their short-term mood. The personality score of a user personality profile (i.e., attribute values (value pairs of profile elements)) may be determined based on mapping rules that define how to calculate the personality score from a set of (optionally, collective) media content descriptors associated with the user (i.e., the user's emotional profile). The mapping rules may define which (collective) media content descriptors contribute to the personality score and how they contribute. For example, the personality score may be determined based on weighted collective numerical media content descriptors. Based on the weighting, different characteristics may contribute to the score to different degrees. Furthermore, the personality score of a personality profile may be determined based on the presence or absence of (collective) media content descriptors. In other words, if (collective) media content descriptors exist, their contribution to the score can be made, for example, by weighting normalized numerical collectivization characteristics. Alternatively, if (collective) media content descriptors are not expected to exist, their contribution to the score can be represented by weighting the difference between 1 and normalized numerical collectivization characteristic values (which range from 0 to 1). The mapping rules may be learned using machine learning techniques similar to those described above for mapping from target profile scores to target content descriptors.
[0025] Further media items corresponding to the target profile may be selected for presentation to the decision user or user group. These media items may be audio or video clips selected for playback or recommendation to the user or user group. Information associated with the media items corresponding to the best-matching media content descriptor set may be provided to the user or a user device associated with the user.
[0026] An electronic message containing information about a product or brand may be automatically generated for a determined user or user group, and the generated message may be electronically transmitted to the user or user group. The electronic message may also contain information about a product or brand associated with a target personality profile, and for example, the electronic message may also contain additional media items selected in accordance with the target profile.
[0027] Comparing a target content descriptor set with multiple media content descriptor sets to determine at least one media content descriptor set that is the best-matching content descriptor set may be repeated, for example, after a determined period or after many media items have been presented to the user (group). In this way, the determination of further media items can be updated periodically (for example, in real time after media items have been presented to the user, based on the most recently determined media content descriptor set). This enables an adaptive media presentation service in which new media items corresponding to the target profile are presented to the user based on the media items the user has consumed in the past.
[0028] Comparing a target content descriptor set with a plurality of media content descriptor sets and determining at least one media content descriptor set that is the best matching content descriptor set may be generated on a server platform. This method may further include transmitting identification information of consumed media items associated with the user from a user device associated with the user to the server platform. For this reason, the server can receive information regarding the user's media consumption (e.g., a playlist) and determine the user's media content descriptor set and personality profile from this information. As described above, this may be executed repeatedly. The user device may be any user device such as a personal computer, a tablet computer, a mobile computer, a smartphone, a wearable device, a smart speaker, a smart home environment, a car radio, etc., or any combination thereof. After determining the best matching content descriptor set, the server can also transmit selected media items corresponding to the target profile to the user device. This information is either received and presented to the user or the selected media item is played.
[0029] In another aspect of the present disclosure, a computer device for executing any of the above methods is proposed. This computer device may be a server computer having a memory for storing instructions and a processor for executing the instructions. The computer device may further include a network interface for communicating with a user device. The computer device may receive information regarding media items consumed by a user from the user device. The computer device may be configured to generate a media content descriptor set and a personality profile as described in the above disclosure. Depending on the use case, the media content descriptor set and / or the personality profile may be used for recommendations of similar media items, determination of media users having a personality profile matching a target profile, or display of advertisements that a media user is likely to be interested in. Information regarding further identified media items corresponding to the target profile may be transmitted to the user device.
[0030] Embodiments of the disclosed device may include, but are not limited to, the use of one or more processors, one or more application specific integrated circuits (ASICs), and / or one or more field programmable gate arrays (FPGAs). Embodiments of the device may also include the use of other conventional hardware and / or customized hardware, such as software programmable processors, such as a graphics processing unit (GPU) processor.
[0031] Another aspect of the present disclosure may relate to computer software, a computer program product, or any medium or data embodying computer software instructions for causing at least one processor to execute at least one of the method steps described in the present disclosure by executing on a programmable computer or dedicated hardware having at least one processor.
[0032] While several exemplary embodiments are described herein with particular reference to the above-mentioned uses, it will be understood that this disclosure is not limited to such articulations but is applicable in a broader context.
[0033] In particular, the methods relating to this disclosure concern methods for operating the apparatus relating to the exemplary embodiments and their variations described above, and each description of the apparatus is similar to the corresponding method, and vice versa. It should be understood that similar descriptions may be omitted for simplification. Furthermore, the above embodiments can be combined in many ways, even if not explicitly disclosed. Those skilled in the art will understand that these embodiments and features / steps can be combined in any way, as long as they do not result in a contradiction that is explicitly excluded.
[0034] Other and further exemplary embodiments of this disclosure will become apparent in the course of the following discussion by referring to the attached drawings. [Brief explanation of the drawing]
[0035] The following describes exemplary embodiments of this disclosure with reference to the attached drawings, but these are only a few examples. [Figure 1] This figure schematically illustrates the operation of one embodiment of the present disclosure. [Figure 2a] This diagram illustrates the generation of semantic descriptors from audio files. [Figure 2b] This figure illustrates the generation of semantic descriptors by the audio content analysis unit. [Figure 3a] This diagram shows the mapping of atmosphere content descriptors to the EI (extroversion-introversion) personality score in the MBTI personality scheme. [Figure 3b] This figure shows the mapping of atmosphere content descriptors to the openness personality score in the OCEAN personality scheme. [Figure 4a] This figure shows an example of a graphical representation of a personality profile in the MBTI Personality Scheme. [Figure 4b] This figure shows an example of a graphic representation of a personality profile in the OCEAN Personality Scheme. [Figure 5] This figure illustrates one embodiment of a method for determining a user or user group. [Figure 6] This diagram shows the mapping from product profile attributes to atmosphere content descriptors. [Modes for carrying out the invention]
[0036] In one broad aspect of this disclosure, a personality profiling engine generates personality or sentiment profiles corresponding to media items being analyzed, such as music, thereby determining the characteristics of a media item. This enables a variety of new applications (also referred to in this disclosure as “Use Cases”) that allow for the classification, searching, recommendation, and targeting of media items or media users. For example, personality profiling or sentiment profiling may be used to recommend similar media items or to display advertisements that media users may be interested in.
[0037] For example, if the input to a personality profiling engine is a user's short-term music listening history, it can determine a personality profile that characterizes the listener's mood based on the songs the user has recently played. If the input is a long-term music listening history, it can determine the listener's general personality profile. It is also possible to calculate the difference between the user's long-term personality profile and their current mood to determine if the user is in an exceptional situation.
[0038] Personality profiles generated by a personality profiling engine can detect emotional signatures of music listeners, focusing on the atmosphere, mood, and values that define a person's multifaceted personality. This allows for addressing questions such as whether a listener is self-conscious or haughty, or whether they enjoy exercise or travel.
[0039] In the case of audio, for example, it is possible to find similar-sounding music tracks based on the emotional descriptor and / or semantic descriptor of an audio file. A media similarity engine using the generated emotional profile may utilize machine learning or artificial intelligence (AI) to match and find musically and / or emotionally similar tracks. Such a media similarity engine can listen to and understand music like a human, then search millions of music tracks for specific sonic or emotional patterns, finding the required music in seconds by matching requirements. Based on the generated profile, it is possible to search for, for example, only instrumental or vocal tracks, or according to other semantic criteria such as genre, tempo, mood, or bass vs. treble.
[0040] The basis of the proposed technology is a personality profiling engine that performs tagging of media items using media content descriptors based on audio analysis and / or artificial intelligence (e.g., deep learning algorithms, neural networks, etc.). The personality profiling engine may tag media tracks by weighted mood, emotion, and musical attributes such as genre, key, and tempo (beats per minute (bpm)) by enriching metadata using AI. The personality profiling engine may analyze the mood, genre, acoustic attributes, and contextual situation of a media item (e.g., a musical track (song)) and obtain weighted values for different "tags" within these categories. The personality profiling engine may analyze a media catalog and tag each media item in the catalog with corresponding metadata. Media items are, for example, • Acoustic attributes (bpm, key, energy, etc.) • Atmosphere / rhythm atmosphere, Genre, • Vocal attributes (instrumental, high notes, low notes), Contextual situation, It may also be tagged with a media content descriptor related to it.
[0041] In the mood category for tagging music from an "emotional" perspective, the personality profiling engine may output up to 35 "complex mood" values that are taxonomically classifiable within, for example, 18 mood subfamilies organized as 6 main families. All 6 main families and 18 subfamilies include human emotions. The level of detail in the mood taxonomy can be arbitrarily narrowed down; that is, the 35 "complex moods" can be further subdivided as needed, or additional "complex moods" can be added.
[0042] Figure 1 schematically illustrates the operation of one embodiment of the present disclosure, which generates personality profiles and determines the similarity of profiles to provide various recommendations for similar media items, or to match users or user groups. The personality profiling engine 10 receives one or more media files 21 from the media database 20. Media files are identified in the media list 30 provided to the personality profiling engine 10 in order to read media items from the database 20. The media list 30 may be the user's playlist read from a playlist database that stores the most recent media items played by the user and user-defined playlists representing the user's media preferences.
[0043] Analysis of the media file 21 determines media content descriptors 43, including the acoustic descriptor, semantic descriptor, and / or emotional descriptor of the audio content. An audio content analysis unit 40, equipped with an acoustic analysis unit 41 that analyzes the acoustic features of the audio content, determines several media content descriptors 43 by analyzing the time-frequency plane in a way that generates a frequency-domain representation, such as a spectrogram, of the audio content and calculates acoustic features such as tempo (bpm) or key. The spectrogram may be transformed according to perspective and / or logarithmic scales, for example, in the form of a Log-Mel spectrogram. The media content descriptors may be stored in a media content descriptor database 44.
[0044] The audio content analysis unit 40 of the personality profiling engine 10 further comprises an artificial intelligence unit 42 that uses an artificial intelligence model to determine media content descriptors 43, such as emotion descriptors and / or semantic descriptors, of the audio content. The artificial intelligence unit 42 may operate on any appropriate representation of the audio content, such as a time-domain representation generated by the acoustic analysis unit 41, a frequency-domain representation of the audio content (e.g., a Log-Mel spectrogram as described above), or intermediate characteristics derived from the audio waveform and / or frequency-domain representation. The artificial intelligence unit 42 may generate, for example, an atmosphere descriptor of the audio content that characterizes the mood of the music and / or rhythm of the audio content. These AI models may be trained on their own large-scale expert data.
[0045] Figure 2a shows an example of semantic descriptor generation from an audio file by the audio content analysis unit. In this embodiment, audio file samples are optionally segmented into audio chunks and converted into frequency representations such as Log-Mel spectrograms. The audio content analysis unit 40 then extracts low-level, mid-level, and / or high-level semantic descriptors from the spectrograms by applying various audio analysis techniques.
[0046] Figure 2b further illustrates an example of semantic descriptor generation by the audio content analysis unit 40. While Figure 2a shows direct audio content analysis using conventional signal processing methods, Figure 2b shows audio content analysis using a neural network, which first requires learning from "ground truth" data ("prior knowledge"). The audio file is converted into a spectrogram, and media content descriptors 43, such as the atmosphere, genre, and context of the audio file, are generated by applying one or more neural networks. The neural networks are trained for this task based on large-scale expert data (large-scale and detailed "ground truth" media annotations for supervised neural network training). In an example of semantic descriptor generation by the artificial intelligence unit 42, spectrogram data of the audio file is fed to the neural network as input, which generates semantic descriptors as output. In embodiments, one or more convolutional neural networks are used to generate descriptors for, for example, genre, rhythm atmosphere, and speech family. Other network configurations and combinations of networks can be used similarly.
[0047] The mapping unit 50 maps media content descriptors 43 for audio files to media personality profiles 61 by applying mapping rules 51 received from the mapping rule database 52. The mapping rules 51 may define the media content descriptors used in calculating the profile score (i.e., the values of the profile attributes) and the weights applied to the media content descriptors. The mapping rules 51 may also be represented as a matrix that links media content descriptors and profile attributes to assign weights to the media content descriptors. The generated personality profiles 61 may be provided to the media similarity engine 70 to determine similar profiles, or they may be stored in the profile database 60 for later use.
[0048] The mapping unit 50 may also operate in reverse to map the personality profile 61 to the media content descriptor 43 using mapping rules 51 from the mapping rule database 52. The mapping rules 51 may define the profile scores (i.e., the values of the profile attributes) used to calculate the properties of the media content descriptor set and the weights to apply to the profile scores in the mapping. The mapping rules 51 may be represented as a matrix that links the profile attributes and media content descriptors to assign weights to the profile scores. For example, the mapping rules may map a target profile defined in the personality scheme to a target content descriptor set.
[0049] When a personality profile is generated for a media item group, media content descriptors 43 are generated for each media item in that group (or read from the media content descriptor database 44), and a collective media content descriptor is generated for the entire group of media items. The collective nature of the numerical media content descriptors can be achieved by calculating the average value of each media content descriptor for the group of media items. Other collective nature algorithms, such as the root mean square (RMS), can also be used. The mapping unit 50 then acts on the collective media content descriptor (e.g., sentiment profile) to generate a personality profile for the entire group of media items.
[0050] The media similarity engine 70 can receive profiles directly from the personality profiling engine 10, or from the profile database 60, as shown in Figure 1. The media similarity engine 70 determines profile similarity based on profile element matching or similarity search, as disclosed below, through profile comparison. If a profile 71 similar to the target profile is determined, corresponding media items or users may be determined, and recommendations may be made. For example, one or more media items matching the user's playlist may be determined and automatically played on the user's terminal device. Other use cases are described in this disclosure.
[0051] The media similarity engine 70 may also perform a search in the content descriptor region. That is, by comparing content descriptor sets, similar content descriptor sets may be determined by matching each characteristic of the content descriptor sets. For example, based on the target profile and mapping rules as described above, the conditions for characteristics in a (collected) media content descriptor set (e.g., sentiment profile) are determined (e.g., whether a particular semantic descriptor or sentiment descriptor exists, has a specific value, or is within a specific range), and by applying these to the media content descriptor sets, matching media content descriptor sets, i.e., media content descriptor sets that satisfy these conditions, are identified.
[0052] Alternatively, this search may be based on a similarity search that calculates matching scores for multiple pairs of content descriptor sets (for example, based on distance criteria such as Euclidean distance), similar to the similarity search in the profile domain described above. For example, a matching score is calculated for each (aggregated) media content descriptor set with respect to the target content descriptor set, and the media content descriptor sets are ranked according to their matching scores. This makes it possible to determine the best-matching media content descriptor set.
[0053] As mentioned above, the personality profiling engine can use machine learning or deep learning techniques to determine the emotional and semantic descriptors of media items. Training may be based on a database consisting of many data points to learn relationships and analyze an individual's music preferences and listening habits. This algorithm can read the user's psychological and emotional portrait and complement existing demographic and behavioral statistics to generate a complete and evolutionary user profile. The output of the personality profiling engine is a profile of the psychologically motivated user ("personality profile") derived from an analysis of the user's music (playlists or listening history).
[0054] A personality profiling engine can derive a user's personality profile from fewer or more media items. For example, based on the most recent 10 or more music items played by the user on a streaming service, the engine can calculate the user's short-term ("instantaneous") profile (reflecting the music listener's "current mood"). If more music items represent the user's long-term listening history or preferred playlists, the engine can calculate the user's intrinsic personality profile.
[0055] The personality profiling engine uses advanced machine learning and deep learning techniques to understand the meaningful content of music from audio signals, achieving a human-like level of comparison beyond simple text language and labels. By extracting essential information about music from audio signals, the algorithm can learn to understand the rhythm, beat, style, genre, and mood of the music. The generated profiles can be applied to music or video streaming services, digital or linear radio, advertising, product targeting, computer games, labels, libraries, publishers, in-store music providers or synchronization agencies, voice assistants / smart assistants, smart homes, and more.
[0056] The personality profiling engine can achieve a human-like level of comparison by applying advanced deep learning techniques to understand the meaningful content of music from audio. This algorithm can analyze and predict relevant mood, genre, context, and other key attributes and assign a weighted relevance score (%).
[0057] The media similarity engine is applicable to recommendation, music targeting, and audio branding tasks and can be used by music or video streaming services, digital or linear radio, fast-moving consumer goods (FMCG) (also known as consumer goods products (CPG)), advertisers, creative agencies, dating companies, in-store music providers, or e-commerce.
[0058] The personality engine may be configured to generate a personality profile based on a group of media items associated with a user by performing the following steps: In the first step, a group list is obtained containing identification information for one or more media items, for example, in the form of a user-defined playlist. Next, a set of media content descriptors is generated for each of the one or more identified media items in the group, or read from a database of previously analyzed media items. The set of media content descriptors includes at least one of the following for each media item: an acoustic descriptor, a semantic descriptor, and an emotional descriptor. This method then includes determining a set of aggregated media content descriptors (i.e., the user's emotional profile) for the entire group of the one or more identified media items, based on each media content descriptor of the individual media items. Finally, the aggregated set of media content descriptors is mapped to the personality profile for the group of media items. The score of the profile elements is calculated from the aggregate properties of the set of aggregated media content descriptors.
[0059] In an exemplary embodiment, the personality profiling engine is applied to determining the mood of a media user. For example, the mood of a music listener is determined based on the input "short-term music listening history." Alternatively, the general personality profile of a music listener is determined from the input "long-term music listening history." In further use cases, an individual's personality profile may be associated with the personality profiles of others to determine individuals with similar profiles at a particular moment (e.g., matching, recommendations for products with similar profiles (e.g., e-commerce), or suggestions for connecting with others (friends, dating, social networks, etc.)).
[0060] Furthermore, by using a personality profiling engine, media items such as music (e.g., current playlists and / or suggestions, or other forms of entertainment (movies, etc.), or environments such as smart homes) may be adapted to a) an individual's current mood, and / or b) an intention to change the individual's mood (intentions explicitly expressed by the individual, implicit intentions of change brought about by the system, e.g., product recommendations or optimization (increase) of user retention on the platform).
[0061] A personality profiling engine can be used to determine the difference between a user's current mood and their general personality by calculating the difference between a user's long-term personality profile and their current (mood) profile. This is useful, for example, for adapting recommendations in the case of a short-term "deviation" from the user's general personality profile that leans towards a particular musical direction (depending on a specific listening context, time of day, user mood, etc.), and for deciding whether to display ads that would normally fit the user's personality profile but are momentarily inappropriate because the current mood profile deviates from the current listening situation. In either case, the placement of recommendations or ads may be adapted to the user's personal situation at that moment.
[0062] The basis of these embodiments is a personality profiling engine that analyzes groups of media items identified by a provided list. For example, audio tracks in a group of songs (from digital audio files) are analyzed. This analysis may be performed by, for example, audio content analysis and / or the application of machine learning (e.g., deep learning) methods. The personality profiling engine is Algorithms may be applied to extract low-level, mid-level, and high-level characteristics from audio. Examples of low-level characteristics are audio waveform / spectrogram-related characteristics (or "descriptors"), examples of mid-level characteristics (or "descriptors") are "fluctuations," "energy," etc., and examples of high-level characteristics are semantic descriptors and emotional descriptors such as genre, mood, or key. Acoustic waveform and spectrogram analysis may be applied to analyze acoustic attributes such as tempo (beats per minute), key, mode, duration, spectral energy, and presence or absence of rhythm. Neural network / deep learning-based models may be applied to analyze high-level descriptors from audio input (e.g., via Log-Mel frequency spectrograms extracted from various segments of an audio track), such as genre, mood, rhythmic mood, presence of sound (instrumental or vocal), and vocal attributes (e.g., bass or treble). The neural network / deep learning models may be trained on a large training dataset containing thousands (hundreds) of annotated examples of the aforementioned categories tagged by music experts. For example, a deep learning convolutional neural network may be used, but alternatively, other types of neural networks (e.g., recurrent neural networks), other machine learning techniques, or any combination thereof may be used. In embodiments, one model is trained for each category group of mood, genre, rhythmic mood, and presence of sound / vocal attributes. In an alternative example, one common model may be trained collectively, or one model may be trained together for multiple moods and rhythmic moods, or one model may be trained for each individual mood or genre itself.
[0063] Audio analysis may be performed on multiple time points within the audio file (for example, three 15-second sections at the beginning, middle, and end of a song), or it may be performed on the entire audio file.
[0064] The output may be stored at the segment level or at the audio track (song) level (e.g., aggregated from segments). Subsequent steps may also be applied at the segment level (e.g., obtaining a list of moods (or mood scores) for each segment (e.g., applicable to long audio recordings such as classical, DJ mixes, or podcasts, or to audio tracks where genre or mood changes)). The personality profiling engine may store all derived song content descriptors, along with their predicted or percentage values, in one or more databases for further use (see below).
[0065] The output of the audio content analysis is: Tempo: For example, 135 bpm, • Key and mode: For example, F# minor, • Spectral energy: For example, 67% (100% is determined by the maximum value listed in the truck's catalog), • Presence of rhythm: For example, 55% (100% is determined by the maximum value in the track catalog), • Genre: As a list of categories (each independent of others with a percentage value from 0 to 100), for example, Pop 80%, New Wave 60%, Electropop 33%, Dance Pop 25%, • Atmosphere: A list of atmospheres included in the song (each with a percentage value from 0 to 100, independent of others), for example, dreamy 70%, intellectual 60%, inspiring 40%, bitter 16%, • Rhythmic atmosphere: A list of atmospheres included in the song (each with a percentage value from 0 to 100, independent of others), for example, elegant 67%, lyrical 53%, • Vocal attribute: Instrument (0 or 100%) or any combination of low and / or high frequencies (50-100%) These are media (e.g., music) content descriptors derived from input audio (also referred to as audio characteristics or music characteristics).
[0066] In one embodiment, audio content analysis is performed. From the audio characteristics extraction, 14 mid-to-high level characteristics + 52 low-level (spectral) characteristics were obtained, From the deep learning model, 67 genres, 35 moods (+24 (by grouping into subfamilies and families) (see below)), 5 rhythmic moods, and 3 vocal attributes were derived. Output each of these.
[0067] Optionally, subsequent post-processing of the values is performed by applying so-called adjustment coefficients (e.g., assigning high or low weights to genre, mood, or some other category). The adjustment coefficients adapt the machine predictions to more closely resemble human perception. The adjustment coefficients may be determined by experts (e.g., music experts) or learned by machine learning. The adjustment coefficients may be defined by one coefficient per semantic descriptor or sentiment descriptor, or by a non-linear mapping from various machine predictions to adjusted output values.
[0068] Furthermore, optionally, the aggregation of music content descriptors may generate values for groups or "families" of music content descriptors that typically follow a taxonomic framework. In one example, 35 atmospheres predicted by a deep learning model are aggregated into 18 parent "subfamilies" and 6 "main families," forming a total of 59 atmospheres (following a taxonomic framework for atmospheres).
[0069] The analysis may be performed at the song level for a set of songs and delivered in audio form (various digital formats, compressed or uncompressed). For the generation of a personality profile, for a group of songs (commonly referred to as a "playlist"), the song content descriptors and their values for multiple songs may be aggregated.
[0070] In some embodiments (use cases), the listener's current mood is determined. In other use cases, the personality profiling engine determines the listener's long-term personality profile. In both cases, the input is a list of songs, and the output is the user's personality profile (according to one or more personality profiling schemes). To determine the listener's mood, the input consists of the most recently listened songs. These songs provide insight into the user's current mood profile. To determine the listener's general (long-term) personality profile, the input consists of songs (usually a larger set) representing the user's (long-term) history.
[0071] The generation of a personality profile may be based on the characteristics of the music the user listens to, including (not limited to) mood, genre, vocal presence, vocal attributes, key, BPM, energy, and other acoustic attributes (= "music content descriptor," "audio characteristics," or "music characteristics"). This may be determined for each characteristic of the song's musical content.
[0072] In one embodiment, the musical content descriptors of n songs are aggregated to generate an aggregated content descriptor, that is, the user's emotional profile is generated, for example, as the average of the numerical values (%) of each song in the aggregate (playlist). Alternatively, more complex aggregation procedures such as median, geometric mean, RMS (root mean square), or various forms of weighted average may be applied.
[0073] In one embodiment, preliminary analysis of songs in the user's playlist or listening history may extract song content descriptors that may contain numerical values (for example, ranging from 0 to 100% for each value). For each content descriptor (for example, mood "sensibility"), the root mean square (RMS) of all "sensibility" values for individual songs may be calculated and stored. The output of this aggregation is a set of song content descriptors having the same number of descriptors (attributes) as each song. This aggregated song content descriptor (sensibility profile) is used in the second stage of the personality profiling engine to determine the user's personality profile.
[0074] In some embodiments, instead of a user's playlist, an album or artist's discography (all tracks by an artist) can be used as input for aggregation. Similarly, aggregation of the musical content descriptors can be performed for many tracks (which may represent albums, artists, or playlists) (by using the various methods disclosed above).
[0075] When the aggregated value of each song content descriptor is calculated, a personality profile is generated. For example, a mapping is performed from elements in an emotional profile (representing aggregated song content descriptors for n songs) to one or more personality profiles. In the mapping, mood, genre, style, etc., are converted into psychological and emotional user characteristics (personality traits). The mapping is performed from the song content descriptors to a score of a personality profile (including personality traits / human characteristics). Rules for mapping song content descriptors and their values to one or more types of personality profiles defined by the personality profile scheme may be defined.
[0076] The output of the personality profiling engine is a range of numerical output parameters, referred to as personality profile attributes and scores, which represent the user's personality profile.
[0077] Personality profiles are, • MBTI (Myers-Briggs Index) ·Ego Equilibrium, • OCEAN (also known as one of the Big Five personality traits) • Enneagram, These may be defined according to various personality profiling schemes.
[0078] Each of these personality profile schemes is composed of personality attributes (for example, "extraversion" or "openness" and assigned scores (values) such as 51% or 88% (specific examples are shown below)).
[0079] For all of these schemes, a mapping from musical content descriptors to profile scores may be used, and vice versa. Figure 3a shows the mapping of atmosphere content descriptors to EI personality scores in the MBTI personality scheme. The mapping may apply a matrix such as the example shown in Figure 3a. The calculation of scores (values) within the personality profile scheme may involve the presence (%) or absence (100%) of (atmosphere or other musical content descriptors).
[0080] Each scheme may have many "scores" to calculate (for example, the MBTI scheme calculates four scores: EI, SN, TF, and JP). One or more mapping rules may be defined for each score, which affect how the score is calculated from the aggregated music content descriptor. For example, the score is equal to the sum of the values calculated by the matrix divided by the number of values to consider (i.e., the usual averaging mechanism).
[0081] For example, the "regressive" atmosphere (included in the music content descriptor) is used in the EI calculation as part of the MBTI scheme. Figure 3a shows an example of a rule matrix applied to the EI calculation from the atmosphere portion of the music content descriptor. The rule matrix shows how the presence or absence of atmosphere can be used in the calculation of the EI score. Other music content descriptors can be similarly included in the calculation.
[0082] In one embodiment, the EI calculation includes 17 rules incorporating 17 values from the music content descriptor. These rules follow a psychological recipe. For example, the rules in the "metal" group psychologically define "closed shoulder," while the rules in the "tree" group define "open shoulder."
[0083] Similar calculations can be performed on other profiling matrices such as OCEAN.
[0084] As mentioned above, the MBTI personality profile has scores for EI, TF, JP, and SN. The following is an example of how the MBTI personality profile and its scores are expressed. “mbti”:{“name”:“INTJ”,“sources”:{ “EI”:33.66403316629877, “SN”:42.419498057065084, “TF”:57.82423612828757, "JP":61.02633025243475}}
[0085] A basic score classification may be performed based on the score value. This classification may be based on a comparison of the score value with a specific threshold. For example, the EI score in the MBTI scheme represents the balance between a user's extraversion (E) and introversion (I). An EI below 50% indicates introversion, while an EI above 50% indicates extraversion. Therefore, if EI < 50%, the user may be assigned to the I (introverted) class, and otherwise, to the E (extroverted) class. Other MBTI scores can be classified similarly.
[0086] The scores are defined as opposites on each axis (EI, SN, TF, JP). For each character pair, the value that determines which trait a person possesses is either <50% or >50%. To exclude characters from the above example, typically, the right-hand character of the character pair is taken if the value is <50%, and the left-hand character is taken if the value is ≥50%.
[0087] The scores for the generated profiles may be further classified into general personality types, for example, based on the basic classification results of the profile scores. For example, the following general personality types may be derived from the basic score classification results. ·ESTJ: Extraversion (E), Sensitivity (S), Thinking (T), Judgment (J) INFP: Introversion (I), Intuition (N), Mood (F), Perception (P)
[0088] The profile in the example above is classified as an INTJ personality type. By classifying the four-dimensional space of profile scores (EI, TF, JP, SN) into personality types, it becomes possible to arrange personality traits in two dimensions within a square that has significant representation.
[0089] Figure 4a shows a schematic representation of a personality profile related to the MBTI scheme, and the classification result (INTJ) for the user's profile can be shown in color. This figure intuitively represents the user's profile along various psychological dimensions. An individual classified as "INTJ" is interpreted as a "planner, scientist." Additional personality traits associated with this MBTI type may be output on the user interface.
[0090] In the OCEAN Personality Profile Scheme, scores are assigned to the "Big Five" mindset, including openness, conscientiousness, extraversion, agreeableness, and neuroticism. Figure 3b shows the mapping of the Atmosphere Content Descriptor to the Openness Personality Score in the OCEAN Personality Scheme. The following is an example of how the OCEAN Personality Profile and its scores are represented. “ocean”:{ “agreeableness”:51.10149671582637, “conscientiousness”:73.42223321884429, “extraversion”:33.66403316629877, “neuroticism”:50.21693055551433, “openness”:39.72017677623826}
[0091] Figure 4b shows a schematic representation of the personality profile related to the OCEAN scheme. This figure intuitively represents the user's profile along various psychological dimensions.
[0092] In some embodiments, the personality profile can optionally be enriched with or associated with additional person-related parameters characterized by additional information (e.g., age, gender, and / or biosignals of the human body via body sensors (smartwatches, sports tracking devices, emotion sensors, etc.)). Alternatively, the personality profile can also optionally be enriched with or associated with additional parameters characterizing the individual's context and environment (location, date and time, weather, other people nearby).
[0093] In this embodiment, the personality profiling engine and the media similarity engine are configured to select the most suitable music for a given target user group. In this embodiment, the target group is defined, and the media similarity engine selects matching music for, for example, broadcast. This makes it possible to suggest music for, for example, an advertising campaign for a brand defined by that target consumer group. Further possible use cases include in-store music, advertisements, etc.
[0094] In these embodiments, as described above, a target group of people is identified by one or more personality profiles following schemes such as MBTI, OCEAN, Enneagram, or Ego-Equilibrium (intended to find music suitable for the target group in the case of music consumption, in-store music, advertising campaigns, and other use cases). Demographic parameters for the target group may also be added.
[0095] In the personality profile space between the target group profile and the "song personality profiles" of individual songs (i.e., a set of content descriptors for songs mapped to personality profiles according to a personality scheme), a search (e.g., similarity search or exact score matching) can be performed. Subsequently, the "song personality profile" from the songs that best match the target group personality profile is identified. In this regard, personality profile scores for various personality profile schemes may be pre-calculated for the candidate songs. The best match for the target group of people is then found by similarity search between the profile score of the defined target group and the personality profile score of each song.
[0096] Alternatively, a media similarity engine can find songs relevant to a target group of people by mapping personality profile schemes to song content descriptors. For this purpose, a mapping from target group personality profiles to song content descriptors ("reverse mapping") may be performed, followed by a search for songs that match the target profiles within the song content descriptor space. In this case, the reverse mapping from target group personality profiles to target song content descriptors is performed first, and then songs that best match these target content descriptors are selected.
[0097] In either case, the output is a list of media items (e.g., music tracks) that match the specified target group.
[0098] The term "similarity search" encompasses a broad range of mechanisms for searching a vast space of objects (in this case, profiles) based on the similarity between any pair of objects (e.g., profiles). Nearest neighbor search and range queries are examples of similarity search. Similarity search can rely on mathematical concepts of distance spaces, which allow for the construction of efficient index structures to achieve scalability in the search domain. Alternatively, non-distance spaces such as the Kullback-Leibler divergence or embeddings learned by neural networks may be used for similarity search. Nearest neighbor search is a form of proximity search and can be expressed as an optimization problem to find the point in a given set that is closest to (or most similar to) a given point. Closeness is usually represented by a dissimilarity function, where the lower the similarity of the objects, the larger the value of the dissimilarity function. In this case, the (dis)similarity of the content descriptor set is the criterion for the search.
[0099] As mentioned above, the search for the best-matching media items for the target group may be performed in the content descriptor set area by comparing the target content descriptor set with the content descriptor set of the media items (for example, the content descriptor set for individual songs). As mentioned above, the target content descriptor set may be derived from the target profile or the product or brand profile. This search is performed • Matching of content descriptor set characteristics (depending on whether or not the content descriptor set has characteristics), • Matching the values of the characteristics of the content descriptor set (numerical search), • Searching for a range of such characteristics (for example, a particular atmosphere is 75% to 100%), • Vector-based matching and similarity calculations (calculating how "close" (similar in terms of numerical distance) the values of the target content descriptor set and the media content descriptor set are by comparing their numerical properties (e.g., using distance measures such as Euclidean distance, Manhattan distance, cosine distance, or other methods such as Kullback-Leibler divergence)) • Learning similarity based on machine learning (a machine learning or deep learning algorithm learns a similarity function based on examples given to the algorithm (in one embodiment, this learned similarity function can then be used permanently)), It may be performed by [this method].
[0100] In other embodiments, the media similarity engine uses one or more of the user's personality profile, the user's current situation or context, and the user's current mood, • Real-time recommendations of songs on online streaming platforms. • Suggesting songs on mobile device applications, and / or • Automatic playback of music based on profile (lean-back radio) You may choose to do so.
[0101] For example, as mentioned above, a personality profiling engine analyzes the user's listening history. In this way, the user's personality profile and / or the listener's emotional profile (including the listener's mood) is determined. Next, similar to determining the target group for a particular song, the media similarity engine may be configured to determine and find the song best suited to the individual (user) based on the individual's (long-term) personal music listening history, personality profile, (short-term) mood profile and / or personality profile, a weighted blend of short-term and long-term personality profiles, and, optionally, the user's context and environment information. The individual's context and environment can be determined by other numerical factors measured from mobile devices or other personal user devices from which location data, weather data, movement data, biosignal data, etc. can be derived. This may be performed instantaneously while the user is listening in a listening session. For example, based on a preliminary analysis of songs the user has listened to previously and songs related to music content descriptors, a song that best matches the user's personality profile is selected. Therefore, as explained above, the user's personality profile is mapped to a target content descriptor set and then compared to a media content descriptor set. For example, a similarity search is performed between the target content descriptor set and the media content descriptor set to determine the best-matching media content descriptor set (and corresponding media items) (which may be ranked according to their matching scores). The output is a list of songs suggested for listening, which can be updated in real time based on new inputs such as updated listening history.
[0102] Similarly, optionally, a set of song content descriptors (as described above) may be compiled from a set of songs (e.g., an album, a playlist, or a set of songs by the same artist) and compared with a target content descriptor set to recommend an artist, album, or playlist to the listener instead of individual songs.
[0103] This system may output a list of tracks, artists, or albums that are matched to an individual's personality profile, along with a matching score (a value indicating the degree of matching for each output item). The matching score may be calculated using a similarity search as described above.
[0104] In other embodiments, the personality profiling engine and media similarity engine are configured to determine a user or group of users for a specific target personality profile, select the best-matching user (user group) for the target profile, and receive further information associated with the target profile. The personality profiling engine may analyze one or more media items associated with the user or user group for their content in terms of acoustic attributes, genre, style, atmosphere, etc. It then generates a description of the user or user group (in the form of a personality profile). After determining a personality profile for each of several users or user groups, the personality profiling engine compares the personality profiles of the several users or user groups to the target personality profile and determines at least one user or user group that has the best-matching personality profile for the target profile. The target profile may be specified by a personality profile following a personality profile scheme such as MBTI, OCEAN, Enneagram, or Ego-Equilibrium, similar to the definition of a user profile. The profile may be optionally enriched with personal-related parameters (e.g., age, gender, etc.).
[0105] More specifically, a song content descriptor, including a semantic descriptor and / or emotional descriptor, is derived through an analysis of the audio in a set of songs associated with the user. Optionally, the descriptors are aggregated (using various methods) for many tracks (which may represent an album or artist), and the user's emotional profile is determined, for example, by calculating the average of the mood and / or other descriptors of multiple songs (possibly using an average, RMS, or weighted average). Then, as described above, a mapping from the song content descriptor to the personality profile is performed. The system then outputs and stores profiles for multiple users, defined by one of several different personality profile schemes. The profiles may be provided in numerical format (e.g., floating-point numbers for different profile scores within the aforementioned schemes).
[0106] In an embodiment, the media similarity engine may use one or more of the user's personality profile, the user's current situation or context, and the user's current mood to search for a user that best matches the target profile. For example, as described above, the personality profiling engine analyzes the user's listening history. In this way, the user's personality profile and / or the music listener's emotional profile (including the listener's mood) are determined. The media similarity engine may then be configured to determine and find the user best suited to the target profile based on the individual's (long-term) personal music listening history, personality profile, (short-term) mood profile and / or personality profile, a weighted mixture of short-term and long-term personality profiles, and, optionally, the user's context and environment information. The individual's context and environment can be determined by other numerical factors measured from a mobile device or other personal user device from which location data, weather data, movement data, biosignal data, etc. can be derived. This may be performed instantaneously while the user is listening in a listening session.
[0107] Alternatively, a search (e.g., similarity search or exact score matching) can be performed in the content descriptor set space by comparing the target content descriptor set with media content descriptor sets corresponding to individual users (which may be aggregated, i.e., sentiment profiles). The media content descriptor set that best matches the target content descriptor set is then identified. Here again, the term “similarity search” includes a broad mechanism for searching a vast space of subjects (e.g., profiles or content descriptor sets) based on the similarity between any pair of subjects (e.g., profiles or content descriptor sets).
[0108] In one embodiment, the media similarity engine is configured to select individuals (users) that match the product or personality profile of a target group. In this embodiment, an advertising client or brand using the disclosure system first defines the target group by setting score values within a specific personality profile scheme (such as MBTI, OCEAN, Enneagram, or Ego-Equilibrium), or provides a target personality profile by defining a brand / product profile that describes the attributes of the brand or product psychologically, emotionally, or marketing-wise.
[0109] This system may have already profiled some users by (preliminarily) analyzing their music preferences (listening history, favorite tracks / albums / artists). The media similarity engine then finds individuals that match a given product profile or target group profile. The output is a list of users suitable for the brand / product profile or specified target group (e.g., by user ID). Identified individuals can then be targeted by specific advertisements.
[0110] In this embodiment, selection based on matching an individual's (user's) personality profile with the target group's personality profile is based on the definition of the target group. The target group may be defined by setting the target profile values within a specific personality profile scheme (such as MBTI, OCEAN, Enneagram, Ego-Equilibrium, etc.).
[0111] Mapping individuals within the target group may be based on a similarity search of the (defined) personality profiles of the target group against a set of user personality profiles (pre-calculated and determined, for example, based on their viewing habits). The comparison of profiles and generation of matching scores via similarity search are as described above. Individuals corresponding to personality profiles may be ranked based on the matching scores of those profiles. A matching score threshold may be applied to select the best-matching group for each individual.
[0112] As mentioned above, matching users can be searched in the content descriptor set area by comparing the target content descriptor set (generated by mapping target group profiles to the content descriptor set area using mapping rules) with the media content descriptor sets corresponding to individual users (i.e., their sentiment profiles), and selecting the best-matching media content descriptor set (and corresponding user).
[0113] In other embodiments, selection based on matching an individual's (user's) personality profile with a product profile may be based on the definitions of the product profile. The advertising customer or brand customer defines the type of emotion they felt towards each product in each advertisement.
[0114] In one example of generating a product profile, a marketing professional defines product attributes and values in a similar way to how target groups are identified by attributes in a personality profile (e.g., MBTI attributes and percentage values ("scores")). Therefore, a product profile may include attributes and scores defined by percentage values, similar to a personality profile. These product attributes may be divided into different groups. For example, in one embodiment, such three groups (also referred to as "evaluations") are "brand evocation," "product symbolism," and "product use." Each of the three groups may have the same elements (attributes) or different elements, and the product profile is defined by setting percentage values for these attributes.
[0115] In one exemplary embodiment, each of the three groups (evaluations) can be defined by one of many terms (for example, the "25 positive emotions" commonly used in marketing: empathy, kindness, respect, love, admiration, dreaminess, aspiration, desire, adoration, ecstasy, joy, entertainment, hope, expectation, surprise, vitality, courage, pride, confidence, inspiration, attraction, fascination, reassurance, relaxation, satisfaction), and a percentage value can be assigned to each.
[0116] In another embodiment, only attribute terms associated with the product are defined, and the associated values are not defined. By selecting one word in each of the three evaluation groups that make up the product profile, it becomes possible to define the corresponding musical content descriptor (in one embodiment, primarily atmosphere) necessary for fitting the product to the target group. For example, searching for an individual to promote a new Harley-Davidson motorcycle: • Brand evokes respect • Product symbol: Entertainment • Product usage: Satisfied This can be done by defining these three attributes (one for each evaluation group).
[0117] There are two methods for mapping product profiles to individual personality profiles. a) Applying mapping rules from attributes and scores defining the product profile to a personality profile (e.g., MBTI). This allows for the derivation of a target personality profile that is compared to an individual's (pre-calculated) personality profile using similarity search as described above. The mapping rules from the product profile to personality profile elements may be defined manually or learned by a machine learning algorithm similar to the mapping rules from content descriptors to personality profile elements as disclosed above. b) Mapping from a defined product profile to a corresponding music content descriptor by specifying a set of mapping rules. The sentiment profile of a user (individual) or user group is calculated by aggregating music content descriptors, as described above. Based on the sentiment profile (ad-hoc or preliminary calculation) represented by values in the music content descriptor space, a similarity search is performed in the music content descriptor space to find the best-matching user or user group.
[0118] Figure 5 shows one embodiment of method 100 for determining a user or user group that matches a target personality profile. This method begins in step 110, where identification information for a user or user group is obtained for a group of media items, including one or more media items. The identification information for a group of media items may be the user's or user group's playlist or media consumption history. For example, the identification information for one or more media items may include the user's (or user group's) short-term media consumption history, and the personality profile may characterize the user's (or user group's) current mood. In step 120, the user's media items (music consumption history) are analyzed, and a set of media content descriptors is obtained for each of the identified one or more media items. The media content descriptors include characteristics that characterize the acoustic descriptor, semantic descriptor, and / or emotional descriptor of each media item, and may be calculated directly from the media items or read from a database. Details regarding the generation of media content descriptors are as described above.
[0119] The music content descriptors of all media items consumed by the user (within a specified time) are aggregated, and the percentage value becomes a "collected music content descriptor" that contains the same attributes as the aggregated music content descriptor. In this way, step 130 determines a set of collected media content descriptors (i.e., the user's emotional profile) for the entire group of one or more identified media items based on each media content descriptor of the individual media items. For example, if one or more identified media items correspond to a playlist, a set of collected media content descriptors is determined for this playlist. If only one media item is identified, a set of collected media content descriptors may be determined from the segment of the media item.
[0120] In step 140, a set of aggregated media content descriptors is mapped to the personality profile defined according to the personality scheme described above. This mapping is optional and may be based on mapping rules. In step 150, the set of aggregated media content descriptors (i.e., sentiment profiles) generated for the user (or group of users) is provided to the media similarity engine.
[0121] The above process is repeated for multiple users or user groups, generating an emotion profile for each of these users or user groups. In this way, multiple emotion profiles are generated, each associated with a corresponding user or user group, characterizing the user or user group in terms of their emotion context.
[0122] Step 160 defines the product or brand profile. For example, for the group "Brand Awakening," the term "Respect" is selected. Then, mapping rules map the product profile to the target content descriptor. For example, the mapping from the product profile attribute "Respect" to the music content descriptor is defined. The term "Respect" is, 〇 In the "Atmosphere" category, "Gentle" and "Brave" have high percentage values. 〇 In the "Rhythmic Atmosphere" category, low percentage values for "Elegant," low percentage values for "Silent," and low percentage values for "Lyric," It can be associated with this.
[0123] Therefore, a set of target content descriptors is generated from the product or brand profile. Step 160 may be performed independently of steps 110-150, before, simultaneously with, or after these steps, to generate the user's sentiment profile.
[0124] In step 170, the emotional profile of a user or user group is compared with a set of target content descriptors to determine at least one user or user group with the best-matching emotional profile. At least one user or user group with the best-matching emotional profile is selected. Thus, the user that best matches the mapping from the product profile to the aggregated music content descriptors in the consumption history based on the similarity search method is identified and output along with the matching score.
[0125] In step 180, a new media item corresponding to the target profile is selected for presentation to at least one determined user or group of users. For example, an electronic message containing the new media item is automatically generated and sent electronically to at least one determined user or group of users. The electronic message (i.e., the new media item) may contain information about a product or brand associated with the target personality profile.
[0126] If the best-matching users are identified, ads may be pushed first to the users who best match the product profile or the brand's target group. The system may also output a list of identified users to the target audience based on, for example, a user identifier, as well as a matching score value indicating the degree to which the user is a good fit for the brand or product.
[0127] Figure 6 shows the mapping from product profile attributes to atmosphere content descriptors. It illustrates the mapping rules for the product attribute "sympathy" to atmospheres such as "sentimental," "cool," and "friendly." In this example, the attribute "sympathy" requires nearly 50% of the atmospheres "sentimental" and "innocent," while requiring nearly 100% of the atmospheres "cool," "friendly," and "gentle." Similar but different rules exist for other evaluations (groups of product attributes). Similarly, it shows the relationship between the product attributes "kindness" and "respect" and atmosphere content descriptors.
[0128] Therefore, only the emotional profiles of users (or user groups) that strictly meet the above atmosphere criteria are considered candidates for matching. Depending on the selected similarity method, the closer the numerical values, the higher the relevance score in the output of the matched user (or user group).
[0129] Similar to the mapping from content descriptors to personality profiles described above, mapping rules define how (collected) media content descriptors can be used to search for matching users. Mapping rules define the product profile attributes that contribute to the media content descriptor, their values, and how they manifest. Again, mapping rules may be learned using machine learning techniques.
[0130] In one embodiment, the media similarity engine is configured to select individuals in real time by matching their current mood with the product or personality profile of a target group. In this case, the brand defines the target group by its target personality profile. The target group may be defined by a personality profile scheme such as MBTI, OCEAN, Enneagram, or Ego-Equilibrium. The system finds individuals who have an instantaneous short-term personality profile (commonly known as an "instantaneous user mood profile") that fits a given target profile. The brand can then target these individuals with specific advertisements.
[0131] In this embodiment, the system analyzes the user's current media consumption (e.g., the last 10 music tracks) in real time to instantaneously profile the user and assign them to a target group. The definition of the brand's target group or product profile, as well as the mapping between music and individuals (listeners), is carried out in the same manner as described above. Advertisements are pushed to users who match the brand / product's target group. While listening to music, individuals who should receive individually targeted advertisements that best match their current short-term personality profile (commonly known as the "instant user mood profile"), or who should be targeted to specific branding or e-commerce campaigns, are selected.
[0132] Similar to the embodiments described above, the search for users matching a product, brand, or target profile can be performed in the content descriptor domain by comparing a set of aggregated media content descriptors associated with the user against a set of target media content descriptors generated from the product, brand, or target profile by mapping rules. In this scenario, the user's aggregated media content descriptor set (emotional profile) is calculated in "real time," meaning that only the most recent tracks (e.g., 10) that the user has watched are used. Using the personality profiling engine described above, the system calculates the user's emotional profile for the most recent short timeframe. By performing this calculation for all users, the system periodically stores all "real-time" user emotional profiles in the database (e.g., every 10 tracks watched). Once the system has grasped the product or brand profile (as described above), it can find groups of users whose emotional profiles align with the product / brand profile, as described above. Based on the emotional alignment between the brand or product and the users, the system outputs a list of users that the brand should target at this time.
[0133] The characteristics of the above-described apparatus (devices, systems) correspond to the characteristics of each method, but it should be noted that for the sake of simplification, they may not be explicitly described. It is understood that the disclosures in this specification also extend to such characteristics of methods. In particular, it should be understood that this disclosure relates to methods for operating the above-described devices and / or to the provision and / or arrangement of the individual elements of these devices.
[0134] Furthermore, it should be noted that the exemplary embodiments of the disclosure can be realized in many ways by using hardware and / or software configurations. For example, the embodiments of the disclosure can be realized by using hardware capable of running dedicated hardware and / or software. The components and / or elements in the figures are illustrative only and do not limit the scope of use or functionality of any hardware, software combined with hardware, firmware, embedded logic components, or combinations of two or more such components that realize a particular embodiment of the disclosure.
[0135] Furthermore, it should be noted that this specification and drawings merely illustrate the principles of the disclosure. Those skilled in the art will be able to realize various configurations that embody the principles of the disclosure and that fall within its conception and scope, although these are not explicitly described or illustrated herein. Moreover, all examples and embodiments outlined in this disclosure are explicitly intended solely for illustrative purposes to assist the reader in understanding the principles of the proposed method. Furthermore, all descriptions in this specification that provide the principles, aspects, and embodiments of the disclosure, as well as their specific examples, are intended to encompass their respective equivalents. Glossary
[0136] The following technical terms are used throughout this specification.
[0137] media Media includes all kinds of media items that can be presented to the user, such as audio (especially music) and video (including embedded audio tracks). Furthermore, photographs, series of photographs, slides, and graphic representations are examples of media items.
[0138] Media content descriptor Media content descriptors (commonly known as "characteristics") are calculated by analyzing the content of a media item. Music content descriptors (commonly known as "music characteristics") are calculated by analyzing digital audio (segments (excerpts) or the entire song). These consist of a set of music content descriptors that include atmosphere, genre, context, acoustic attributes (key, tempo, energy, etc.), vocal attributes (presence of vocals, vocal family, vocal gender (low or high)), etc. Each of these includes the range of a descriptor or characteristic. Characteristics are defined by a name and a floating-point or percentage value (e.g., bpm: 128.0, energy: 100%).
[0139] Song A musical piece is an example of a media item that consists of single lines (melody) or double lines (harmony) and represents audio data containing timbres or sounds produced by one or more voices or instruments, or both. A media content descriptor for a musical piece is also called a musical content descriptor or musical profile.
[0140] Emotional Profile An emotional profile can be determined for many media items, and it includes one or more sets of media or musical content descriptors associated with a mood or emotion. In this case, it is a collection of content descriptors for individual media items. These are typically derived by aggregating media / musical content descriptors from a set of media items associated with a person or individual (for example, consumed by that person or individual). These contain the same elements as the media / musical content descriptors, and their values are determined by aggregating the individual content descriptors (depending on the aggregation method used).
[0141] Persons (users, individuals) and personality profiles A person (also referred to as a user or individual) is characterized by an emotional profile or personality profile. An emotional profile is characterized by elements of a media content descriptor (see above). A personality profile, on the other hand, includes many different elements with percentage values. The elements of a personality profile are weighted elements within a personality profile scheme (defined by a name or attribute and a percentage value, e.g., MBTI: "EI: 51%"). Personality profiles are defined by personality profile schemes such as MBTI, OCEAN, and Enneagram. - User mood (instantaneous, short-term) (i.e., a personality profile interpreted as the user's short-term emotional status (also called the user mood profile)), or - User personality type (long term) (i.e., a personality profile derived from long-term observation of the user's media consumption behavior), This may be related.
[0142] Target group The target group represents a group of individuals. This is identified as one or a combination of "personality profiles." Optionally, it may be enriched with individual-related parameters (e.g., age, gender, etc.).
[0143] product A product profile includes product attributes expressed in a psychological, emotional, or marketing manner. These attributes may be associated with a percentage value of importance.
[0144] brand Product profiles may be related to a brand. Brand profiles include brand attributes expressed in psychological, emotional, or marketing terms. Attributes may be associated with percentage values of importance.
[0145] mapping Mapping is an algorithmically implemented set of rules that translate profiles from one entity (e.g., media item, song) to another entity (e.g., person, product, or brand) (or vice versa). For example, mapping can be applied between a set of content descriptors (emotional profiles) and personality profiles from a personality profiling scheme.
[0146] Similarity search Similarity search is an algorithmic procedure that calculates the similarity, proximity, or distance between two or more "profiles" of any kind (e.g., emotion profiles, personality profiles, product profiles). The output is a ranked list of profile items with matching scores (values indicating the degree of matching between profiles).
Claims
1. A method for determining the best matching media content descriptor set, - Obtaining a target profile with multiple profile scores, - Mapping the target profile to a set of target content descriptors having multiple characteristics by applying at least one mapping rule that defines how to calculate the characteristics of a set of target content descriptors from the profile score, - Obtaining a set of media content descriptors, each associated with a media item or group of media items, and having the characteristic of including a semantic descriptor for each of the media items or group of media items, wherein the semantic descriptor includes at least one emotion descriptor for the media item or group of media items, - With respect to the target content descriptor set, search for at least one media content descriptor set that has the best matching media content descriptor set, Methods that include...
2. The method according to claim 1, wherein the media item includes a musical portion, and is preferably a musical piece.
3. The method according to claim 1 or 2, wherein the characteristics of a media content descriptor set for a media item include one or more acoustic descriptors of the media item determined based on an acoustic analysis of the media item.
4. The method according to any one of claims 1 to 3, wherein the characteristics of a set of media content descriptors for a media item are determined based on an artificial intelligence model that determines a semantic descriptor for the media item.
5. The method according to claim 4, wherein the semantic descriptor includes one of genre, presence of voice, gender of voice, mood of music, and rhythmic mood.
6. The method according to claim 4 or 5, wherein the emotion descriptor is determined by the artificial intelligence model.
7. The method according to any one of claims 1 to 6, wherein obtaining the plurality of media content descriptor sets includes reading media content descriptor sets for media items or groups of media items from a database.
8. The method according to any one of claims 1 to 7, wherein the mapping rules are learned by machine learning techniques.
9. The method according to any one of claims 1 to 8, wherein searching for the best-matching content descriptor set is based on matching the target content descriptor set with a media content descriptor set having the same or similar characteristics as the target content descriptor set.
10. The method according to any one of claims 1 to 8, wherein the search for the best matching content descriptor set is based on a similarity search in which the corresponding characteristics of the content descriptor sets are compared and a matching score indicating the similarity of each pair of content descriptor sets is calculated.
11. The method according to claim 9 or 10, further comprising ranking the media content descriptor sets according to their matching scores.
12. The method according to any one of claims 1 to 11, wherein the target profile is based on a personality scheme that defines many profile scores for target profile elements representing personality traits, corresponding to individual users or user groups.
13. The method according to claim 12, wherein the target profile is determined based on the short-term media consumption history of the user or the user group and characterizes the current mood of the user or the user group.
14. The method according to claim 12 or 13, wherein the target profile is determined based on the media consumption history or playlists of the user or the user group.
15. The method according to any one of claims 12 to 14, wherein a media item corresponding to the best-matching media content descriptor set is selected for playback or recommendation to the user or the user group.
16. The method according to any one of claims 12 to 15, wherein information associated with media items corresponding to the best-matching media content descriptor set is provided to the user or a user device associated with the user.
17. The method according to any one of claims 12 to 16, wherein searching the content descriptor set depends on the context or environment of the user or user group corresponding to the target profile.
18. The aforementioned target profile corresponds to a product or brand, The method according to any one of claims 1 to 11, further comprising determining a user or user group associated with the best-matching media content descriptor set.
19. The method according to claim 18, wherein each media content descriptor set is associated with a user or a group of users.
20. The method according to claim 18 or 19, wherein the media content descriptor set includes a collection characteristic that characterizes a group of media items presented to the user or user group, in particular a collection characteristic that characterizes media items identified in a playlist associated with the user or user group.
21. The method according to any one of claims 18 to 20, wherein a user personality profile is provided to the user or user group associated with the media content descriptor set, and the personality profile includes a plurality of personality scores for elements of the profile that represent the personality traits of the user or user group based on a personality scheme.
22. The method according to claim 21, wherein the user personality profile is generated by mapping the characteristics of the associated media content descriptor set to personality scores, and characterizes the user's personality or the user's demeanor.
23. The method according to claim 22, wherein mapping the characteristics of the media content descriptor set to a personality score is based on a mapping rule that defines a method for calculating a personality score from the characteristics of the media content descriptor set associated with the user.
24. The method according to claim 23, wherein the mapping rule is learned by machine learning technology.
25. The method according to any one of claims 18 to 24, wherein the search for the best-matching media content descriptor set depends on the context or environment of the user or user group associated with the media content descriptor set.
26. The method according to any one of claims 18 to 25, wherein a media item corresponding to the target profile is selected for presentation to the user or user group.
27. The method according to any one of claims 18 to 26, wherein an electronic message containing information about the product or brand is automatically generated for the user or user group, and the generated message is electronically transmitted to the user or user group.
28. A computer device having a memory and a processor configured to perform the method according to any one of claims 1 to 27.