Method and apparatus for identifying patterns in interactions with medicines

By analyzing media items from human authors using classifiers and clustering analysis, this method identifies patterns in medicine interactions without relying on clinical data, offering an efficient and effective solution for assessing medicine safety, tolerability, and usage trends.

WO2025093878A1PCT designated stage expired Publication Date: 2025-05-08TALKING MEDICINES LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2024/052769
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-10-31
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Current methods for identifying patterns in interactions between people and medicines rely on clinical research data and confidential patient records, which are not always available or feasible to use.

Method used

A method that analyzes media items created by human authors, such as social media posts, using classifiers to identify author types, medicine references, sentiments, and opinions, and performs clustering analysis to determine patterns in author involvement with medicines.

Benefits of technology

This approach allows for the efficient analysis of large volumes of data without relying on clinical research or patient records, enabling the identification of patterns and trends in medicine interactions and improving clustering and pattern determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2024052769_08052025_PF_FP_ABST
    Figure GB2024052769_08052025_PF_FP_ABST
Patent Text Reader

Abstract

An aspect of the invention provides a method of identifying patterns in author involvement with a medicine, comprising: (a) accessing a database of media items created by human authors; (b) processing the media items with an author type classifier to identify at least one media item created by an author type; (c) processing said at least one media item with a named entity recognition classifier to identify references in said at least one media item to the medicine; (d) processing said at least one media item referring to the medicine with a sentiment classifier to identify at least one of: (i) an author sentiment regarding said medicine; and (ii) an author opinion regarding the medicine; (e) processing the media items to identify further media items by the author; (f) processing the further media items to identify personal characteristics of the author; (g) repeating steps (b)-(f) for a plurality of authors to create aggregate author data; (h) processing the aggregate author data to perform a clustering analysis in dependence on said author sentiments and / or said author opinions, and said personal characteristics; and (i) analysing the result of the clustering analysis to determine patterns in author involvement with the medicine.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method and apparatus for identifying patterns in interactions with medicines

[0002] Field of the invention

[0003] The invention relates to the field of identifying patterns in interactions between people and medicines. to the invention

[0004] It is useful to determine interactions between people and medicines, to assess factors such as the safety and tolerability of medicines, compliance with treatment regimens and medical treatment practices, and to determine trends in the use and prescription of medicines.

[0005] Such interactions are typically identified using traditional medical research programs which output clinical research data, and / or by analysis of confidential patient record data. It would be advantageous to assess such factors and trends, and to make predictions, without clinical research data or confidential patient record data.

[0006] It is in this context that the present invention has been devised. of the invention

[0007] According to an aspect of the invention, there is provided a method of identifying patterns in author involvement with a medicine, comprising: (a) accessing a database of media items created by human authors. It may be that the method comprises (b) processing the media items with an author type classifier to identify at least one media item created by an author type. It may be that the method comprises (c) processing said at least one media item with a named entity recognition classifier to identify references in said at least one media item to the medicine. It may be that the method comprises (d) processing said at least one media item referring to the medicine with a sentiment classifier to identify at least one of: (i) an author sentiment regarding said medicine; and (ii) an author opinion regarding the medicine. It may be that the method comprises (e) processing the media items to identify further media items by the author. It may be that the method comprises (f) processing the further media items to identify personal characteristics of the author. It may be that the method comprises (g) repeating steps (b)-(f) for a plurality of authors to create aggregate author data. It may be that the method comprises (h) processing the aggregate author data to perform a clustering analysis in dependence on said author sentiments and / or said author opinions, and said personal characteristics. It may be that the method comprises (i) analysing the result of the clustering analysis to determine patterns in author involvement with the medicine.

[0008] Thus, it is possible to efficiently analyse large volumes of media items as only those by selected human authors are analysed. It is possible to isolate the commentary of specific author types. It is possible to determine patterns in author involvement with medicines using additional data from media items which do not directly reference the medicine. Furthermore, we have found that using author sentiments and / or opinions significantly improve clustering and pattern determination.

[0009] The media items are typically discrete digital media comprising one or more of text, audio and video, created by an author. The media items created by human authors may, for example, comprise, or consist of, social media posts. The media items created by human authors may be sourced from open (e.g. online forums or boards) or closed communities (e.g. an advisory board or market research group). The media items created by human authors may be sourced from presentation slides or notes. The media items created by human authors may be sourced from market research (e.g. surveys). The media items created by human authors may be survey responses. The media items created by human authors may be interview responses. The media items created by human authors may be transcripts of audio data. The media items created by human authors may be transcripts of speeches (e.g. from conferences, for example health care professional conferences). The media items created by human authors may be transcripts of market research or focus groups (e.g. market and commercial reports). The media items created by human authors may be transcripts of interviews. The media items may comprise thick data which provides qualitative context regarding a medicine. Thick data is understood within the art to mean information which provides qualitative data which provides context surrounding a target (in this case a medicine of interest) and indicates feelings, interactions, thoughts and behaviours associated with the target. The media items may be collected via software applications (e.g. apps running on mobile electronic devices such as mobile telephones).

[0010] It will be appreciated that this list is not exhaustive and other types of media items which are human authored and refer to a target medicine, for which is it possible to apply a sentiment classifier to identify at least one of an author sentiment and an author opinion regarding the medicine, will be envisaged.

[0011] Typically, the media items relate to non-clinical data. Typically, the media items relate to non-biometric data.

[0012] The media items created by human authors may comprise, or consist of, text media (e.g. posts, articles). The media items created by human authors may comprise audio or video media. Audio and / or video media (e.g. social media posts) may be processed by speech recognition to generate text data corresponding to what is said in the media. The invention seeks to analyse media items created by human authors although it may be that in practice some media items created by computers (e.g. small or large language models) is included due to the practical difficulties in excluding such media items.

[0013] Typically, the media items comprise media items by different author types and processing the media items with an author type classifier to identify at least one media item created by an author type selects a subset of the media items, typically fewer than 90%, fewer than 75% or fewer than 50% of the media items. Identifying at least one media item created by an author type and identifying references in said at least one media item to the medicine may take place in either order. The term medicine may refer to a drug that has been released to market or a drug that is in development (e.g. in the test phase).

[0014] Author involvement with a medicine may be author interaction with a medicine. Author involvement with a medicine may be patient interaction with a medicine. The medicine of interest may be referred to as a “target” of the method.

[0015] The author type may be patients. Thus, the invention is useful to analyse patients and their involvement (e.g. interactions) with medicines. It may be that the identity of the majority or all of the authors is not known. For example, the authors may be identified by pseudonyms (as is normal on social media sites).

[0016] The author type may be health-care professionals. Thus, the invention may be useful to analyse health-care professionals and what they are saying about medicines. The author type may be carers for patients.

[0017] The author type may be digital opinion leaders (DOL). By digital opinion leaders we refer to individuals who are deemed to be influential to others. That is, the digital opinion leaders are individuals whose opinion leads others due to their influence. It may be that digital opinion leaders are sponsored, financially or in kind, for promoting specific products. However, this need not be the case. An example of a digital opinion leader may be a social media influencer.

[0018] The author type may be key opinion leaders (KOL). By key opinion leaders we refer to individuals who are deemed to be experts in a relevant field. That is, the key opinion leaders are individuals whose opinion leads others due to their expertise on a particular subject (e.g. medicine or disease). Typically, key opinion leaders are not sponsored, financially or in kind, for promoting specific products. An example of a key opinion leader may be an academic or researcher.

[0019] This author type may be a policy author (i.e. a person who writes policies for treatment of illnesses or medicines). The author type may be an investor (e.g. in a medical institution or medical company). The invention is especially useful for the analysis of social media posts, because there are large volumes of social media posts available, but in this case there are significant challenges presented by the fact that social media posts are made by different author types. Accordingly, in some embodiments, the media items are social media posts, the authors are patients, the author sentiments and / or opinions are patient sentiments and / or opinions, the personal characteristics are personal characteristics of the patient, the aggregate author data is aggregate patient data and the author involvement is patient interaction.

[0020] Typically, the author sentiment refers to a parameter defining an author’s overall general feeling towards a particular medicine (e.g. positive or negative). Typically, the author opinion refers to a parameter defining the more specific feeling towards a particular medicine. Typically, the author opinion may be an adjective.

[0021] It may be that author sentiment refers to a categorisation of a media item at a post or sentence level. It may be that author opinion refers to a labelling of a media item by the author at a sentence or word level.

[0022] Typically, author sentiment comprises a single value that encapsulates an author’s overall feeling towards a medicine. As an example, the author sentiment may be one of: positive, neutral or negative. It may be that the method comprises deriving the author sentiment from the at least one media item by the author. It may be that the method comprises processing the at least one media item with the sentiment classifier to identify a proportion of the at least one media item having one or more author sentiments. As an example, the method may comprise processing the at least one media item with the sentiment classifier to identify one or more of: a percentage of the at least one media item having a “positive” author sentiment; a percentage of the at least one media item having a “negative” author sentiment; a percentage of the at least one media item having a “neutral” author sentiment. The method may comprise categorising the overall author sentiment towards a medicine based on the percentage(s) identified by the author sentiment. For example, the overall author sentiment may be identified as the author sentiment having the largest percentage. It may be that parts of the at least one media item with “neutral” author sentiment are disregarded and only the relative proportion of “positive” author sentiment and “neutral” author sentiment are used to determine the author sentiment. The author sentiment may for example be a number, label, descriptor (e.g. word), reference to a selected one of a group of possible author sentiments etc. Typically, author opinion comprises a specific word or category of words that the author has used in the at least one media item in conjunction with the medicine. As an example, the author may use the descriptive terms “effective”, “tiring”, “awful” in conjunction with the medicine. Using this example, the author opinion may be determined as the word directly e.g. “effective”, “tiring”, “awful” or the author opinion may be determined as being a category (which could be numbered) in which the descriptive terms are categorised. It may be that the method comprises deriving the author opinion from the at least one media item by the author. It may be that the method comprises processing the at least one media item with the sentiment classifier to identify which descriptive words have been used in conjunction with the medicine.

[0023] Typically, the author sentiment may be derived from the author opinion. That is, the author opinion may be indicative of the author sentiment. As an example, a media item may be as follows: “My experience with DrugA is that it was very ineffective.” In this example, the target is the medicine called “DrugA”, the author opinion is “ineffective” and the author sentiment is “negative”.

[0024] Typically, the author sentiment refers to a category from a list of categories having less than 5 items in the list, less than 10 items in the list or less than 15 items in the list. Typically, the author sentiment is used for coarse grouping of authors.

[0025] Typically, the author opinion refers to a category from a list of categories having more than 25 items in the list, more than 50 items in the list or more than 70 items in the list. Typically, the author opinion is used for fine grouping of authors.

[0026] Typically the author sentiment refers to a category from a first list of categories and the author opinion refers to a category from a second list of categories, wherein the first list has fewer items, for example half or fewer items, than the second list of categories.

[0027] Typically, since the author sentiment has less categories than the author opinion, authors may be categorised as having the same author sentiment but a different author opinion.

[0028] It may be that the media items are social media posts. It may be that the processing of said at least one social media post with a sentiment classifier identifies both (i) an author sentiment regarding said medicine; and also (ii) an author opinion regarding the medicine. It may be that the processing the aggregate author data to perform a clustering analysis is carried out in dependence on both said author sentiments and said author opinions, as independent variables.

[0029] It may be that the authors are patients. It may be that the social media items are social media posts by patients. It may be that the processing with a sentiment classifier identifies both (i) a patient sentiment regarding said medicine; and also (ii) a patient opinion regarding the medicine. It may be that the processing the aggregate patient data to perform a clustering analysis is carried out in dependence on both said patient sentiments and said patient opinions, as independent variables. It may be that analysing the result of the clustering analysis determines patterns in patient involvement with the medicine. Thus, it may be that there is provided a method of identifying patterns in patient involvement with a medicine, comprising: (a) accessing a database of media items created by patients; (b) processing the media items with an author type classifier to identify at least one media item created by a patient type; (c) processing said at least one media item with a named entity recognition classifier to identify references in said at least one media item to the medicine; (d) processing said at least one media item referring to the medicine with a sentiment classifier to identify at least one of: (i) a patient sentiment regarding said medicine; and (ii) a patient opinion regarding the medicine; I processing the media items to identify further media items by the patient; (f) processing the further media items to identify personal characteristics of the patient; (g) repeating steps (b)-(f) for a plurality of patients to create aggregate patient data; (h) processing the aggregate patient data to perform a clustering analysis in dependence on said patient sentiments and / or said patient opinions, and said personal characteristics; and (i) analysing the result of the clustering analysis to determine patterns in patient involvement with the medicine.

[0030] It may be that the further media items by the author comprise or consist of media items which do not mention or reference the medicine. It is advantageous to use such media items as they enable personal characteristics of the author to be determined, improving clustering and pattern determination, but because it is known that there is at least one media item by the author which mentions or references the medicine, there is no need to analyse personal characteristics of authors who are not relevant to the clustering and pattern determination.

[0031] It may be that the identified personal characteristics include the age and / or gender of the author. It may be the personal characteristics include one or more of, at least five or more of, for example at least ten or more of, typically at least fifteen or more of, for example at least twenty or more of: author type, age, gender, medicine, opinion, side effects, comorbidities, behaviour, physical conditions, type of disease / illness experienced by author, education, migration, ethnicity, religious affiliation, marital status, household size, employment, income, location, generation, language, field of study, school, industry, home ownership and type, interests, early / late adopters of technology, alcohol use, drug use, caffeine consumption, sugar consumption, dietary preferences, dietary restrictions, beauty product consumption, pain relief products, over-the-counter medications, pets, commute to work, activity level, diet history, diet, weight trends, height trends, interaction with healthcare providers, interaction with caregivers, influencer behaviour, pharmaceutical representative interaction, interaction with pharmaceutical manufacturers, interaction with insurance carriers, interaction with national health services, interaction with advertising, emotional stress, stress, sleeping patterns, sleeping times, sleeping levels, exercise levels, exercise times, eating patterns, TV patterns, health status data (e.g. “deceased” or “alive”), participation in clinical trials, interaction with advocacy groups, health literacy, disease severity, medicine dosage, medicine tolerability, quality of life, feelings, smoking status (e.g. “non-smoker”, “previous smoker”, “current smoker”), symptoms, non-clinical genetic references, duration of illness and (e.g. target) medicine-related personal characteristics (e.g. medicine mode of action, medicine class in relation to mechanism of actions, medicine target, medicine safety, medicine efficacy and medicine accessibility).

[0032] It may be that the personal characteristics comprise one or more classifications selected from a plurality of classifications. The plurality of classifications may comprise structured medical data. The plurality of classifications may comprise official classifications for medical and health parameters, for example disease state, disease progression, medical procedures, medicines or drugs, and medical terminology. The plurality of classifications may be sourced from trusted external sources (e.g. registries (e.g. medical registries) and / or ontologies). The plurality of classifications may be selected from a structured group of classifications. The structured group of classifications may be classifications from a published standard.

[0033] It may be that the disease and / or treatment progression is classified informally, by unstructured data, using content from one or more media items. It may be that the disease and / or treatment progression is classified formally, with reference to structured data of the one or more classifications selected from the plurality of classifications.

[0034] Advantageously, analysing the further media items allows determination of classifications which could not be determined from an individual media item which mentions a medicine (e.g. which are not mentioned at all in an individual media item).

[0035] Advantageously, the personal characteristic comprising one or more classifications selected from a plurality of classifications provides a mechanism for linking the content of media items by an author (e.g. mention of disease state, disease progression, medical procedures, medicines or drugs) to widely recognised official terminology. This means that colloquial language or informal terminology used by authors in media items can be associated with medically correct language and terminology. Another advantage is that more reliable connections, inferences and relationships to other areas of interest can be obtained from the data.

[0036] It may be that the further media items may each have one or more dates associated with them. It may be that the one or more dates indicate key dates in an author’s health journey. The one or more dates may indicate at least one of the following key dates: when the author began experiencing symptoms, when the author had a particular procedure, when the author began taking a particular medicine. As part of processing the further media items to identify personal characteristics, the disease state or disease stage of progression of the author can be identified in dependence on the key dates. Thus, the disease state or disease stage of progression can be linked to one or more classifications which formally define the disease state or disease stage of progression.

[0037] Advantageously, by processing multiple posts to identify key dates, information such as disease state or disease stage of progression can be identified, which may not have been possible from processing a single media item in isolation.

[0038] The trusted external sources may comprise data sources from official health or drug authorisation bodies such as the US Food and Drug Administration (FDA), National Health Service (NHS) in the UK or the World Health Organisation (WHO). Other such trusted external sources will be envisaged.

[0039] The datasets defined by external sources may comprise a US Food and Drug Administration (FDA) dataset, Anatomical Therapeutic Chemical (ATC) Classification codes, Systematized Nomenclature of Medicine Clinical Terms (SNOMED CT), International Classification of Diseases (ICD), Human Phenotype Ontology (HPO). Other relevant datasets may also be envisaged.

[0040] The plurality of classifications may be stored on a computer system used to perform the method. The plurality of classifications may be stored external to the computer system used to perform the method. The method may comprise obtaining the the plurality of classifications from an external source and / or a storage device of the computer system.

[0041] The method may comprise linking content of the further media items to the one or more classifications selected from a plurality of classifications (for example as part of processing the media items to identify personal characteristics of the author). The method may comprise identifying content of the further media items which is indicative of a type of classification. The method may comprise selecting the one or more classifications from the plurality of classifications corresponding to the content of the one or more media items indicative of a type of classification. The method may be performed using a machine learning model (e.g. a neural network language model such as a Convolutional Neural Network (CNN) model, or a Transformer Language model, or a Large Language Model (LLM)).

[0042] The method may comprise identifying and verifying which medicine that the at least one media item refers to using the one or more classifications selected from the plurality of classifications. For example, a patient my refer to a brand name of medicine in their media item(s), and the one or more classifications selected from the plurality of classifications may identify the medicine by its drug class or chemical name.

[0043] It may be that the creating the aggregate author data comprises processing the one or more classifications selected from the plurality of classifications (e.g. as part of processing the personal characteristics of aggregate author data). It may be that method step (h) comprises processing the aggregate author data to perform a clustering analysis in dependence on said author sentiments and / or said author opinions, and one or more classifications selected from the plurality of classifications.

[0044] Advantageously, by using the one or more classifications selected from the plurality of classifications (e.g. where the one or more classifications selected from the plurality of classifications is a personal characteristic) in the clustering analysis, the medical accuracy of the data used in the clustering analysis is improved because the one or more classifications selected from the plurality of classifications is part of established medical ontologies and frameworks.

[0045] Performing the clustering analysis without the one or more classifications selected from the plurality of classifications can be advantageous in that organic emergence of patterns and trends from the data in the media items can be observed without restricting the data to clearly defined ontologies and frameworks of the plurality of classifications. This provides the potential to uncover true patterns and trends from the data in the media items that have not been previously observed by other studies.

[0046] For some author types, for example when the author type is a KOL, HCP or DOL, a personal characteristic may include the speciality or area of expertise of the KOL, HCP or DOL. For example, the personal characteristic may include the relevant field in which a health care provider, expert, academic, researcher or influential person (i.e. KOL or DOL) specialises in. The relevant field may be a specific type of disease or group of diseases or specific medicine or group of medicine.

[0047] It may be that the method comprises processing the at least one media item to identify linked data to the at least one media item (and / or the further media items). It may be that the method comprises processing the linked data to identify the personal characteristics of the author. It may be that the method comprises repeating steps of i) processing the at least one media item (and / or the further media items) to identify linked data to the at least one media item and ii) processing the linked data to identify the personal characteristics of the author for a plurality of authors to create the aggregate author data. It may be that the linked data comprises data relating to the author from a source other than a media item (e.g. registries (e.g. medical registries) and / or ontologies).

[0048] It may be that the method comprises identifying media item data of at least one media item. It may be that the method comprises identifying media item data of at least one media item for a plurality of authors to create the aggregate author data. It may be that the method comprises processing the aggregate author data to perform a clustering analysis in dependence on said author sentiments and / or said author opinions, said personal characteristics, and said media item data. For example, the media item data may comprise one or more of: a source of media item (e.g. Reddit (RTM) or SocialGist Boards (RTM)), site of source of the media item (e.g. specific subreddit or specific board), data type with subset, message, sender, receiver, and channel. The media item data may be media item type data.

[0049] It may be that the analysing of the result of the clustering analysis to determine patterns in author involvement with the medicine comprises identifying changes in the clustering of the authors with time, age and / or disease or treatment progression.

[0050] It may be that analysing the result of the clustering analysis to determine patterns in patient involvement with the medicine comprises identifying changes in time in the clustering of the authors.

[0051] It may be that the method comprises predicting one or more effects of a medicine on an individual patient based on the clustering analysis.

[0052] It may be that the method comprises the step of determining one or more properties of a representative author within a cluster identified by the clustering analysis. It may be that the method comprises the step of providing the determined one or more properties of a representative author to one or more processors executing a neural network language model, the one or more processors responding to user queries taking into account the determined one or more properties of the representative author.

[0053] The one or more properties of a representative author might comprise the personal characteristics of a representative author. The one or more properties of a representative author might comprise the author sentiment and / or author opinion of the representative author. The representative author may be a selected author from amongst the authors in a cluster. The representative author may have properties determined as typical, average or otherwise representative of authors in a cluster.

[0054] The method may further comprises providing media items by human authors in a cluster to the one or more processors executing a neural network language model, the one or more processors responding to user queries taking into account the provided media items. Providing may comprise providing references to media items enabling analysis or retrieval of media items.

[0055] The neural network language model may be formed using a deep learning model. The neural network language model may be formed using a transformer model. It may be that the method comprises the step of assessing the effectiveness of a communication path for communication of information about a medicine responsive to the clustering analysis. Effectiveness may be determined taking into account when information was communicated about a medicine and / or to whom. It may be that the step of assessing the effectiveness of a communication path for communication of information about a medicine responsive to the clustering analysis uses one or more models that identify for effective forms of messaging depending on a given audience (e.g. the author type). Typically, the one or more models use one or more of the personal characteristics discussed above. For example, it may be that in a patient population, the source of information about the medicine that a patient receives may be inferred from the media items authored by the author.

[0056] It may be that the method comprises the step of using the results of the clustering analysis to define, measure the success of, and / or control, the progress of a clinical trial for a medicine. It may be that the step of using the results of the clustering analysis to define, measure the success of, and / or control, the progress of a clinical trial for a medicine uses one or more models that models that determine the effects of a particular medicine on patient populations (e.g. patients having a particular disease and / or patients of a particular demographic. Typically, the one or more models use one or more of the personal characteristics discussed above. For example, it may be that in a patient population, patients can be identified in various stages of a clinical trial for a medicine. It may be that the success of the clinical trial is inferred from the personal characteristics of the author. Typically, the success of the clinical trial is defined by determining the author sentiment and opinion that an author taking part in the clinical trial has after taking the medicine of the clinical trial.

[0057] It may be that the method comprises the step of using the results of the clustering analysis to assess one or more of the group consisting of: the safety, efficacy, patient tolerance and accessibility of a medicine, and the quality of life of patients taking a medicine. It may be that the step of using the results of the clustering analysis to assess one or more of the group consisting of: the safety, efficacy, patient tolerance and accessibility of a medicine, and the quality of life of patients taking a medicine uses one or more models that determine and output the ability of an intervention or drug, to produce a desired effect (e.g. improvement in health). Typically, the one or more models use one or more of the personal characteristics discussed above. For example, it may be that in a patient population, the safety, efficacy, patient tolerance and accessibility of a medicine, and the quality of life of patients taking a medicine is inferred from the author sentiment and author opinion regarding the medicine.

[0058] It may be that the method comprises the step of using the results of the clustering analysis to predict that an individual patient or a group of patients will start to take, or continue to take, a medicine. It may be that the method comprises the step of using the results of the clustering analysis to identify patterns in or predict whether and / or when an individual patient or group of patients will start to take and / or continue to take a medicine. For example, it may be that in a patient population, within the same therapeutic domain, patients can be identified in various stages of their journey. Typically, each stage requires a different drug. It may be that the stage each patient is at is inferred from their personal characteristics. Typically, the prediction of whether an individual patient or a group of patients will start or continue taking a drug and / or the patterns in when an individual patient or group of patients will start to take and / or continue to take a medicine may be made by identifying what stage the patient is in their patient journey. For some author types, such as patients, those patients further into their disease progression may have more diseases than those patients earlier in their disease progression. Therefore, when a patient is further into their disease progression journey, their disease profile, and the medicines the patient takes for the diseases, is more complex. Therefore, it may be that the method comprises the step of using the results of the clustering analysis to predict that an individual patient or a group of patients will start to take, or continue to take, a medicine by using a control variable of a patient early on in their disease progression. This is because patients early on in their disease progression are less likely to have developed further diseases.

[0059] When the clustering analysis is performed in dependence on the one or more classifications selected from the plurality of classifications (e.g. where the one or more classifications selected from the plurality of classifications is a personal characteristic), the prediction that an individual patient or a group of patients will start to take, or continue to take, a medicine may be more reliable. This is because the prediction uses the one or more classifications selected from the plurality of classifications which can define formal terminology and official classification of medicines and drugs.

[0060] It may be that the method comprises the step of using the results of the clustering analysis to identify patterns in or predict whether and / or when an individual patient will suffer a specific disease, or that a disease which they have will progress to a specific state. For example, it may be that in a patient population, subpopulations with similar characteristics are identified through clustering analysis. If subpopulation A has the same personal characteristics as subpopulation B, but subpopulation B has an additional characteristic being suffering from disease A, the prediction that subpopulation A will progress to disease A can be made.

[0061] When the clustering analysis is performed in dependence on the one or more classifications selected from the plurality of classifications (e.g. where the one or more classifications selected from the plurality of classifications is a personal characteristic), the classification and grouping of authors into subpopulations may be more reliable. This is because the grouping and classification depends on the one or more classifications selected from the plurality of classifications which can define formal terminology and official classification of illnesses experienced by the author and medications taken by the author.

[0062] When the clustering analysis is performed in dependence on the one or more classifications selected from the plurality of classifications (e.g. where the one or more classifications selected from the plurality of classifications is a personal characteristic), the identification of patterns in or prediction whether and / or when an individual patient will suffer a specific disease, or that a disease which they have will progress to a specific state may be more reliable. This is because the prediction uses one or more classifications selected from the plurality of classifications which can define formal terminology and official classification of disease progression.

[0063] It may be that the patterns or predictions relate to two or more different medicines.

[0064] It may be that the method does not include the step of processing electronic healthcare records relating to individual patients.

[0065] Advantageously, the method and system disclosed herein can be used to understand interactions with medicines and analyse or predict disease states and health outcomes without using health records for each patient. This can avoid the legal and operational complexities associated with obtaining access to health records. In addition, since the method uses content generated by the patient directly, self-reported health conditions can also be considered in the analysis.

[0066] It may be that the method comprises the step of generating aggregate author data which is segregated by one or more demographic characteristics of authors. It may be that the method comprises the step of processing the further media items by an author to identify one or more words which are relevant to the further media items by the author and taking into account the identified one or more words when performing the clustering analysis.

[0067] It may be that the method comprises the step of processing the further media items by an author to identify one or more words which are relevant to media items by the authors in a cluster of authors.

[0068] Typically, the method may use term frequency-inverse document frequency (TF-IDF) to determine which words are relevant to media items. That is, TF-IDF may be used to assign a numerical weighting on the importance of a word depending on the frequency of use of the word in a set of media items. TF-IDF may also be used after the clustering analysis is performed. For example, TF-IDF may be used on the results of the clustering analysis to determine which words are to be associated with the representative author.

[0069] It may be that one or more of the words which are identified are words which relate to symptoms. It may be that one or more of the words which are identified are words which relate to medicines. It may be that one or more of the words which are identified are words which relate to health conditions. It may be that one or more of the words which are identified are words which relate to lifestyle factors which affect disease. It may be that one or more of the words which are identified are words which relate to feelings.

[0070] It may be that the method comprises the step of performing a dimensionality reduction procedure on the aggregate author data which is the subject of the clustering analysis and representing the data in a two or three dimensional representation, with authors in different clusters displayed differently.

[0071] It may be that the method comprises the step of receiving a query, the query identifying a cluster output from the clustering analysis, and selecting media items by one or more authors within the identified cluster.

[0072] The selected media items may be output. The selected media items may be provided to a one or more processors implementing a neural network language model. The one or more processors implementing a neural network language model may responding to user queries taking into account the content of the selected media items. For example, the neural network language model may be trained on the selected media items.

[0073] It may be that the number of independent personal characteristics processed by the clustering analysis is in the range from 10 to 1000 inclusive. It may be that the number of independent personal characteristics processed by the clustering analysis is in the range from 10 to 500 inclusive. It may be that the number of independent personal characteristics processed by the clustering analysis is in the range from 10 to 250 inclusive. It may be that the number of independent personal characteristics processed by the clustering analysis is in the range from 10 to 100 inclusive. It may be that the number of independent personal characteristics processed by the clustering analysis is in the range from 10 to 20 inclusive. It may be that the number of independent personal characteristics processed by the clustering analysis is in the range from 10 to 25 inclusive.

[0074] It may be that either the method step of processing the aggregate author data to perform a clustering analysis in dependence on said author sentiments and / or said author opinions, and said personal characteristics and or the method step of analysing the result of the clustering analysis to determine patterns in author involvement with the medicine comprises analysing the further media items by the author to assess alignment between the further media items and a statement of interest.

[0075] Advantageously, assessing alignment between the further media items and a statement of interest allows for author agreement with a statement of interest to be determined. This can be useful information for a variety of purposes, such as identifying a viewpoint of a patient population or sub-population towards a medicine and / or identifying strategies for marketing campaigns.

[0076] It may be that the statement of interest is a target statement which defines a target textual expression related to topic, such as a target medicine, symptom, patient experience, company or brand. The statement of interest may be a predefined textual expression chosen to assess the agreement of an author type with a particular statement.

[0077] It may be the method step of analysing the further media items by the author to assess alignment between the further media items and a statement of interest comprises applying a measurement or scoring model configured to calculate an agreement score between the statement of interest relating to a topic and an expression of the author relating to the same topic. The expression of the author may be determined from the further media items. The agreement score may be a score between a predetermined range, such as 1 to 5. The agreement score may be a continuous or discrete scale.

[0078] Where the method step of processing the aggregate author data to perform a clustering analysis in dependence on said author sentiments and / or said author opinions, and said personal characteristics comprises analysing the further media items by the author to assess alignment between the further media items and a statement of interest, the agreement score may be used as a personal characteristic. That is, the personal characteristics used to perform the clustering analysis may comprise the agreement score. In this way, the clustering analysis takes into account the alignment between the further media items and the statement of interest.

[0079] When the method step of analysing the result of the clustering analysis to determine patterns in author involvement with the medicine comprises analysing the further media items by the author to assess alignment between the further media items and a statement of interest, the agreement score may be used to determine a statistical relevance of the strength and relationship between clusters.

[0080] This in itself is believed to be novel, therefore according to an aspect of the present invention, there is provided a method of identifying patterns in author alignment with a statement of interest. It may be that the method comprises (a) accessing a database of media items created by human authors. It may be that the method comprises (b) processing the media items with an author type classifier to identify at least one media item created by an author type. It may be that the method comprises (c) processing the media items to identify further media items by the author. It may be that the method comprises (d) analysing the further media items by the author to assess alignment between the further media items and the statement of interest. It maybe that the method comprises (e) identifying patterns in the author alignment with the statement of interest.

[0081] In this method, analysing the further media items by the author to assess alignment between the further media items and the statement of interest is performed independently of any clustering analysis. This is advantageous as it means that author agreement with a statement of interest can be determined without also carrying out more intensive analysis on the media items, such as clustering analysis. In addition, in this method, the media items are not restricted to having a reference to a target medicine. That is, this method allows the alignment between the further media items and the statement of interest to be assessed for any media item for any author type.

[0082] It may be that the patterns comprise correlations or trends between two or more variables. It may be that the patterns relate to author type or groups (e.g. clusters) of authors of the author type.

[0083] An example embodiment of the present invention will now be illustrated with reference to the following Figures in which:

[0084] Figure 1 is a flowchart of a method of identifying patterns in author involvement with a medicine;

[0085] Figure 2 is a schematic of a computer system operable to carry out the process shown in Figure 1 ;

[0086] Figure 3 is a more detailed schematic of the computer system of Figure 2;

[0087] Figure 4 is a schematic of the data flows in the computer system of Figure 2;

[0088] Figures 5a and 5b are graphs representing clustering analysis;

[0089] Figure 6 is a flowchart of a method of determining one or more properties of a representative author;

[0090] Figure 7 is an example of an output provided by the system;

[0091] Figure 8 is a schematic of a clustering analysis result analyser; and

[0092] Figure 9 is a schematic of a prediction module.

[0093] Detailed Description of an Example Embodiment An embodiment of a process and system for identifying patterns in author involvement with a medicine will now be described.

[0094] In overview, the process involves analysing media items of an author to identify i) an author sentiment and author opinion towards a particular medicine, and ii) the personal characteristics of that author. The personal characteristics of the authors are identified using other media items from the author, not necessarily relating to the medicine, for example the disease / illness experienced by the author. The process is repeated for a plurality of authors and then a clustering analysis on the data, from all of the authors and using these parameters, is performed to determine patterns in author involvement with a medicine. Advantageously, the present process determines patterns in author involvement with a medicine based on patient data relating to the “whole” patient (i.e. not only the data in which the patient describes medication, but also data providing additional contextual information relating to a patient). The media items are often derived from a range of sources, including posts on websites, social media posts generally, and more structured sources, such as apps for monitoring medicine use. The process will now be described in more detail.

[0095] Figure 1 is a flowchart of a method of identifying patterns in author involvement with a medicine. In this example, the medicine of interest is DrugA (where DrugA is a medicine for the purpose of this description). It will be appreciated that the method of Figure 1 can be performed for any medicine.

[0096] In step S100, a database is accessed. The database stores media items (for example, social media posts such as Twitter(RTM) tweets, Reddit(RTM) forum posts, and so on) which are posted by human authors. The media items are accessed and / or provided in any appropriate form, and may, for example, be converted from the original source into a standardized format, such as XML, CSV, HTML, and so on. The media items are processed and stored in a granular format, or may be treated as a single block of data. Typically the media items are selected from specific forums, or filtered using keywords, some combination of the two, or otherwise pre-filtered to return at least approximately relevant subject-matter in general.

[0097] In step S102, the media items, which will herein be referred to as social media posts, are processed using an author type classifier to identify at least one social media post created by an author type. In other examples, the media items may be a sentence, word, paragraph or image translation. The author type distinguishes between different authors (e.g. patient and health care professional). In this example, the social media posts are processed to identify at least one media item created by a “patient” author type.

[0098] In step S104, the social media posts are processed using a named entity recognition classifier to identify references in the at least one social media post to the medicine. The named entity recognition classifier distinguishes between different types of medicines. In this example, the at least one social media post identified in step S102 as having been created by a patient author type are processed to identify references to DrugA in the media item.

[0099] In step S106, the at least one social media post is processed using a sentiment classifier to identify: author sentiment regarding the medicine and / or author opinion regarding the medicine. Author sentiment refers to the general view that a patient has of a particular medicine. Author opinion refers to what specific adjectives the patient associates with the particular medicine. In this example, the social media posts which are created by a patient author type and include references to DrugA are processed using the sentiment classifier. As an example of a social media post, a patient may be identified as having made the following post “I’m on DrugA and I feel awful today”. The identified patient sentiment for this post is negative and the identified patient opinion for this post is “awful”.

[0100] In step S108, the social media posts are processed to identify further social media posts by the author. In this example, the social media posts identified as being created by the patient author type and including references to DrugA are made specifically by PatientA. The further social media posts include posts which were created (i.e. posted) before the initial social media post identified as having a patient author and relating to the medicine of interest. The further social media posts are social media posts also stored on the database which are also created by PatientA. The further social media posts include posts which don’t mention the DrugA and are not related to medical subject-matter. In some embodiments, this step includes retrieving the further social media posts having the same patient author from a further database.

[0101] In step S110, the further social media posts are processed to identify personal characteristics of PatientA. The personal characteristics of PatientA include the age and gender of the author. The age of the patient is determined by processing social media posts to identify age mentions in all posts created by the patient. The gender of the patient is determined by processing social media posts to identify mentions of gender in all posts created by the patient. The processing includes calculating word similarity to predefined gender terms such as ‘femaleV’maleV’other. The gender with highest similarity sum is assigned to patient. Term frequency-inverse document frequency (TF-IDF) is used to identify important terms used in the social media posts. The TF-IDF analysis allocates a weighting to words which occur most frequently in the social media posts as being particularly relevant to the characteristics of patients taking a particular medicine.

[0102] In step S112, the method steps S102 to S110 are repeated for a plurality of authors to create aggregate author data. As the author type in S102 is a patient author type, the method steps S102 to S110 are repeated for a plurality of patients (e.g. PatientB, PatientC and PatientD). As the medicine of interest for the method shown in Figure 1 is DrugA, each of the patients for which the aggregate author data relates are patients with an involvement with the same medicine, DrugA. The aggregate author data is data relating to all of the plurality of patients having created social media posts referencing DrugA.

[0103] In step S114, the aggregate author data is processed to perform a clustering analysis in dependence on the author sentiments and / or the author opinions, and the personal characteristics. The clustering analysis is a technique of data analysis in which data points are grouped into clusters. In a preferred embodiment, the data points are representative of patient sentiment, patient opinion and personal characteristics of the author. Therefore, the plurality of patients are grouped into different clusters depending on their author sentiment, author opinion and personal characteristics. TF-IDF is used on the results of the clustering analysis to determine which words are to be associated with a representative author of patient subpopulation.

[0104] After the clustering analysis is performed on the aggregate author data, the results of the clustering analysis are processed using dimensionality reduction to enable clustering analysis results to be easily displayed. Typically, dimensionality reduction is employed to reduce the data points to 2 or 3 dimensions to provide a visual representation of clusters (see Figures 5a and 5b).

[0105] In step S116, the clustering analysis result if analysed to determine patterns in author involvement with the medicine. The clustering analysis result is a grouping of patients into clusters, which are represented on a graph, such as that shown in Figure 5a. The patterns in author involvement with the medicine refer to relationships and links between patient intake of a medicine and other parameters, e.g. age, progression of disease, gender and / or intake of another medicine.

[0106] Advantageously, the method can be applied to patients having specific personal characteristics, which can be helpful when assessing the effects of medicine on a particular group of patients. It can also be used in the opposite way to identify which sociodemographic groups of patients taking a medicine experience a particular effect from the medicine, or the opinion of the medicine by patients within that group. The patients can also be grouped in accordance with their diagnosis timeline (e.g. prediagnosed, diagnosed and treatment management).

[0107] Figure 2 is a schematic of a computer system operable to carry out the process shown in Figure 1. In its simplest form, the computer system 200 includes at least one processor 202, local data storage 204 for computer program code, temporary data, and any other volatile or non-volatile storage needs of the or each processor. The system 200 also includes at least one network adaptor 206, and appropriate input / output means 208, such as a display, a keyboard and a pointing device, and the like. The computer system 200 may connected via the Internet 210 or any other appropriate network or local connection to various remote data sources 212 or to local data sources (not shown). The various aspects described above may, as appropriate, be distributed in any appropriate fashion, and need not all be provided in the same location.

[0108] Figure 3 is a more detailed schematic of the computer system of Figure 2. A server 302 in this case essentially performs all of the functions of the computer system 200 shown in Figure 2, but separate servers may be provided in respect of different aspects of the system’s functionality, or any other appropriate architecture may be used. The server 302 is accessed by one or more client devices 304, whether a customer device or otherwise. Like the client device 304, information sources 306a, 306b are accessible via a network 350 such as the Internet. It is possible also to have locally connected information sources 306c, for example providing data from an app associated with the provider of the server 302 and optionally running under the control of, or otherwise with the assistance of the server 302. Classifier trainers 308a, 308b are provided to assist with training the aforementioned classifiers, either remotely (308a) or locally (308b). The server 302 has access to multiple databases, including the raw post store 310 for storing downloaded and / or normalized social media posts and other media items. The server 302 accesses a medicine database 312, which contains a list of brand names and generic names for medicines provided within particular jurisdictions (such as the UK and the US). The medicine database 312 is a medicine knowledge database in that it includes data related to the medicine, such as the name and prevalence of the disease treated by the medicine, mode of action of the medicine and approval status and approval details of the medicine.

[0109] An opinion tagger 313 is also accessible for marking up individual posts as having, for example an “awful”, “tired”, “healthy” opinion towards a medicine. The opinion tagger 313 may alternatively be provided within the system (for example on or via the server 302), and different opinion scoring methods are of course possible. An alternative form of opinion tagger 313 is also possible which marks up individual parts of a post separately.

[0110] A sentiment tagger 314 is also accessible for marking up individual posts as having, for example, a positive, negative or neutral sentiment. The sentiment tagger 314 may alternatively be provided within the system (for example on or via the server 302), and different sentiment scoring methods are of course possible. An alternative form of sentiment tagger 314 is also possible which marks up individual parts of a post separately.

[0111] A characteristic identifier 315 is also accessible for identifying personal characteristic of each patient. The characteristic identifier 315 may alternatively be provided within the system (for example on or via the server 302). An as example, a named entity recognition (NER) model is used to detect mentions of age in patients’ posts. Out of the many ages mentioned, empirically, the age the patients mention the most, is the patient’s own age. Therefore, the most commonly mentioned age is determined to be the patient’s age. A NER model is also used to detect mentions of gender in patients’ posts. Out of the many genders mentioned, empirically, the most frequent and similar genders to a first-person mention of a gender, e.g. ‘female’, is the patient’s own gender. A patient’s gender is calculated by determining the similarity between the author’s mentions of gender in media items authored by the patient and the predetermined gender categories of the NER model. This similarity is calculated using language model embeddings.

[0112] A clustering analysis module 316 is also accessible for performing the clustering analysis, in dependence on the patient sentiments and / or the patient opinions, and the personal characteristics, by processing the aggregate patient data. The clustering analysis module 316 may alternatively be provided within the system (for example on or via the server 302).

[0113] A clustering analysis result analyser 317 is also accessible for analysing the result of the clustering analysis performed by the clustering analysis module 316 to determine patterns in patient involvement with the medicine, e.g. DrugA. As an example, the clustering analysis result analyser 317 may determine that patients with DiseaseA, who respond well to DrugA, will also respond well to DrugB. The clustering analysis result analyser 317 may alternatively be provided within the system (for example on or via the server 302). The clustering analysis is performed on data having 10 to 1000 dimensions which include the personal characteristics of the patient, their sentiment and their opinion.

[0114] A prediction module 318 is also accessible for determining predictions associated with a patient (or a group of patients) and / or a medicine (or medicines) of interest. As an example, the prediction module 318 predicts that patients with DiseaseA will start to take DrugB after 6 years of suffering with DiseaseA when processing of existing social media posts from patients with DiseaseA indicates that have started taking DrugB after 6 years of suffering with DiseaseA. The prediction module 318 may alternatively be provided within the system (for example on or via the server 302).

[0115] An output module 319 is also accessible for providing an output to a user of the system. The output module 319 conveys results of the clustering analysis to the user, for example in a format like the graph shown in Figure 5a below. The output module 319 may alternatively be provided within the system (for example on or via the server 302).

[0116] A dimensionality reduction module 321 is also accessible for reducing the dimensions of the clustering analysis to two or three dimensions. The dimensionality reduction module 321 may alternatively be provided within the system (for example on or via the server 302).

[0117] A processor 323 is also accessible to execute a neural network language model (NNLM). The NNLM enables communication between a user of the system and the system itself. The user inputs queries (via voice or text) and the NNLM determines a suitable response for to the query. The processor 323 may alternatively be provided within the system (for example on or via the server 302). A display screen 325 is also accessible to display the output of the clustering analysis (following dimensionality reduction) to a user.

[0118] The central repository 320 stores other discrete data sets such as a processed post store 322, a medicine filter list 324 with a list of keywords or other search terms for carrying out the initial filter on the downloaded posts (or, rather, before the posts are downloaded), per post annotation data 326 including the outputs generated by the classifiers for each post / item.

[0119] Also provided in database 320 are the aggregate author data store 328, which stores the patient data for the plurality of patients having created social media posts referencing the medicine e.g. DrugA, and the author involvement pattern store 330 which stores the patterns in patient involvement with the medicine, as determined by the clustering analysis result analyser 317. The aggregate author data store 328 and the author involvement pattern store 330 store data produced for each medicine of interest (either as a regular task, or on demand, or both).

[0120] As regards the medicine database and medicine data, there are for example different ways to classify medicines. The World Health Organisation (WHO) Anatomic Therapeutic Chemical (ATC code) classifies medicines by therapeutic indication, whereas the FDA classifies medicines by mode of action. Other systems such as ICD- 9, ICD-10 and SNOMED-CT also exist and can be used as and where appropriate. The choice of classification is preferably dependent on the selected jurisdiction or jurisdictions, but different classifications can be combined as and where appropriate also. The plurality of classifications may be obtained from and / or stored on the medicines database 312. The plurality of classifications may be obtained from and / or stored on the information sources 306a, 306b, 306c.

[0121] The server 302 can access a measurement module 340 for analysing the further media items by the author to assess alignment between the further media items and a statement of interest. The measurement module 340 stores the statement of interest, though this can be stored elsewhere in other embodiments, for example in the central database 320. The measurement module 340 applies a measurement or scoring model configured to calculate an agreement score between the statement of interest relating to a topic and an expression of the author relating to the same topic. The measurement module 340 accesses the processed post store 322 and / or per post annotation data 326 to retrieve the media items from which the expression of the author is determined. In some embodiments, the measurement module 340 accesses the clustering analysis module 316 to retrieve the results of the clustering analysis in order to perform analysis on the further media items by the author to assess alignment between the further media items and a statement of interest.

[0122] Figure 4 is a schematic of the software architecture and data flows in the computer system of Figure 2. Entities in the system are shown with normal numerals, and data flows are indicated with the prefix ‘F’.

[0123] The computer system includes 400 includes data sources 402 which includes major social media websites and apps, such as Reddit(RTM) and social network aggregator SocialGist (RTM) which aggregates posts from social networks including Twitter(RTM), custom apps reporting anonymised medical commentary and / or structured data, and specialist forums frequented (at least) by patients. A raw downloader 404 handles the downloading of posts (FOO) from the data sources 402. The downloader is customised to deal with a particular API and / or post format, and output (F02) the posts in a normalized format, tagged with raw data source specific data. For example, a Twitter(RTM) post downloader may store specific data relating to hashtags used, and numbers of likes and retweets. Downloaded and normalised posts are stored in the raw post and normalised post store 406. Downloaded and normalised posts are stored in the raw post and normalised post store 406, which in turn outputs (F04) the normalised posts to the post processor 408.

[0124] The post processor 408 receives (F06) data encoding the trained author type classifier and the trained named entity recognition (NER) model from a model store 410. The post processor 408 also receives (F10) a medicine list (which may be a plurality of classifications relating to a medicine list from an official health or medical body) from a medicine data store 414. The post processor 408 classifies each post according to an identified author type (patient, Health Care Professional (HCP) or non-relevant using a Large Language Model (LLM) of the trained author type classifier. The post processor 408 identifies references to a medicine of interest, from the medicine list, using the NER in posts which are classified as having an author type of interest, in this case a ‘patient’ author type. In some examples, the model store 410 is an external language processing module. The post processor 408 also receives (F08) potential author sentiment and author opinion tags from the author sentiment and opinion tag store 412. The post processor 408 identifies an author sentiment and author opinion regarding the medicine of interest. Aspect-based sentiment analysis is applied to individual words within the post that are mentioned in relation to an adjective or descriptor and is used to determine author sentiment. Aspect-based sentiment analysis is a technique known to those skilled in the art which allocates a sentiment to individual aspects, e.g. words, within a sentence. In an example, the proportion of aspects within a post which are categorised as having positive, negative or neutral sentiment is determined and the relative amount of each sentiment is used to allocate an overall sentiment.

[0125] The post processor 408 receives (F06) a personal characteristic identifier model from the model store 410. The post processor 408 processes additional posts from the same author. In particular, the post processor 408 processes the additional posts from the same author to identify one or more personal characteristics of the author, such as age and gender. The post processor 408 uses the personal characteristic identifier model to select one or more classifications from the plurality of classifications as a personal characteristic. The post processor receives (F38) the plurality of classifications from an information source 430 and selects one or more classifications from the plurality of classifications which corresponds to content of the one or more media items which is indicative of a type of classification.

[0126] For example, the content of the media items may be processed to identify a personal characteristic of a disease state of the author, e.g. stage 1 B lung cancer. The method may comprise identifying that stage 1 B is indicative of a TNM Classification of Malignant T umours (TNM) stage which is a type of classification. The method may that select that the classification of T2a, NO, M0 corresponds to stage 1B lung cancer. In this way, the personal characteristic of the TNM classification of T2a, NO, M0 as a classification selected from the plurality of classifications is determined.

[0127] The post processor 408 repeats the post processing for a number of authors having the same author type and generates aggregate author data which is output (F12) to a database 416 which stores the aggregate author data. The aggregate author data includes demographic characteristics (e.g. age, gender).

[0128] The aggregate author data is transmitted (F34) to a dimensionality reduction module 428. The dimensionality reduction module 428 performs a dimensionality reduction procedure on the aggregate author data. The output of the dimensionality reduction procedure is a representation of the data in a two or three dimensional representation. The output of the dimensionality reduction is provided (F36) to the database 416.

[0129] The aggregate author data and the author sentiments, author opinions, and personal characteristics (including the one or more classifications selected from the plurality of classifications) are sent (F14) to a clustering analysis module 418, which performs clustering analysis on the aggregate author data. The aggregate author data which is the subject of the clustering analysis is the dimensionally reduced aggregate author data. In this example, the clustering analysis module 418 performs a multi-stage clustering analysis. In a first stage, the clustering analysis module 418 performs K- modes clustering and uses a Kneedie algorithm to find the optimal number of clusters. K-modes clustering is performed with this optimal number of clusters using all the personal characteristics, which is high-dimensionality data clustering. This outputs the optimal number of clusters, each of which is associated with personal characteristics. In order for the second stage of clustering analysis to be performed, multiple correspondence analysis is used for dimensionality reduction to reduce the number of personal characteristics to two dimensions which allows K-means clustering to be utilised. In the second stage of clustering analysis, K-means clustering is performed with the optimal number of clusters and the reduced dimensions. The resultant output is the optimal number of clusters identified in the first stage of clustering analysis which have a matching X and Y coordinate along the reduced dimensions corresponding to the centre point of each cluster (i.e. a “persona” as discussed below). K-means clustering analysis provides the ability to plot the patients in a two-dimensional graph and assign them their closest cluster while K-modes clustering analysis provides the personal characteristics representative of those clusters.

[0130] The results of the clustering analysis are returned (F16) to the database 416. The clustering analysis results are sent (F18) to a clustering analysis result analyser 420. The clustering analysis result analyser 420 determines patterns in patient involvement with the medicine and the determined patterns for patient involvement with specific medicines are returned (F20) to the database 416.

[0131] The clustering analysis result analyser 420 also determines one or more properties of a representative author within the cluster and transmits (F20) the one or more properties to the database 416 with the determined patterns for patient involvement with specific medicines. In some embodiments, this analysis and the data which is transmitted, is filtered to, or broken down by category of author, e.g. the analysis is carried out for, or broken down into separate analysis for, patients with a specific disease.

[0132] The clustering analysis result analyser 420 also assesses the effectiveness of a communication path for communication of information about a medicine responsive to the clustering analysis results. The clustering analysis result analyser 420 also defines a measure the success of the progress of a clinical trial for the medicine of interest using the results of the clustering analysis. The clustering analysis result analyser 420 also assesses: the safety, efficacy, patient tolerance and accessibility of the medicine of interest, and the quality of life of patients taking the medicine of interest. This information is transmitted (F20) to the database 416. Advantageously, the clustering analysis can be used to identify the quality of life of patients taking a particular medicine.

[0133] The clustering analysis results are transmitted (F24) to a prediction module 424 which predicts one or more effects of a medicine on a particular individual patient and transmits (F26) the predicted effect(s) of the medicine data to the database 416. The prediction module 424 also predicts that an individual patient or a group of patients will start to take, or continue to take, the medicine of interest. The prediction module 424 also predicts whether an individual patient will suffer a specific disease, and if so, when, or that a particular disease from which an individual suffers will progress to a specific state. The prediction module 424 also predicts if, and if so, when an individual patient or group of patients will start to take and / or continue to take a medicine. This information is transmitted (F26) to the database 416.

[0134] The one or more properties of a representative author are transmitted (F28) to a processor 426 on which a neural network language model (NNLM) is executed. The processor 426 uses the one or more properties to respond (F32) to queries submitted (F30) by a system user. The processor 426 also receives (F28) the results of the clustering analysis and a query which identifies a cluster output from the clustering analysis. The processor 426 selects social media posts from the one or more authors within the identified cluster.

[0135] The results of the clustering analysis (dimensionally reduced) and determined patterns in author involvement with the medicine are output (F22) to an output module 422 along with any other appropriate or requested statistics, e.g. the one or more properties of a representative author within the cluster, the effectiveness of a communication path for communication of information about a medicine, the measure the success of the progress of a clinical trial for the medicine of interest. The output module 422 provides an output to the user through e.g. a display screen. The output module 422 shows the cluster results of the clustering analysis in which the data is represented in two or three dimensions with authors in different clusters displayed differently.

[0136] Figures 5a and 5b are graphs representing clustering analysis. Figure 5a shows the clustering when the author sentiment and the author opinion are independent variables in the clustering analysis. Each point on the graph represents a patient, having an associated patient sentiment and patient opinion. Figure 5b shows the clustering when the author sentiment and the author opinion are not considered in the clustering analysis. Both figures display the same patient population, although Figure 5a appears to have more patients than Figure 5b, since there are very few overlapping data points compared to Figure 5b. The variable “sentiment” has a small number of unique values (e.g. “negative”, “indifferent”, “positive”), which contributes to a coarse grouping of patients. A change in a sentiment value is therefore quite significant in a patient’s (i.e. author’s) (e.g. health) journey. In contrast, the variable “opinion” has a lot of unique values (e.g. “awful”, “itchy”, “miserable”, “tired”, “happy”, “healthy”, “refreshed”), which contributes to a finer grouping of patients. A change in “opinion” of patient occurs more frequently than a change in “sentiment” of a patient. A shared opinion between patients is rarer and results in a stronger similarity between patients in a cluster.

[0137] As shown in Figure 5b, the clustering analysis has fewer overlapping points, with a high concentration of patients within clusters and separation between clusters of patients. When there are too few independent variables, the clusters become distorted by extremes, such as many patients concentrated too far away from each other as shown in Figure 5b, and cluster populations being imbalanced, making it challenging to discern between clusters. Without using patient opinion and patient sentiment, the patient clusters are not sufficiently distinct to draw a meaningful conclusion from the clustering analysis. Therefore, it is clear from Figures 5a and 5b that clustering analysis with author sentiment and author opinion as independent variables produces more defined clusters.

[0138] As well as patient opinion and / or patient sentiment, the cluster analysis uses identified personal characteristic as additional variables such that the number of independent personal characteristics processed by the clustering analysis is in the range from 10 to 1000 inclusive.

[0139] Figures 5a and 5b show the results of the dimensionality reduction procedure on the aggregate author data. In order to convey the results of the cluster analysis to the user, the data, which has more than 3 dimensions, is reduced to two dimensions. The data could be represented in three dimensions. In addition, the different clusters of patients are displayed in different colours to easily distinguish between the clusters.

[0140] The clustering analysis may be repeated over time to determine whether the results of the clustering analysis change. For example, the clusters may merge, separate or otherwise move depending on the patient sentiment, the patient opinion and personal characteristics of taking a particular medicine.

[0141] Figure 6 is a flowchart of a method of determining one or more properties of a representative author. The centre of the clusters (in the multidimensional space used for the clustering analysis, not the dimension reduced data for display) represents a representative author of that cluster. The representative author of the cluster is a patient subpopulation representative, which is also referred to as a “persona”. The persona defined by a cluster is a generic patient which is representative of patients within the cluster. The one or more properties of the patient subpopulation representative are determined by selecting S600 a random point in the cluster. In an iterative process S606, the distance between each patient within the cluster and the randomly point is determined S602, these distances are then summed S604 and a new random point is selected. The persona of a cluster is defined by the single point within that cluster that has the smallest total distance to each patient. The properties of the persona of that cluster can be used to communicate with users of the system.

[0142] Figure 7 is an example of an output provided by the system. In Figure 7, the system uses a processor to execute a NNLM or provides data to an external NNLM through an API. The one or more properties of PersonaA are provided to the NNLM. The processor uses the NNLM to communicate with the user, which in this case is a healthcare professional (HCP). The user communicates with the system using chat function 700. The user inputs a query into the input bar 720 and sends the query 710 to the system. The system receives the query 710 and uses the NNLM to determine the query and a suitable response. The user has previously selected which patient persona the response to the query should be from, in this case PersonaA. The NNLM uses the one or more properties of PersonaA to provide a response 730 to the user. Typically, the NNLM will also be provided with social media posts of authors who were included in the respective cluster, of which PersonA is representative. This enables the NNLM to provide better output representing typical sentiment and / or opinions of authors in the cluster. The persona may be implemented as an avatar. For example, the persona may be presented to the user as an avatar and the user may be able to interact with the avatar to thereby interact with the persona.

[0143] For example, an Application Programming Interface (API) is used to interface between the software used to receive and output communication with the user and the NNLM. The API receives a textual input. The textual input includes a question received from the user. The textual input also includes an instruction to the NNLM to consider and provide an answer to the question from the user. The textual input also includes one or more extracts from media items which will be used by the NNLM to formulate an answer to the question. The media items provided to the NNLM are collated by an information retrieval system, which retrieves the most relevant extracts of media items amongst a database of curated media items, based on the question from the user.

[0144] The user can also ask 740 what exactly patients are saying about a medicine, in this example DrugA, using the input bar 920. The user can also ask 740 what exactly patients are saying about a medicine, in this example DrugA. The system receives the message 740 and selects a relevant social media post from the central database to provide as an output to the user. The message 750 is the response generated by the system in combination with the NNLM to output the quotation of the social media post from a patient taking DrugA. In a preferred embodiment, the user communicates with the system using Drug-GPT. The output may be provided on any output device. The inputs received from the user may be received using any input device. For example, the output device may be a display screen implemented on any form of device, for example a computer monitor or a portable computing device. The output device may provide other forms of output such as audio output. The input device may be a keyboard and mouse. The input device may accept other forms of input such as audio input. The input and output devices may be provided as a single device such as a touchscreen. The input and output device may be provided on a robot, for example.

[0145] Figure 8 is a schematic of a clustering analysis result analyser 800. The clustering analysis result analyser 800 includes a comparison module 810 which compares results of clustering analysis having different parameters. The comparison module 810 compares the clustering analysis results over time to identify changes in the clustering analysis results, for example the clustering of patients into different groups. The comparison module 810 identifies changing in clustering as a disease progresses or treatment progresses. The comparison module 810 includes a parameter selector 811 which selects different clustering parameters for clustering analysis comparison, a cluster comparison analyser 812 which analyses and compares clustering results based on selected parameters and a change detector 813 which identifies changes in clustering over time, age, and disease or treatment progression.

[0146] The clustering analysis result analyser 800 includes a communication path effectiveness determiner 820 to assess how effective a particular communication path is, or would be, for providing patients with information about a particular medicine. As an example, the clustering analysis result analyser 800 determines a pattern that patients suffering from DiseaseC respond particularly well or prefer messages having a particular form (e.g. tone and style). The communication path effectiveness determiner 820 therefore determines that using messages having a particular tone and style is the most effective way to communicate information about DrugD to patients suffering from DiseaseC. The communication path effectiveness determiner 820 uses characteristics of: source or sender, receiver, channel and statement of interest. The communication path effectiveness determiner 820 includes a personal characteristic extractor 821 which extracts relevant personal characteristics from the aggregate author data. These personal characteristics may be personal characteristics of a persona which is representative of a cluster. The communication path effectiveness determiner 820 also includes a communication path analyser 822 which analyses the effectiveness of different communication paths based on extracted personal characteristics and an effectiveness score calculator 823 which calculates a score for each communication path based on the analysis.

[0147] The clustering analysis result analyser 830 includes a clinical trial progress analyser 830 which measures the success of the progress of a clinical trial for a medicine of interest using the clustering analysis results. As an example, the clustering analysis result analyser 830 uses a model which identifies and analyses the effect of DrugE on patients taking this medicine. The clustering analysis result analyser 830 uses personal characteristics of: patients, drug, physiological elements (age, feelings, symptoms etc), comparison statements of interest to identify and analyse the effect of DrugE. The clinical trial progress analyser 830 can then use this information to determine whether the clinical trial for DrugE is progressing well. The clinical trial progress analyser 830 includes a trial data extractor 831 which extracts relevant trial data and personal characteristics from the aggregate author data. The relevant trial data may include inclusion criteria data, exclusion criteria data, which trial a patient is in, which treatment (e.g. medicine and optionally dosage regime) is being used in the trial etc. The clinical trial progress analyser 830 includes a trial progress analyzer 832 which analyses the progress and success of the clinical trial based on extracted trial data and personal characteristics and a progress measurement calculator 833 which calculates a measurement of the clinical trial progress and success based on the analysis.

[0148] The clustering analysis result analyser 800 includes an assessment module 840 which assesses the quality of life of patients taking a medicine using the clustering analysis results. As an example, the clustering analysis results have a significant cluster of patients taking DrugF in a position on the graph indicating that patient sentiment is positive and patient opinion towards DrugF is that DrugF has made patients feel “fit” and “healthy” (based on comments made in the social media posts analysed in the process of obtaining this clustering analysis). The assessment module 840 uses this information to assess that patients taking DrugF have a good quality of life. The assessment module 840 uses personal characteristics of: patients, multiple drug, side effects, comorbidities, age, behaviour and physical conditions, different settings. The assessment module 840 includes a quality of life data extractor 841 which extracts relevant quality of life data and personal characteristics from the aggregate author data. The assessment module 840 also includes a quality of life analyser 842 which analyses the quality of life of patients based on extracted data and personal characteristics and a quality of life score calculator 843 which calculates a quality-of- life score for patients based on the analysis.

[0149] The clustering analysis result analyser 800 includes a statistical analysis module 850 which performs statistical analysis at the end of each of the methods performed by the modules 810, 820, 830, 840.

[0150] Figure 9 is a schematic of a prediction module 900. The prediction module 900 includes a medicine effect determiner 910 which predicts one or more effects of the medicine of interest on a patient based on results of the clustering analysis. As an example, the clustering analysis result analyser 800 determines a pattern that patients suffering from DiseaseB and having an involvement with DrugC for 1 year are alleviated of the symptom of headaches. The medicine effect determiner 910 uses this pattern to determine that a new patient suffering from DiseaseB and having just started to take DrugC, will be alleviated of the symptom of headaches after taking DrugC for 1 year. The prediction module 900 includes an effect prediction data extractor 911 which extracts relevant data and personal characteristics from the clustering analysis results, an effect prediction analyzer 912 which analyses the potential effects of the medicine on patients based on extracted data, personal characteristics and identified patterns and an effect prediction calculator 913 which calculates the predicted effects of the medicine on patients based on the analysis.

[0151] The prediction module 900 includes a medicine intake determiner 920 which predicts that an individual patient will start to take to take a medicine and when they will start to do so and disease analyser 930 which determines whether and when an individual patient will suffer a particular disease and / or whether a patient is suffering for more than one disease at the same time (i.e. a comorbidity). As an example, the clustering analysis results may show that, for a patient population within the same therapeutic domain, patients can be identified as being at various stages of their journey of disease progression. Each stage requires a different medication in order to treat and / or manage the symptoms of that stage of the disease. The stage each patient is at is inferred by their personal characteristics. The medicine intake determiner 920 predicts whether a patient or a group of patients will start or continue taking a medicine by identifying what stage they are in their patient journey. The medicine intake determiner 920 determines that an individual patient will start taking a medicine and when the individual patient will start taking the medicine by performing pattern analysis on other patients who had the same journey as the individual patient and making a prediction that patients going through the same journey of having a disease will take similar medicines. In this example, the pattern analysis is performed using a combination of personal characteristics of the patients within the same patient subpopulation.

[0152] As a further example, the personal characteristics of the patients falling into subpopulation G (e.g. patients taking DrugG) may be the same as patients falling into another subpopulation H (e.g. patients taking DrugH), except for a single additional personal characteristic (e.g. suffering from DiseaseG) that patients in subpopulation G have that patients in subpopulation H do not. The disease analyser 930 predicts that patients in subpopulation H will progress to DiseaseG. The disease analyser 930 determines that an individual patient will suffer a specific disease, or that a disease which they have will progress to a specific state by performing pattern analysis on other patients who had the same journey as the individual patient and making a prediction that patients going through the same journey of having a disease will develop similar subsequent diseases or disease progression. In this example, the pattern analysis is performed using a combination of personal characteristics of the patients within the same patient subpopulation.

[0153] The medicine intake determiner 920 includes a patient journey data extractor 921 which extracts relevant patient journey data and personal characteristics from the clustering analysis results, a patient journey analyzer 922 which analyses the patient journey data to identify the stage of the patient's journey and an intake prediction calculator 923 which calculates the predicted medicine intake for patients based on the patient journey analysis.

[0154] The disease analyzer 930 includes a disease prediction data extractor 931 which extracts relevant disease data and personal characteristics from the clustering analysis result, a disease prediction analyzer 932 which analyses the disease data to identify potential disease progression or onset based on extracted data and personal characteristics and a disease prediction calculator 933 which calculates the predicted disease progression or onset for patients based on the analysis.

[0155] The prediction module 900 includes a statistical analysis module 940 which performs statistical analysis at the end of each of the methods performed by the modules 910, 920, 930.

[0156] The prediction module 900 includes a communication module 950 which is used to communicate with the information sources 306a, 306b. This allows for the prediction module 900 to use the one or more classifications (e.g. medical ontologies) when making predictions.

[0157] The prediction module 900 may operate independently of the clustering analysis result analyser 800.

[0158] Advantageously, the results of the clustering analysis can be used to predict the ideal marketing campaign (weight, targeting and type of spend) at a medicine level to optimise patient brand reputation / Rx scripts pick up.

[0159] Advantageously, the clustering analysis can provide improved understanding or the ability to predict how to activate patient segmentation by gender, age, location, language within a market through identification of sociodemographic groups for those patients taking medicines and understanding how they behave / how they can be activated / what good patient experience looks like.

[0160] The method may be used to determine a patient confidence score. The patient confidence score provides an indication of the confidence a group of patients has in a medicine. Advantageously, the patient confidence score can be used to determine the impact of weight of ad spend / type of marketing spend and / or be used to predict impact of spend geographically and by representative of the cluster to be able to direct future marketing plans.

[0161] Advantageously, the method can be used to identify and predict the impact of marketing on patient group at a particular point on the diagnosis timeline, i.e. prediagnosed / diagnosed / in treatment, for each medicine of interest.

[0162] Advantageously, the method can be used to build mock marketing campaigns and the machine learning and artificial intelligence models disclosed herein can be used to assess the benefit of said marketing campaign.

[0163] Throughout the description and claims of this specification, the words “comprise” and “contain” and variations of them mean “including but not limited to”, and they are not intended to and do not exclude other components, integers, or steps. Throughout the description and claims of this specification, the singular encompasses the plural unless the context otherwise requires. In particular, where the indefinite article is used, the specification is to be understood as contemplating plurality as well as singularity, unless the context requires otherwise.

[0164] Features, integers, characteristics, or groups described in conjunction with a particular aspect, embodiment, or example of the invention are to be understood to be applicable to any other aspect, embodiment or example described herein unless incompatible therewith. All of the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. The invention is not restricted to the details of any foregoing embodiments. The invention extends to any novel one, or any novel combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of any method or process so disclosed.

Claims

Claims1. A method of identifying patterns in author involvement with a medicine, comprising:(a) accessing a database of media items created by human authors;(b) processing the media items with an author type classifier to identify at least one media item created by an author type;(c) processing said at least one media item with a named entity recognition classifier to identify references in said at least one media item to the medicine;(d) processing said at least one media item referring to the medicine with a sentiment classifier to identify at least one of: (i) an author sentiment regarding said medicine; and (ii) an author opinion regarding the medicine;(e) processing the media items to identify further media items by the author;(f) processing the further media items to identify personal characteristics of the author;(g) repeating steps (b)-(f) for a plurality of authors to create aggregate author data;(h) processing the aggregate author data to perform a clustering analysis in dependence on said author sentiments and / or said author opinions, and said personal characteristics; and(i) analysing the result of the clustering analysis to determine patterns in author involvement with the medicine.

2. A method according to claim 1 , wherein the media items are social media posts and the processing of said at least one social media post with a sentiment classifier identifies both (i) an author sentiment regarding said medicine; and also (ii) an author opinion regarding the medicine, and wherein the processing the aggregate author data to perform a clustering analysis is carried out in dependence on both said author sentiments and said author opinions, as independent variables.

3. A method according to claim 2, wherein the authors are patients, the social media posts are social media posts by patients, the processing with a sentiment classifier identifies both (i) a patient sentiment regarding said medicine; and also (ii) a patient opinion regarding the medicine, and wherein the processing the aggregate patient data to perform a clustering analysis is carried out in dependence on both said patient sentiments and said patient opinions, as independent variables; and whereinanalysing the result of the clustering analysis determines patterns in patient involvement with the medicine.

4. A method according to claim 1 , wherein the further media items by the author comprise or consist of media items which do not mention or reference the medicine.

5. A method according to claim 1 , wherein the identified personal characteristics include the age and / or gender of the author.

6. A method according to claim 4, wherein the analysing of the result of the clustering analysis to determine patterns in author involvement with the medicine comprises identifying changes in the clustering of the authors with time, age and / or disease or treatment progression.

7. A method according to claim 1 , wherein analysing the result of the clustering analysis to determine patterns in patient involvement with the medicine comprises identifying changes in time in the clustering of the authors.

8. A method according to claim 1 , comprising predicting one or more effects of a medicine on an individual patient based on the clustering analysis.

9. A method according to claim 1 , comprising the step of determining one or more properties of a representative author within a cluster identified by the clustering analysis.

10. A method according to claim 8, comprising the step of providing the determined one or more properties of a representative author to one or more processors executing a neural network language model, the one or more processors responding to user queries taking into account the determined one or more properties of the representative author.

11. A method according to claim 1 , comprising the step of assessing the effectiveness of a communication path for communication of information about a medicine responsive to the clustering analysis.

12. A method according to claim 1 , comprising the step of using the results of the clustering analysis to define, measure the success of, and / or control, the progress of a clinical trial for a medicine.

13. A method according to claim 1 , comprising the step of using the results of the clustering analysis to assess one or more of the group consisting of: the safety, efficacy, patient tolerance and accessibility of a medicine, and the quality of life of patients taking a medicine.

14. A method according to claim 1 , comprising the step of using the results of the clustering analysis to predict that an individual patient or a group of patients will start to take, or continue to take, a medicine.

15. A method according to claim 1 , comprising the step of using the results of the clustering analysis to identify patterns in or predict whether and / or when an individual patient will suffer a specific disease, or that a disease which they have will progress to a specific state.

16. A method according to claim 1 , comprising the step of using the results of the clustering analysis to identify patterns in or predict whether and / or when an individual patient or group of patients will start to take and / or continue to take a medicine.

17. A method according to claim 16, wherein the patterns or predictions relate to two or more different medicines.

18. A method according to claim 1 , which does not include the step of processing electronic healthcare records relating to individual patients.

19. A method according to claim 1 , comprising the step of generating aggregate author data which is segregated by one or more demographic characteristics of authors.

20. A method according to claim 1 , comprising the step of processing the further media items by an author to identify one or more words which are relevant to the further media items by the author and taking into account the identified one or more words when performing the clustering analysis.

21. A method according to claim 1 , comprising the step of performing a dimensionality reduction procedure on the aggregate author data which is the subject of the clustering analysis and representing the data in a two or three dimensional representation, with authors in different clusters displayed differently.

22. A method according to claim 1 , comprising the step of receiving a query, the query identifying a cluster output from the clustering analysis, and selecting media items by one or more authors within the identified cluster.

23. A method according to claim 1 , wherein the number of independent personal characteristics processed by the clustering analysis is in the range from 10 to 1000 inclusive.

24. A method according to claim 1 , wherein either the method step of processing the aggregate author data to perform a clustering analysis in dependence on said author sentiments and / or said author opinions, and said personal characteristics and or the method step of analysing the result of the clustering analysis to determine patterns in author involvement with the medicine comprises analysing the further media items by the author to assess alignment between the further media items and a statement of interest.

25. A method according to claim 1 , wherein the personal characteristics comprise one or more classifications selected from a plurality of classifications.

26. A method according to claim 25, wherein the method step of processing the further media items to identify personal characteristics of the author comprises linking content of the further media items to the one or more classifications selected from a plurality of classifications.

27. A method according to claim 25, comprising identifying content of the further media items which is indicative of a type of classification.

28. A method according to claim 27, comprising selecting the one or more classifications, from a plurality of classifications, corresponding to the content the further media items which is indicative of a type of classification.

29. A method of identifying patterns in author alignment with a statement of interest, comprising:(a) accessing a database of media items created by human authors;(b) processing the media items with an author type classifier to identify at least one media item created by an author type;(c) processing the media items to identify further media items by the author;(d) analysing the further media items by the author to assess alignment between the further media items and the statement of interest; and(e) identifying patterns in the author alignment with the statement of interest.