Generative community network builder
The system uses machine learning to create and manage community groups for clinical trials, addressing the challenge of community identification and engagement, enhancing diversity and inclusivity, and providing a centralized management solution with duplicate detection.
Patent Information
- Application Number
- GB2023019736
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-07-02
AI Technical Summary
Clinical trial sponsors face challenges in identifying and engaging diverse communities for clinical trials, leading to difficulties in recruitment and inclusivity, and there is a lack of efficient systems for managing community groups and avoiding duplication.
A computer-implemented method and system for creating and managing community groups using machine learning models to identify patient attributes, forming groups based on common attributes, and a cloud-based system to manage and search these groups, with a duplicate detection mechanism to avoid redundancy.
Enhances community engagement, promotes diversity and inclusion, improves recruitment, and provides a centralized system for managing community groups, ensuring accurate identification and reducing duplication.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field of the disclosure The present disclosure relates to a computer-implemented method for the processing and / or creation of community groups for clinical trials and community engagement. Background Clinical trials provide evidence for evaluation whether a medicament is safe and effective. To ensure that any clinical trial is effective, it needs to be tested on a community for which the medicament is intended. This is even more relevant where populations are becoming increasingly diverse. Community-based clinical trials may offer tremendous value. For example, community-based clinical trials can help with diversity and recruitment, two areas where trials often struggle. Community engagement is also extremely important. Community engagement is desirable because it ensures participation is inclusive and scalable. However, it is often difficult for clinical trial sponsors to identify communities within which to perform the clinical trial, and even if those communities are identified, for there to be real engagement with those communities. Aspects of the disclosure seek to address problems in selecting and identifying communities relevant for clinical trials. Summary of the invention Aspects of the invention are as set out in the independent claims and optional features are set out in the dependent claims. Aspects of the invention may be provided in conjunction with each other and features of one aspect may be applied to other aspects. Aspects of the disclosure provide a computer-implemented method and system for creating and managing community groups for clinical trials and community engagement. Community groups are groups that a clinical trial may recruit from. Improving community engagement may mean that improved relationships may be formed, and diversity, equity and inclusion promoted. Furthermore, it may allow patients / participants with complex needs to be included. Community engagement also means that other trials sponsors can communicate and share ideas with each other, enabling more inclusive and diverse multi level engagement, reaching patients / partici pants that would not otherwise be able to participate. Furthermore, managing community groups may enable tracking of community engagement to be monitored. Advantageously this may provide a single system that manages the whole lifecycle of community groups and engagement. The inventors have recognised that it may be desirable to create “smart” networks of community / diversity groups - for example, a user (such as a clinical trial sponsor) may want to be able to enter a keyword like “cancer” and find all community groups relating to cancer and want to be able to filter down within these. This enables users to better select customers / users / patients. To be able to identify and filter down within community groups, community groups preferably associate patients or groups of patients with attributes - such as therapeutic areas, diseases, and symptoms. A user may then be able to filter down community groups based on one or more attributes. The inventors have also recognised that it may be desirable to avoid duplication of community groups, so that users (such as clinical trial sponsors) may be more effectively and accurately identify the correct community group for their desired use case. Aspects of the disclosure may address both these problems. Accordingly, in a first aspect there is provided a computer-implemented method of creating community groups for clinical trials and community engagement, each community group comprising a plurality of patients grouped based on a common attribute. The method comprises obtaining at least one source of information relating to patients. In the event that the at least one source of information is in an unstructured format, the method comprises converting the information to a structured format. The method then comprises applying a machine learning model to the information to determine one or more attributes of the patients from the information and grouping the patients into community groups based on the determined attributes. The structured format may comprise a json / xml / csv data type. The term community groups may be understood to encompass groups that share a common attribute that a clinical trial may recruit from. Applying a machine learning model to the information to determine attributes of the patients from the information may comprise applying at least one of (i) a large language model, LLM, and (ii) a natural language processing, NLP, model, to the information. Applying a machine learning model to the information to determine attributes of the patients from the information may comprise determining health conditions of the patients. Applying a machine learning model to the information to determine attributes of the patients from the information may comprise identifying patients, or a subset of patients, from the information. The method may further comprise associating each of the identified patients with attributes. The attributes may comprise at least one of: symptoms, demographics, therapeutic areas, non-clinical areas, minimum group size, maximum group size, category. A non-exhaustive list of categories may include: charity, advocacy, influencer, support, hospice, religious group, government, hospital or practice, medical support, social care, other, corporate, fitness group, neighbourhood, pharmacy. Min and max group size may be used to search for community groups based on the group size, for example when a user is using the advanced search on the front end. The obtained information relating to patients may comprise any of: website information, marketing collateral, surveys, clinical trial master file, electronic clinical trial master file, documentation relating to a clinical trial. Grouping the patients into community groups based on the determined attributes may comprise grouping each patient into a plurality of community groups. For example, a patient may be in both a “stroke” community group, and a “high blood pressure” community group. The computer-implemented method may further comprise creating a searchable database of community groups based on the grouping of patients into community groups, wherein the searchable database of community groups is searchable based on determined attributes. The computer-implemented method may further comprise ascribing at least one feature to each community group in the searchable database based on the attribute or attributes in common for that community group. For example, the at least one feature may comprise community group name, community group contact details, community group website, community group address, which may all be determined based on the attributes. The computer-implemented method may further comprise parsing the database of community groups to check for duplicates. Parsing the database of community groups to check for duplicates may comprise applying a machine learning model and process automation to the database of community groups to determine community groups in the database that have at least one feature in common; and in the event that features are found to be common between at least two community groups, providing an indication that at least one of the at least two community groups is a duplicate. Providing an indication may comprise providing an alert to a user via a graphical user interface that at least one of the at least two community groups is a duplicate. This may comprise flagging both community groups as being identical and asking the user to select which of these should be retained and which one should be deleted. In some examples the community groups may be merged. Features may be found to be in common when at least one of: (i) the website links have the same domain, (ii) email addresses have the same domain, (iii) the name of the community group is identical, (iv) the size of the community group is identical. Features may additionally, or alternatively, be found to be in common when (i) the website links are similar, (ii) email addresses are similar, (iii) the name of the community groups are similar, (iv) the size of the community group are similar. Being similar may comprise at least a portion being identical, for example a portion of the name, email address or website link may be identical but not the entire name, email address or website link. In another aspect there is provided a computer-implemented method of determining duplicates in a database of community groups for a clinical trial. The method comprises obtaining a database of community groups, wherein each community group comprises a plurality of patients having at least one attribute in common, and wherein each community group has at least one feature; parsing the database of community groups to determine community groups that have at least one feature in common; and in the event that at least one feature is found in common between at least two community groups, providing an indication that at least one of the at least two community groups is a duplicate. Determining community groups that have at least one feature in common may comprise at least one of: (i) determining community groups that have at least one identical feature in common; and (ii) determining community groups that have at least one similar feature. Similar features may be determined using a machine learning model, for example groups that have at least one similar feature may be groups that have a similar name. Additionally, or alternatively, community groups that have at least one similar feature may comprise community groups that have a portion of the feature that is identical, such as American Foundation for the Blind, and American Foundation for the Blind (AFB). Determining community groups that have at least one feature in common may comprise determining community groups that have at least one identical feature in common, and at least one other feature that is similar in common. In another aspect there is provide a system for creating and managing community groups for clinical trials and community engagement, the system comprising: a web application module; a community group manager module; a duplicate manager module; a cloud-based database; and a cloud-based search module configured to search the cloud-based database. The web-application module is configured to interact with the community group manager module and provide a user interface; the community group manager module is configured to communicate with the cloud-based database to enable community groups to be added, edited or deleted from the cloud-based database interface, and to communicate with the duplicate manager module; wherein the duplicate manager module is configured to check for duplicates of community groups in the cloud-based database; and wherein the cloud-based search module is configured to enable searches of community groups in the cloud-based database. The cloud-based database may be configured to store community groups in the database along with at least one feature and at least one attribute, and wherein the could-based search module is configured to enable searches of the cloud-based database based on at least one of: (i) features, and (ii) attributes. In another aspect there is provided a computer readable non-transitory storage medium comprising a program for a computer configured to cause a processor to perform the method of any of the aspects described above. Drawings Embodiments of the disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which: Fig. 1 shows an example screenshot of a graphical user interface for use in a computer-implemented method of creating and managing community groups. Fig. 2 shows another example screenshot of a graphical user interface for use in a computer-implemented method of creating and managing community groups. Fig. 3 shows another example screenshot of a graphical user interface for use in a computer-implemented method of creating and managing community groups. Fig. 4 shows another example screenshot of a graphical user interface for use in a computer-implemented method of creating and managing community groups. Fig. 5 shows another example screenshot of a graphical user interface for use in a computer-implemented method of creating and managing community groups. Fig. 6 shows an exemplary Level 1 architecture diagram of a system for use with a computer-implemented method of creating and managing community groups. Fig. 7 shows another exemplary Level 1 architecture diagram of a system for use with a computer-implemented method of creating and managing community groups. Fig. 8 shows another exemplary Level 1 architecture diagram of a system for use with a computer-implemented method of creating and managing community groups. Fig. 9 shows a process flow diagram 900 of an exemplary computer-implemented method of creating and managing community groups. Fig. 10 shows an exemplary process flow diagram of a computer-implemented method of creating and managing community groups. Fig. 11 shows an exemplary process flow diagram of a computer-implemented method of creating and managing community groups. Fig. 12 shows an exemplary Level 1 architecture diagram of a computer system for running a computer-implemented method of creating and managing community groups. Soecific description Embodiments of the claims relate to a computer-implemented method and system for creating and managing community groups for clinical trials and community engagement. Community groups are groups that a clinical trial may recruit from. Fig. 1 shows an example screenshot of a graphical user interface 100 for use in a computer-implemented method of creating and managing community groups. The graphical user interface 100 shows a plurality of different community groups in a tabular format. The graphical user interface 100 comprises a search box 101 for searching a database of community groups for community groups having something in common, such as an attribute in common. The attributes may comprise at least one of: symptoms, demographics, therapeutic areas, non-clinical areas, minimum group size, maximum group size, category. A non-exhaustive list of categories may include: charity, advocacy, influencer, support, hospice, religious group, government, hospital or practice, medical support, social care, other, corporate, fitness group, neighbourhood, pharmacy. Min and max group size may be used to search for community groups based on the group size, for example when a user is using an advanced search feature as will be discussed in more detail below with respect to Fig. 4. Accordingly, the user may be able to search for community groups by searching for a therapeutic area or disease (e.g., “stroke”). The community groups displayed in the graphical user interface 100 include information such as the community group name 103, an email address for the community group 105, the category 107 of the community group, the size of the community group 109, the location 111 of the community group (such as the city and state or country), an indication of whether the community group is a funder 113 or not, a link to the website 115 of the community group, and options 117 to edit or delete a community group. This information may be referred to as features of the community group. The user of the graphical user interface 100 may also be able to filter community groups based on any of these features and / or attributes and / or perform an advanced search, for example based on two or more features and / or attributes (such as therapeutic area and group size). Community groups may be entered or imported manually, automatically, and / or a combination of manually and automatically. Fig. 2 shows another example screenshot 200 of a graphical user interface for use in a computer-implemented method of creating and managing community groups. Fig. 2 shows a screenshot 200 of the graphical user interface where a user may be able to enter in select details about the community group, such as name, email, alternate email, location, address, zip code, category, description, website, size and whether it is a funder or not. In some examples, a user may enter in only some of these details, and a machine learning method may be applied to determine any missing fields. For example, if no category or conditions are entered, the machine learning mode may determine this information from other information sources. For example, if a website is entered, the machine learning model may be configured to trawl that website and produce a description of the community group and / or determine categories for that community group. Community groups may be entered or imported automatically from e.g., Excel® or csv files. In this way, a searchable database of community groups may be created, wherein the searchable database of community groups is searchable based on determined attributes. A machine learning model may also be used to determine attributes of the community group, and / or associate community groups and / or specific patients within those community groups with attributes. It will be understood that patients may be a member of more than one community group. Because the attributes may include symptoms, demographics, therapeutic areas, non-clinical areas, minimum group size, maximum group size, and / or category, applying a machine learning model to the information to determine attributes of the patients from the information may therefore comprise determining health conditions of the patients. The machine learning model may be configured to determine attributes, for example, based on other information already entered for that community group. For example, if a website is entered, the machine learning model may be configured to trawl that website and determine attributes of the community group from information on that website. It will be appreciated that the information on that website may be freetext. The machine learning model may therefore be configured to determine attributes of the community group based on freetext information. Additionally, or alternatively, if a description of the community group is provided, the machine learning model may be configured to determine attributes of the community group based on the freetext information provided in the description. The obtained information relating to patients comprises any of: website information, social media profile or page, marketing collateral, surveys, clinical trial master file, electronic clinical trial master file, documentation relating to a clinical trial. Accordingly, the machine learning model may be configured to perform natural language processing and / or perform word organisation and / or be a large language model, which will be known to a person skilled in the art. In some examples a combination of different algorithms or machine learning models may be used. For example, key detection may be used to analyse information to find relevant sections of the information, which may then be processed using natural language processing (NLP) word tokenizer algorithms or Large Language Models (LLM) may be used to optimize the performance and the accuracy of the results, and to further segregate the data according to what output is needed. In some examples, the machine learning model run on AWS® Lamba®. For example, the user may input a website for a community group and the system automatically acquires unstructured data (freetext) from that website. The information or unstructured data may then be parsed, and algorithms automatically execute an algorithm that converts the unstructured data into structured data. This may include, for example, using an OCR model (such as Tesseract OCR) to sequence inputs. For example, the PyTesseract library may be used for OCR. Another machine learning model may (such as a long short term memory, LSTM, model) then be used to take the recognised text as inputs and outputs. For example, when a user enters a website / URL, the basic details against the website may be extracted from the meta data of that URL, that may include but not limited to, the description of the website, name or any image or logo related to that website / URL. In other examples, the user may input a social media profile or page for a community group, and the system automatically acquires unstructured data (freetext) from that website. The freetext may be transformed using a number of known machine learning methods including (but not limited to) bert, biobert, fasttext and named entity recognition (NER) with performance on a given feature set evaluated downstream of the classification task. The natural language processing model may include, for example bert, biobert and / or fasttext. In some examples a combination of natural language processing models may be used. The natural language processing models may be pre-trained, for example on clinical datasets. For example, applying a natural language processing model comprises applying a plurality of natural language processing models, comprising a first specialised model trained on text available from the data sources, with a second general model trained on Wikipedia. In some examples, applying a natural language processing model comprises applying BiObert (a model pretrained on biomedical text) which converts into a 768 dimension vector from input biomedical free text and a modified version of fasttext, which combines a specialised model trained on the available free-text with a generalised model trained on Wikipedia which converts both to respective 300 dimension vectors. Accordingly, the graphical user interface shown in Figs. 1 and 2 may be used to perform a method of creating community groups for clinical trials and community engagement, as shown for example in the exemplary process flow diagram of Fig. 10. Each community group may comprise a plurality of patients grouped based on a common attribute. The method 900 shown in Fig. 10 comprises obtaining 910 at least one source of information relating to patients. This may be, for example, from a description of a community group entered by a user, or another source of information related to information provided by the user, such as a website. The method then comprises converting 920 the information to a structured format in the event that the at least one source of information is in an unstructured format. The method then comprises applying 930 a machine learning model to the information to determine one or more attributes of the patients from the information and grouping 940 the patients into community groups based on the determined attributes. This may involve creating a searchable database of community groups based on the grouping of patients into community groups, wherein the searchable database of community groups is searchable based on determined attributes. It will be understood that it may be undesirable to have duplicates in the database of community groups. Accordingly, in some examples the method may comprise checking for duplicates. This is shown in Figs. 3 and 10. Fig. 3 shows another example screenshot of a graphical user interface 300 for use in a computer-implemented method of creating and managing community groups, and Fig. 11 shows an exemplary process flow diagram of a computer-implemented method of creating and managing community groups, and in particular for detecting and removing duplicates. The graphical user interface 300 shown in Fig. 3 presents the user with a list of community groups that are suspected as being duplicates, as determined automatically by the system. This may be performed by a duplicate manager. The duplicate manager may nest entries together that it believes to be duplicates. Duplicates may be determined based on: - name the same; - contact details (e.g., emails) the same; - website the same; - address the same; and - other identity-related parameters / inputs. The duplicate manager may output e.g., a list of possible duplicates that the user can either approve or reject as duplicates, as shown in the graphical user interface 300 of Fig. 3. The duplicate manager may do this by first obtaining 1010 a database of community groups, wherein each community group comprises a plurality of patients having at least one attribute in common, and wherein each community group has at least one feature. The duplicate manager may ascribe at least one feature to each community group in the searchable database based on the attribute or attributes in common for that community group. Additionally or alternatively, features may have already been entered at the time the community group was created, for example by the user. The feature may comprise, for example, community group name, community group contact details (such as telephone number and / or email address), community group website, community group address, which may all be determined based on the attributes. The duplicate manager may then parse 1020 the database of community groups to determine community groups that have at least one feature in common. Parsing 1020 the database of community groups to check for duplicates may comprise applying a machine learning model and process automation to the database of community groups to determine community groups in the database that have at least one feature in common. In the event that features are found to be common between at least two community groups, the duplicate manager provides 1030 an indication that at least one of the at least two community groups is a duplicate. Features may be found to be in common when at least one of: (i) the website links have the same domain, (ii) email addresses have the same domain, (iii) the name of the community group is identical, (iv) the size of the community group is identical. As noted above, in some examples the user may be able to perform an advanced search on community groups, as shown in the graphical user interface 400 of Fig. 4. The user may be able to search community groups based on attributes. For example, the user may be able to search community groups based on any of the following: symptoms, therapeutic areas, group size, demographics, non-clinical areas, category, funders. A user may also be able to request a new condition / attributes be added, and / or add a community group via the interface 400. The results of these searches may be shown in another graphical user interface, such as the graphical user interface 500 of Fig. 5. As well as listing the results of the search in a list / tabular format, the graphical user interface 500 also presents the user with the ability to build relationships between community groups. For example, the user may send a request via a web application, which sends an https request to an API application to either create actions for each community group and / or set permissions for each community group. These actions and / or permissions may be saved in a database. It will be understood that the web application may be the web application 605 shown in Fig. 7, the API application may be the API application 755 shown in Fig. 7, and the database may be the cloud-based database 759 shown in Fig. 7, and as described in more detail below. Fig. 6 shows an exemplary Level 1 architecture diagram of an exemplary system 600 for use with a computer-implemented method of creating and managing community groups, such as the method above described with respect to Figs. 1 to 5. The system 600 shows a web application that may act as an interface between the user 603 and the community group system 607. The web application 605 is an angular framework that delivers static content and features of the graphical user interface. The community group system manages the community groups database. The user 603 may make a request to the web application 605 via a browser, and the web application 605 may interact with the community group system 607 by validating the permission of the user / employee, authenticate the user / employee, and give them access to the community group system 607. Fig. 7 shows another exemplary Level 1 architecture diagram of a system 700 for use with a computer-implemented method of creating and managing community groups. The system 700 comprises the system 600 of Fig. 6 and shows how it interacts with the duplicate manager discussed above. The system 700 shown in Fig. 7 comprises a web application module 605 communicatively coupled to a community group manager module 607, an API application 755 communicatively coupled to the community group manager module 607, a duplicate manager module 750 communicatively coupled to a cloud-based database 759 and a cloud-based search module 753 communicatively coupled to the cloud-based database 759 and configured to search the cloud-based database. The web-application module 605 is configured to interact with the community group manager module 607 and provide a user interface, such as the user interface discussed above with respect to Figs. 1 to 5. The community group manager module 607 is configured to communicate with the cloudbased database 759 via the API application 755 to enable community groups to be added, edited or deleted from the cloud-based database 759, and to communicate with the duplicate manager module 750. The duplicate manager module 750 is configured to check for duplicates of community groups in the cloud-based database 759. The cloud-based search module 753 is configured to enable searches of community groups in the cloud-based database 759, which may be AWS® DynamoDB®, although it will be understood by a person skilled in the art that other database technologies may be used. The cloud-based search module 753 may, for example, be AWS® OpenSearch® and configured to perform an elastic search, although the skilled person will understood that other smart search databases may be used. The cloud-based database 759 may be configured to store community groups in the database along with at least one feature and at least one attribute, and wherein the could-based search module is configured to enable searches of the cloud-based database based on at least one of: (i) features, and (ii) attributes. Fig. 8 shows another exemplary Level 1 architecture diagram of a system 800 for use with a computer-implemented method of creating and managing community groups. The system 800 comprises the system 600 of Fig. 6 and the system of 700 of Fig. 7 and shows how it is broken down into an “Internal Admin Panel”, “Backend” and “Frontend” or “External User”. As described above, the methods described herein may involve converting unstructured data (i.e., freetext from a description or website, for example) into structured data. This may involve a transformation into a json or xml format, for example. The data may be multi-hierarchical, hence tabular results would not establish the multi-level relationships. The unsorted data is transformed into sorted data (for example, via machine learning models described below). Fig. 9 shows a process flow diagram 900 of an exemplary computer-implemented method of creating and managing community groups, for example using the system of Fig. 8. The process 900 begins at step 901 where a user, such as sponsor, wishes to add or update a community group. This initiates a process automation model and data structureisation, which performs a duplicate check 903. In the event that a duplicate is identified, then the process resolves, deletes or updates 905 the duplicate. Alternatively, if no duplicate is found, then the process automation and data structureisation model updates the count, stores the record and triggers a sync 907. This results in the record being stored in a database, such as the database 759 described above. The sync triggers a level 1 quality assurance check that checks the existence of a community group name 909 and checks the existence of a community group description 911 (community group name and description is required for prediction, hence its checking their existence and if so it goes forward with the prediction, else it skips it). Either a user can enter these details at step 913, or a first machine learning model, which may for example be a large language model, may predict 915 the health assets (i.e., the relevant health conditions / therapeutics etc.), and a second machine learning model, which may for example be a natural language processing model, may predict 917 the demographic assets (i.e., the relevant demographics of the members of that community group). As part of this, health conditions data may be obtained from the database. As an output from this, a level 2 quality assurance check is performed, which either structurizes 919 the user provided health condition data or structurizes 921 the user provided and / or predicted data. The outputs from these are then stored in a separate database, such as an elastic database. A third machine learning model, which may be a machine learning micro-service script, running on an HTTP server, may accumulate / request / response interface 925 (which may for example be the API application 755 described above with reference to Fig. 7) and an advanced search query generator 923, operable via an interface 927 for search for community groups via a web application (which may, for example, be the web application SPA 605 described above with reference to Fig. 7). It will be understood that any of the machine learning models described herein may comprise a neural network. The neural network may comprise at least one of a deep residual network, a highway network, a densely connected network and a capsule network. For any such type of network, the network may comprise a plurality of different neurons, which are organised into different layers. Each neuron is configured to receive input data, process this input data and provide output data. Each neuron may be configured to perform a specific operation on its input, e.g. this may involve mathematically processing the input data. The input data for each neuron may comprise an output from a plurality of other preceding neurons. As part of a neuron’s operation on input data, each stream of input data (e.g. one stream of input data for each preceding neuron which provides its output to the neuron) is assigned a weighting. That way, processing of input data by a neuron comprises applying weightings to the different streams of input data so that different items of input data will contribute more or less to the overall output of a neuron. Adjustments to the value of the inputs for a neuron, e.g. as a consequence of the input weightings changing, may result in a change to the value of the output for that neuron. The output data from each neuron may be sent to a plurality of subsequent neurons. The neurons are organised in layers. Each layer comprises a plurality of neurons which operate on data provided to them from the output of neurons in preceding layers. Within each layer there may be a large number of different neurons, each of which applies a different weighting to its input data and performs a different operation on its input data. The input data for all of the neurons in a layer may be the same, and the output from the neurons will be passed to neurons in subsequent layers. The exact routing between neurons in different layers forms a major difference between capsule networks and deep residual networks (including variants such as highway networks and densely connected networks). For a residual network, layers may be organised into blocks, such that the network comprises a plurality of blocks, each of which comprises at least one layer. For a residual network, output data from one layer of neurons may follow more than one different path. For conventional neural networks (e.g. convolutional neural networks), output data from one layer is passed into the next layer, and this continues until the end of the network so that each layer receives input from the layer immediately preceding it and provides output to the layer immediately after it. However, for a residual network, a different routing between layers may occur. For example, the output from one layer may be passed on to multiple different subsequent layers, and the input for one layer may be received from multiple different preceding layers. In a residual network, layers of neurons may be organised into different blocks, wherein each block comprises at least one layer of neurons. Blocks may be arranged with layers stacked together so that the output of a preceding layer (or layers) feeds into the input of the next block of layers. The structure of the residual network may be such that the output from one block (or layer) is passed into both the block (or layer) immediately after it and at least one other later subsequent block (or layer). Shortcuts may be introduced into the neural network which pass data from one layer (or block) to another whilst bypassing other layers (or blocks) in between the two. This may enable more efficient training of the network, e.g. when dealing with very deep networks, as it may enable problems associated with degradation to be addressed when training the network (which is discussed in more detail below). The arrangement of a residual neural network may enable branches to occur such that the same input provided to one layer, or block of layers, is provided to at least one other layer, or block of layers (e.g. so that the other layer may operate on both the input data and the output data from the one layer, or block of layers). This arrangement may enable a deeper penetration into the network when using back propagation algorithms to train the network. For example, this is because during learning, layers, or blocks of layers, may be able to take as an input, the input of a previous layer / block and the output of the previous layer / block, and shortcuts may be used to provide deeper penetration when updating weightings for the network. For a capsule network, layers may be nested inside of other layers to provide ‘capsules’. Different capsules may be adapted so that they are more proficient at performing different tasks than other capsules. A capsule network may provide dynamic routing between capsules so that for a given task, the task is allocated to the most competent capsule for processing that task. For example, a capsule network may avoid routing the output from every neuron in a layer to every neuron in the next layer. A lower level capsule is configured to send its input to a higher level (subsequent) capsule which is determined to be the most likely capsule to deal with that input. Capsules may predict the activity of higher layer capsules. For example, a capsule may output a vector, for which the orientation represents properties of an object in question. In response, each subsequent capsule may provide, as an output, a probability that the object that capsule is trained to identify is present in the input data. This information (e.g. the probabilities) can be fed back to the capsule, which can then dynamically determine routing weights, and forward the input data to the subsequent capsule most likely to be the relevant capsule for processing that data. For either type of neural network, there may be included a plurality of different layers which have different functions. The neural network may include at least one convolutional layer configured to convolve input data across its height and width. The neural network may also have a plurality of filtering layers, each of which comprises a plurality of neurons configured to focus on and apply filters to different portions of the input data. Other layers may be included for processing the input data such as pooling layers (to introduce non-linearity) such as maximum pooling and global average pooling, Rectified Linear Units layer (ReLU) and loss layers, e.g. some of which may include regularization functions. The final block of layers may receive input from the last output layer (or more layers if there are branches present). The final block may comprise at least one fully connected layer. The final output layer may comprise a classifier, such as a softmax, sigmoid or tanh classifier. Different classifiers may be suitable for different types of output; for example, a sigmoid classifier may be suitable where the output is a binary classifier. The output of the neural network may provide an indication of a probability. Inputs may be fed into a set of 3D layers in the neural network. There are several features of this network which may be varied as training of the network proceeds. For each neuron, there may be a plurality of weightings, each of which is applied to a respective input stream for output data from neurons in preceding layers. These weightings are variables which can be modified to provide a change to the output of the neural network. These weightings may be modified in response to training so that they provide more accurate data. In response to having trained these weightings, the modified weightings are referred to as having been ‘learned’. Additionally, the size and connectivity of the layers may be dependent upon the typical input data for the network; although, these too may be a variable which may be modified and learned during training, including the reinforcement of connections. To train the network, e.g. to learn values for the weightings, these weightings are assigned an initial value. These initial values may essentially be random; however, to improve training of the network, a suitable initialisation for the values may be applied such as a Xavier / Glorot initialisation. Such initialisations may inhibit situations from occurring in which initial random weightings are too great or too small, and the neural network can never properly be trained to overcome these initial prejudices. This type of initialisation may comprise assigning weightings using a distribution having a zero mean but a fixed variance. Once the weightings have been assigned, training object data may be fed or input into the neural network. This may comprise operating the neural network on the results of previous clinical trial adjudication decisions, as referenced above when describing Fig. 6. Algorithms such as mini-batch gradient descent, RMSprop, Adam, Adadelta and Nesterov may be used during this process. This may enable an identification of how much each different point (neuron) or path (between neurons in subsequent layers) in the network is contributing to determining an incorrect score, thus enabling a determination of any weight adjustments that need to be made. The weightings may then be adjusted according to the error calculated. For example, to minimise or remove the contribution from neurons which contribute, or contribute the most, to an incorrect determination. After an iteration of training the network, the weightings may be updated, and this process may be repeated a large number of times. To inhibit the likelihood of overtraining the network, training variables such as learning rate and momentum may be varied and / or controlled to be at a selected value. Additionally, regularisation techniques such as L2 or dropout may be used which reduce the likelihood of different layers becoming over-trained to be too specific for the training data, without being as generally applicable to other, similar data. Likewise, batch normalisation may be used to aid training and improve accuracy. In general, the weightings are adjusted so that the network would, if operated on the same training data again, produce the expected outcome. Although, the extent to which this is true will be dependent on training variables such as learning rate. It is to be appreciated that increasing the depth of neural networks may cause problems when training, e.g. due to vanishing gradient problems, and it may also provide slower networks. However, the present disclosure may enable the provision of a network having increased depth and accuracy without sacrificing the ability to adequately train the network. The depth of the network used may be selected to provide a balance between accuracy and the time taken to provide an output. Increasing the depth of the network may provide increased accuracy although it may also increase the time taken to provide an output. Use of a branched structure (as opposed to in a convolutional neural network) may enable sufficient training of the network to occur as depth of the network increases, which in turn provides for an increased accuracy of the network. FIG. 12 is a block diagram of a computer system 1200 suitable for implementing one or more embodiments of the present disclosure, including for example any of the methods described above with reference to Figs. 1 to 5 and 9 and 10. The computer system 1200 may form the function of any of the elements of the system described in Figs. 6, 7 and 8, although it will be appreciated that many of those elements (such as, for example, the web application 605, community group manager 607, API application 755, duplicate manager 750, database 759 and / or search module 753 may be cloud-based, as would be known to a person skilled in the art. The computer system 1200 includes a bus 1212 or other communication mechanism for communicating information data, signals, and information between various modules of the computer system 1200. The modules include an input / output (I / O) module 1204 that processes a user (i.e., sender, recipient, service provider) action, such as selecting keys from a keypad / keyboard, selecting one or more buttons or links, etc., and sends a corresponding signal to the bus 1212. It may also include a camera for obtaining image data. The I / O module 1204 may also include an output module, such as a display 1202 and a cursor control 1208 (such as a keyboard, keypad, mouse, etc.). The display 1202 may be configured to present a login page for logging into a user account, and is configured to display a user interface such as the user interface described above with reference to Figs. 1 to 5. An optional audio input / output module 1206 may also be included to allow a user to use voice for inputting information by converting audio signals. The audio I / O module 1206 may allow the user to hear audio. A transceiver or network interface 1220 transmits and receives signals between the computer system 1200 and other devices, such as another user device, a merchant server, or a service provider server via network 1222. In one embodiment, the transmission is wireless, although other transmission mediums and methods may also be suitable. A processor 1214, which can be a microcontroller, digital signal processor (DSP), or other processing module, processes these various signals, such as for display on the computer system 1200 or transmission to other devices via a communication link 1224. The processor 1214 may also control transmission of information, such as cookies or IP addresses, to other devices. The modules of the computer system 1200 also include a system memory module 1210 (e.g., RAM), a static storage module 1216 (e.g., ROM), and / or a disk drive 1218 (e.g., a solid-state drive, a hard drive). The computer system 1200 performs specific operations by the processor 1214 and other modules by executing one or more sequences of instructions contained in the system memory module 1210. For example, the processor 1214 can run the systems 100, 200 described above. It will also be understood that any or all of the modules described above with reference to Figs. 6 to 8 may be implemented in software or hardware, for example as dedicated circuitry. For example, the modules may be implemented as part of a computer system. The computer system may include a bus or other communication mechanism for communicating information data, signals, and information between various modules of the computer system. The modules may include an input / output (I / O) module that processes a user (i.e., sender, recipient, service provider) action, such as selecting keys from a keypad / keyboard, selecting one or more buttons or links, etc., and sends a corresponding signal to the bus. The I / O module may also include an output module, such as a display and a cursor control (such as a keyboard, keypad, mouse, etc.). A transceiver or network interface may transmit and receive signals between the computer system and other devices, such as another user device, a merchant server, or a service provider server via a network. In one embodiment, the transmission is wireless, although other transmission mediums and methods may also be suitable. A processor, which can be a micro-controller, digital signal processor (DSP), or other processing module, processes these various signals, such as for display on the computer system or transmission to other devices via a communication link. The processor may also control transmission of information, such as cookies or IP addresses, to other devices. The modules of the computer system may also include a system memory module (e.g., RAM), a static storage module (e.g., ROM), and / or a disk drive (e.g., a solid-state drive, a hard drive). The computer system performs specific operations by the processor and other modules by executing one or more sequences of instructions contained in the system memory module. Logic may be encoded in a computer readable medium, which may refer to any medium that participates in providing instructions to a processor for execution. Such a medium may take many forms, including but not limited to, non-volatile media, volatile media, and transmission media. In various implementations, non-volatile media includes optical or magnetic disks, volatile media includes dynamic memory, such as a system memory module, and transmission media includes coaxial cables, copper wire, and fibre optics. In one embodiment, the logic is encoded in non-transitory computer readable medium. In one example, transmission media may take the form of acoustic or light waves, such as those generated during radio wave, optical, and infrared data communications. Some common forms of computer readable media includes, for example, floppy disk, flexible disk, hard disk, magnetic tape, any other magnetic medium, CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with patterns of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer is adapted to read. In various embodiments of the present disclosure, execution of instruction sequences to practice the present disclosure may be performed by a computer system. In various other embodiments of the present disclosure, a plurality of computer systems 1200 coupled by a communication link to a network (e.g., such as a LAN, WLAN, PTSN, and / or various other wired or wireless networks, including telecommunications, mobile, and cellular phone networks) may perform instruction sequences to practice the present disclosure in coordination with one another. It will also be understood that aspects of the present disclosure may be implemented using hardware, software, or combinations of hardware and software. Also, where applicable, the various hardware modules and / or software modules set forth herein may be combined into composite modules including software, hardware, and / or both without departing from the spirit of the present disclosure. Where applicable, the various hardware modules and / or software modules set forth herein may be separated into sub-modules including software, hardware, or both without departing from the scope of the present disclosure. In addition, where applicable, it is contemplated that software modules may be implemented as hardware modules and vice-versa. Software in accordance with the present disclosure, such as program code and / or data, may be stored on one or more computer readable mediums. It is also contemplated that software identified herein may be implemented using one or more general purpose or specific purpose computers and / or computer systems, networked and / or otherwise. Where applicable, the ordering of various steps described herein may be changed, combined into composite steps, and / or separated into sub-steps to provide features described herein. The various features and steps described herein may be implemented as systems including one or more memories storing various information described herein and one or more processors coupled to the one or more memories and a network, wherein the one or more processors are operable to perform steps as described herein, as non-transitory machine-readable medium including a plurality of machine-readable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform a method including steps described herein, and methods performed by one or more devices, such as a hardware processor, user device, server, and other devices described herein. Reference throughout this specification to “some embodiments,” or “an embodiment,” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in some embodiment,” or “in an embodiment,” in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. As utilized herein, terms “module,”, “controller”, “module”, “system,” “interface,” “unit” and the like are intended to refer to a computer-related entity, hardware, software (e.g., in execution), and / or firmware. For example, a module can be a processor, a process running on a processor, an object, an executable, a program, a storage device, and / or a computer. By way of illustration, an application running on a server and the server can be a module. One or more modules can reside within a process, and a module can be localized on one computer and / or distributed between two or more computers. Further, these modules can execute from various computer readable media having various data structures stored thereon. The modules can communicate via local and / or remote processes such as in accordance with a signal having one or more data packets (e.g., data from one module interacting with another module in a local system, distributed system, and / or across a network, e.g., the Internet, a local area network, a wide area network, etc. with other systems via the signal). As another example, a module can be an apparatus with specific functionality provided by mechanical parts operated by electric or electronic circuitry; the electric or electronic circuitry can be operated by a software application or a firmware application executed by one or more processors; the one or more processors can be internal or external to the apparatus and can execute at least a part of the software or firmware application. As yet another example, a module can be an apparatus that provides specific functionality through electronic modules without mechanical parts; the electronic modules can include one or more processors therein to execute software and / or firmware that confer(s), at least in part, the functionality of the electronic modules. In some cases, a module can emulate an electronic module via a virtual machine, e.g., within a cloud computing system. As will be understood by one of ordinary skill in the art, each embodiment disclosed herein can comprise, consist essentially of or consist of its particular stated element, step, ingredient or module. Thus, the terms “include” or “including” should be interpreted to recite: “comprise, consist of, or consist essentially of.” The transition term “comprise” or “comprises” means has, but is not limited to, and allows for the inclusion of unspecified elements, steps, ingredients, or modules, even in major amounts. The transitional phrase “consisting of’ excludes any element, step, ingredient or module not specified. The transition phrase “consisting essentially of’ limits the scope of the embodiment to the specified elements, steps, ingredients or modules and to those that do not materially affect the embodiment. Moreover, the word “exemplary” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, use of the word exemplary is intended to present concepts in a concrete fashion. As used in this application, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless specified otherwise, or clear from context, “X employs A or B” is intended to mean any of the natural inclusive permutations. That is, if X employs A; X employs B; or X employs both A and B, then “X employs A or B” is satisfied under any of the foregoing instances. In addition, the articles “a” and “an” as used in this application and the appended claims should generally be construed to mean “one or more” unless specified otherwise or clear from context to be directed to a singular form. Groupings of alternative elements or embodiments of the invention disclosed herein are not to be construed as limitations. Each group member may be referred to and claimed individually or in any combination with other members of the group or other elements found herein. It is anticipated that one or more members of a group may be included in, or deleted from, a group for reasons of convenience and / or patentability. When any such inclusion or deletion occurs, the specification is deemed to contain the group as modified thus fulfilling the written description of all Markush groups used in the appended claims. It will be appreciated from the discussion above that the embodiments shown in the Figures are merely exemplary, and include features which may be generalised, removed or replaced as described herein and as set out in the claims. In the context of the present 5 disclosure other examples and variations of the apparatus and methods described herein will be apparent to a person of skill in the art.
Claims
1. A computer-implemented method of creating community groups for clinical trials and community engagement, each community group comprising a plurality of patients grouped based on a common attribute, the method comprising:obtaining at least one source of information relating to patients;in the event that the at least one source of information is in an unstructured format, converting the information to a structured format;applying a machine learning model to the information to determine one or more attributes of the patients from the information;grouping the patients into community groups based on the determined attributes.
2. The computer-implemented method of claim 1 wherein applying a machine learning model to the information to determine attributes of the patients from the information comprises applying at least one of (i) a large language model, LLM, and (ii) a natural language processing, NLP, model, to the information.
3. The computer-implemented method of claim 2 wherein applying a machine learning model to the information to determine attributes of the patients from the information comprises determining health conditions of the patients.
4. The computer-implemented method of any of the previous claims wherein applying a machine learning model to the information to determine attributes of the patients from the information comprises identifying patients, or a subset of patients, from the information.
5. The computer-implemented method of claim 4 further comprising associating each of the identified patients with attributes.
6. The computer-implemented method of any of the previous claims wherein the attributes comprise at least one of: symptoms, demographics, therapeutic areas, non-clinical areas, minimum group size, maximum group size, category.
7. The computer-implemented method of any of the previous claims wherein the obtained information relating to patients comprises any of: freetext, website information,description of community group, marketing collateral, surveys, clinical trial master file, electronic clinical trial master file, documentation relating to a clinical trial.
8. The computer-implemented method of any of the previous claims wherein grouping the patients into community groups based on the determined attributes comprises grouping each patient into a plurality of community groups.
9. The computer-implemented method of any of the previous claims further comprising creating a searchable database of community groups based on the grouping of patients into community groups, wherein the searchable database of community groups is searchable based on determined attributes.
10. The computer-implemented method of claim 9 further comprising ascribing at least one feature to each community group in the searchable database based on the attribute or attributes in common for that community group.
11. The computer-implemented method of claim 10, further comprising: parsing the database of community groups to check for duplicates.
12. The computer-implemented method of claim 11, wherein parsing the database of community groups to check for duplicates comprises:applying a machine learning model and process automation to the database of community groups to determine community groups in the database that have at least one feature in common; andin the event that features are found to be common between at least two community groups, providing an indication that at least one of the at least two community groups is a duplicate.
13. The computer-implemented method of claim 12 wherein features are found to be in common when at least one of: (i) the website links have the same domain, (ii) email addresses have the same domain, (iii) the name of the community group is identical, (iv) the size of the community group is identical.
14. A computer-implemented method of determining duplicates in a database ofcommunity groups for a clinical trial, the method comprising:obtaining a database of community groups, wherein each community group comprises a plurality of patients having at least one attribute in common, and wherein each community group has at least one feature;parsing the database of community groups to determine community groups that have at least one feature in common;and in the event that at least one feature is found in common between at least two community groups, providing an indication that at least one of the at least two community groups is a duplicate.
15. The computer-implemented method of claim 13 or 14 wherein determining community groups that have at least one feature in common comprises at least one of:(i) determining community groups that have at least one identical feature in common; and(ii) determining community groups that have at least one similar feature.
16. The computer-implemented method of claim 13, 14 or 15 wherein determining community groups that have at least one feature in common comprises determining community groups that have at least one identical feature in common, and at least one other feature that is similar in common.
17. A system for creating and managing community groups for clinical trials and community engagement, the system comprising:a web application module;a community group manager module;a duplicate manager module;a cloud-based database; anda cloud-based search module configured to search the cloud-based database; wherein:the web-application module is configured to interact with the community group manager module and provide a user interface;the community group manager module is configured to communicate with the cloud-based database to enable community groups to be added, edited or deleted from the cloud-based database interface, and to communicate with the duplicate managermodule;wherein the duplicate manager module is configured to check for duplicates of community groups in the cloud-based database; andwherein the cloud-based search module is configured to enable searches of 5 community groups in the cloud-based database.
18. The system of claim 17, wherein the cloud-based database is configured to store community groups in the database along with at least one feature and at least one attribute, and wherein the could-based search module is configured to enable searches of the cloud-10 based database based on at least one of: (i) features, and (ii) attributes.
19. A computer readable non-transitory storage medium comprising a program for a computer configured to cause a processor to perform the method of any of claims 1 to 16.15