Method for recommending a qualified health study

The method uses machine learning algorithms to extract and recommend qualified health studies from scattered and poorly referenced sources, addressing the challenge of identifying and validating studies for reuse in new health studies, and improving the efficiency of study design.

WO2025132922A1PCT designated stage expired Publication Date: 2025-06-26SKEZI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/087603
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-12-19
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing health studies are scattered, poorly referenced, and often cited in scientific studies, making it difficult to identify and validate qualified studies for reuse in designing new health studies.

Method used

A computer-implemented method that uses machine learning algorithms to extract and recommend qualified health studies by processing free text data, extracting relevant articles from databases, classifying articles containing health studies, and generating a recommendation index based on criteria such as date, frequency of use, and similarity index.

Benefits of technology

The method effectively identifies and prioritizes qualified health studies, reducing the time and effort required to design new studies by providing a structured and organized database of reusable health studies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024087603_26062025_PF_FP_ABST
    Figure EP2024087603_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for recommending a qualified health study, comprising:  receiving (REC1) a first set of data (ENS1) encoding a free text describing the objective of a study characterizing the health status of a set of patients;  extracting (EXT1), from at least one database (BDa), a first list (LIST1) of articles (ARTi);  extracting (EXT2) a second list of health studies (LIST2) based on the execution of a second function (F2) so as to classify the articles (ARTi) in the first list (LIST1);  comparing (COMP1) each health study (PROi) in the extracted second list (LIST2) with a set of studies (PROk) present within a first memory in order to generate an availability label (D0);  generating (GEN1) a fourth list (LIST4) according to a recommendation index (INDR), said recommendation index (INDR) being computed based on a date criterion (C1).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PROCESS FOR RECOMMENDING A QUALIFIED HEALTH STUDY

[0002] Field of invention

[0003] The field of the invention relates to that of computer-implemented methods for generating an ordered list of documents corresponding to health studies with the aim of recommending at least one of them to a user. The field of the invention relates to solutions for extracting, identifying and exploiting data in natural languages ​​to generate recommendations for exploiting health studies.

[0004] State of the art

[0005] Currently, there is a need to create health studies that allow the monitoring of patients with a given pathology or disease. More and more health studies are based on questionnaires to be sent to patients. Some of these questionnaires have a particularly validated metrology and can be used as tools to be implemented within other larger questionnaires. These studies are difficult to carry out because it is necessary to establish a study with patients over a given period and require a certain homogeneity in the types of responses to the questions asked, and therefore satisfy a metrology, although the responses themselves may differ. In other words, the design of a monitoring form containing a set of questions is a particularly sensitive and difficult exercise to implement.Indeed, the questions must be unambiguous, cover all possible cases and be organized and structured in an optimal way to collect the maximum amount of useful information that can be used for a given patient. Furthermore, the outline of a study with regard to a pathology or disease of a group of patients is not always easy to establish because it is necessary to have a sufficient, heterogeneous and statistically representative panel of a phenomenon.

[0006] These tracking forms therefore involve a great deal of design work and often require validation by a health organization or other body. These studies are then generally referenced and stored in a database. One problem is that these studies are associated with medical data that cannot be made accessible without data processing. References to these studies and other studies not referenced in such a database are most often found in all available resources. These resources can be, for example, scientific articles. Indeed, these studies are generally scattered, poorly referenced, and often cited in scientific studies in support of a scientific publication.

[0007] However, there are many resources of scientific documents such as articles, theses, scientific abstracts, patent documents, etc. that sometimes mention health studies that can be more or less serious. The seriousness of a study can depend on the medical body that carried out the study, the number of patients studied, their representativeness of a pathology or not, the methodologies used, the conclusions of the study and many other criteria. It is therefore important that the studies identified can be validated or certified by a third-party organization.

[0008] There is a need to define a solution to identify existing studies that can be used or reused in the design of a new health study.

[0009] Summary of the invention

[0010] Process for recommending a qualified health study including:

[0011] ■ Receipt of a first set of data encoding free text describing the objective of a study characterizing the health status of a group of patients;

[0012] ■ Extracting a first list of articles from at least one database, from the execution of a first function implementing a first machine learning algorithm trained to generate a set of semantic similarity indices between the first data set and each identified article from said at least one database, said first list of articles corresponding to a set of articles having a similarity index greater than a predefined threshold, said first data set defining an input of the first function;

[0013] ■ Extraction of a second list of health studies from the execution of a second function implementing a second machine learning algorithm trained to classify the articles of the first list, a first class corresponding to the articles comprising first data describing at least one health study, a health study characterizing the health status of a cohort of patients and being called “health study”, said first data being automatically extracted to define a second list of health studies;

[0014] ■ Comparison of the first data of each health study from the second extracted list with a set of studies present within a first memory of a data server, the result of said comparison making it possible to generate an availability label when a study from the second list is present in said first memory;

[0015] ■ Extraction of a third list of health studies with an availability label;

[0016] ■ Generation of a fourth list of health studies according to a recommendation index, said recommendation index being calculated from a first criterion of date of each study in the third list, a second criterion of frequency of use of each study in the third list and the similarity index of each study in the third list.

[0017] According to one embodiment, the first machine learning algorithm comprises a software component for automatically generating a translation of the first data set into a given language. An advantage is that it makes it possible to obtain a greater number of references of interest.

[0018] According to one embodiment, the semantic similarity index is calculated from a definition of a semantic distance.

[0019] According to one embodiment, the calculation of the semantic distance between two data sets is implemented by the first machine learning algorithm, two input vectors being constructed from the first data set and the data characterizing an article, said semantic distance making it possible to calculate a distance between said two vectors encoding data specific to two distinct data sets.

[0020] According to one embodiment, the first machine learning algorithm comprises the implementation of a generative autoregressive language model of the Transformers type trained on the one hand from a set of data describing in natural language a health study and on the other hand from data quantifying the state of health of a given population, said data being organized according to different questions and answers.

[0021] According to one embodiment, the second function implements a syntactic analysis comprising the execution of the second machine learning algorithm trained to classify a set of articles, a class of articles being generated for which a citation to a health study has been identified in said articles of the class, said second machine learning algorithm comprising the identification and extraction of named entities characterizing a citation to a health study. An advantage is to make it possible to select health studies from a set of articles being preselected in the field of interest of a user. Thus, the invention makes it possible to reduce calculation times and to obtain more relevant results.

[0022] According to one embodiment, the second machine learning algorithm comprises calculating a confidence score associated with each article classified as the first class.

[0023] According to one embodiment, the second machine learning algorithm is of the convolutional neural network type.

[0024] According to one embodiment, the data extracted from the articles corresponding to the health studies include metadata comprising titles, identifiers, author names, themes, author names.

[0025] According to one embodiment, the second machine learning algorithm comprises a second class comprising articles not comprising health studies, a selection of said articles of the second class being made accessible to a user by means of a display. An advantage is to make it possible to provide articles of interest to a user even when they do not comprise a health study.

[0026] According to one embodiment, the method comprises acquiring a classification indicator making it possible to reinforce the training of the first algorithm and / or the second algorithm, said classification indicator making it possible to indicate that a classified health study does not correspond to the first data set or that an article does not include a health study while said article has been classified in the first class.

[0027] One advantage is to strengthen machine learning by taking into account user feedback on the relevance of a health study that has been classified according to the method of the invention.

[0028] According to one embodiment, the method comprises:

[0029] ■ Extraction of health studies from each article in the fourth list defining new entries from a first database;

[0030] ■ Adding the said new entries to the database.

[0031] One advantage is maintaining an up-to-date database of health studies that can be reused by other services.

[0032] According to one embodiment, a study database is updated, said update comprising:

[0033] ■ Recording the fourth list within said studies database; said recording comprising a first recording of all the articles in the fourth list defining a first entry in the database and a second recording of all the health studies of each article in the fourth list defining a second entry in the first database and recording the associations between the health studies extracted from the articles and the articles themselves.

[0034] One advantage is to organize a knowledge base comprising a set of health studies that can be reused for medical purposes, for example. According to one embodiment, each article comprises a set of discrete symbols encoding data in a natural language.

[0035] According to one embodiment, the recommendation index is calculated from at least one other recommendation criterion including: the publication date, the similarity index, the presence of a right associated with a study, the citation of an article or an author, a citation index of the article. An advantage is to provide data of interest in an organized and ordered manner taking into account a sorting key for the data in order to improve the relevance of the results displayed.

[0036] According to one embodiment, the fourth list is ordered according to the value of the recommendation index. One advantage is to facilitate the choice of a study of interest for a user.

[0037] According to another aspect, the invention relates to a system comprising a server hosting a database and a terminal configured to implement the method of the invention.

[0038] Brief description of the figures

[0039] Other characteristics and advantages of the invention will emerge on reading the detailed description which follows, with reference to the appended figures, which illustrate:

[0040] ■ Figure 1 is a diagram of the main steps of an embodiment of the method of the invention;

[0041] ■ Figure 2 is an example of implementation of the method of the invention when different article databases are used;

[0042] ■ Figure 3 is an example of a system architecture for implementing the method of the invention.

[0043] First ENS1 dataset

[0044] Figure 1 represents an exemplary embodiment of the method of the invention. It comprises a first step of receiving, noted RECi, a first set of data ENSi encoding a free text describing the objective of a study characterizing the state of health of a group of patients. This free text can correspond to a few lines expressed by a language of a doctor or a scientist or any individual expressing which study he seeks to carry out. For example, such a text can be: "I am seeking to carry out a study on the quality of life of women aged between 20 and 40 years suffering from Myasthenia Gravis".

[0045] A health study may advantageously include a questionnaire intended for patients and patient responses associated with the questionnaire questions. A health study is noted PROi in this application. Some of these questionnaires have particularly validated metrology and can be used as tools to be implemented within other questionnaires concerning other health studies or can be used for statistical purposes. Finally, validated metrology also allows more reliable exploitation by computer systems allowing calculations to be carried out or using the data for the purpose of training a machine learning model.

[0046] One of the advantages of PROi validated health studies is the quality of the quantified and qualified information that allows for example to obtain numerous indicators for different operations. For example, the following treatments are cited, but not limited to: quantifications of side effects of a disease in a cohort, statistics on the prevalence of a disease, the specificity of a symptom of a given disease, the sensitivity of a symptom present in a disease, or indicators on remissions, cure or recurrences of a cohort or data characterizing a treatment of a disease.

[0047] Finally, PROi health studies allow us to obtain numerous indicators of a disease during all the phases in which it is addressed, from its detection to the different stages of treatment.

[0048] According to one embodiment, the first ENSi data set comprises a predefined number of sequences of discrete symbols in a natural language, each sequence defining a word, the predefined number being between 10 and 50 sequences.

[0049] According to one embodiment, the predefined number being between 10 and 400 sequences.

[0050] The method of the invention is capable of processing a text in any language and with any grammar. Indeed, the invention implements a pre-trained machine learning algorithm making it possible to transpose a first matrix of statistical scores of a first set of words in a given language to a second matrix of statistical scores of a second set of words in another given language, the second set of words corresponding to an automatic translation or not of the first set of words.

[0051] In one example, free text is translated into a given language using a software component that can generate a machine translation. Such a translation can be achieved using a dictionary, a thesaurus, and a machine learning algorithm.

[0052] The first ENSi dataset may also include queries reformulated from a "Transformer" type algorithm for regenerating text in the form of a question or query from natural language text. A model of such a machine learning algorithm may be of the large language model type, autoregressive model, generative model or pre-trained or pre-trained generative transformer type.

[0053] The first ENSi dataset can also encode a list of words according to their index or reference in a dictionary.

[0054] First function F1 for identifying similar items

[0055] According to a first example, the first dataset ENSi defines an input of a function denoted Fi implementing a first machine learning algorithm MLAi trained to calculate a similarity score between the first dataset ENSi and each article of an article database, for example denoted BDa. According to an example, different article databases are consulted BDa, BDb, BDc.

[0056] Such a score can also be called a proximity score between an article on the one hand and the first ENSi data set on the other hand.

[0057] The execution of the first function Fi makes it possible to extract ARTi articles from at least one database BDa to generate a first list LISTi whose similarity score is greater than a given threshold. The list of articles LISTi can be ordered so as to index the articles according to their similarity score. The threshold can be fixed or considered so as to retain a given number of ARTi articles similar to the set ENSi. The method of the invention therefore comprises the extraction, denoted EXTi in figure 1, of a set of ARTi articles or more generally of digital documents corresponding to scientific articles, abstracts, patent documents, theses and any other document accessible within a first database of articles BDa.

[0058] Each item contains a set of discrete symbols encoding data in a natural language that can be analyzed by a machine learning algorithm, denoted MLAi.

[0059] To this end, a uniform resource locator, commonly referred to as URL, makes it possible to perform a search in at least one remote BDa database. The method of the invention can allow the configuration of a plurality of URLi in the case where a search for articles in several databases is configured.

[0060] The first ENSi dataset may also include concepts extracted from the received text in natural language. Concept extraction can be performed, for example, using a semantic graph and calculating the proximity of concepts in the graph by defining a distance in the graph. Such a distance is generally used to calculate the distance between different nodes in the semantic graph and therefore their proximity due to the fact that a given number of edges in the graph connects two concepts in the graph together.

[0061] Article databases

[0062] According to one embodiment, the method of the invention comprises a configuration for querying a plurality of article bases BDa, BDb, BDc. Figure 2 represents an example in which a plurality of article bases are queried in order to process a large number of scientific articles from different sources and to extract a maximum of PROi health studies.

[0063] It is understood that this phase can be carried out independently of the reception of a new query comprising a first set of ENSi data. Thus, a database BD2, described below, grouping the references to the extracted PROi health studies and the associated ARTi articles can be created and this database BD2 can then be updated when the article bases are enriched over time with new articles. However, during a new query and the reception of a new ENSi set, an update of the extraction of new articles comprising data characterizing health studies can be carried out so that the enrichment of the database BD2 is continuously carried out. According to one embodiment, the first algorithm MLA1 is executed to identify a set of articles regardless of whether or not they may contain data describing studies.Thus, in this case, the first MLA1 algorithm makes it possible to identify articles close to the first ENS1 set by implementing a predefined semantic distance. In a second step, the articles containing data characterizing health studies will be selected from a second MLA2 algorithm applied to all the selected articles.

[0064] The EXT1 extraction can be carried out on only part of the BDa database, or any other BDb or BDc database based on a date criterion in order to select the most recent articles and apply the MLA1 algorithm or the F1 function to these new articles from the BDa database or another one.

[0065] First machine learning model

[0066] A pre-trained MLA1 machine learning model is then implemented from the ENS1 input data organized in the form of an input vector. According to one embodiment, the MLA1 machine learning model is configured to calculate a similarity score between each ARTi article in the BDa database and the first data set.

[0067] According to another example, an algorithm implementing a technique other than a machine learning model can be implemented to calculate the similarity according to a predefined distance between the first set ENS1 and an article. According to different examples, a function of the Cosine similarity type, the Jaccard distance, or even the N-gram technique.

[0068] In order to calculate an IND similarity score s, the method of the invention makes it possible, for example, to use a semantic distance based on a vector product. A semantic distance can then be defined to calculate a distance between two vectors encoding data specific to two distinct data sets, such as the data of the first set ENS1 and all or part of the data of an ARTi article. According to one embodiment, the MLAi machine learning model implements a regression which makes it possible to generate a proximity score between each ARTi article and the first ENSi data set. The articles having a score above a given threshold can be retained and ordered in an order going from closest to furthest from the first ENSi data set. The proximity is defined by a similarity score IND Swhich can be calculated from a regression of the MLAi learning model. The MLAi machine learning model can be configured to define a semantic distance by training said model.

[0069] According to an exemplary embodiment, before processing each ARTi article, a language detection of the ARTi article is performed. When an identified language does not correspond to the language of the data of the first ENSi data set, then a preprocessing aims to automatically translate the first ENSi data set into the language of the article. A new ENSi' data set is then considered to identify the presence of relevant studies in the article considered. In the remainder of the description, the ENSi notation is retained, regardless of the language considered.

[0070] An advantage of the invention is to integrate pre-processing capable of processing natural language and identifying not only PROi-relevant health studies but also their language.

[0071] Second function F2 for study extraction

[0072] When the first LISTi list is generated, a second function F2 can be implemented to extract ARTi articles containing data characterizing scientific studies rated PROi.

[0073] According to one embodiment, a machine learning algorithm MLA2 is configured to classify the articles according to classes. A first class CL1 corresponds to the articles comprising data characterizing PROi health studies. To this end, the algorithm is trained from training data comprising articles. A first set of articles comprises data characterizing health studies and a second set of articles does not comprise data characterizing health studies. Thus, supervised training implementing a cost function makes it possible to train the machine learning algorithm in order to classify the articles defining the input data, for example, of a neural network.A CNN-type neural network, meaning "convolutional neural networks" in English terminology and called a convolutional neural network, can be implemented to implement a classifier and extract data characterizing health studies.

[0074] Data characterizing health studies may include, for example, a reference to a health study, test results, questionnaires and associated responses for a given cohort of patients, study authors, titles, abstracts or summaries, keywords, and conclusions.

[0075] The health study can be named by a given designation, it can in particular be identified thanks to the recognition of terms specific to health trials on a given population. Finally, elements of context, references to patients, to a test protocol, to dates or types of trials can be used to identify the presence of a reference to a health study.

[0076] According to one embodiment, the function F2 implements a syntactic analysis of the articles. This syntactic analysis can be carried out automatically by means of a machine learning algorithm, such as a neural network,

[0077] A neural network can be trained to classify a set of ARTi articles. In one example, a class of CL1 articles is generated for which a citation to a health study has been identified in said articles of the class.

[0078] Other characteristic data may characterize these health studies. We remind you that these studies are rated PROi.

[0079] Alternatively, this classification can be carried out from a function implementing an expert system comprising a knowledge base and rules.

[0080] According to another embodiment which may be complementary, the machine learning model MLA2 is configured to calculate a proximity score for each article or part of the ARTi articles in the BDa database comprising a reference to a health study, noted PROi. According to this example, a single machine learning model grouping the functions F1 and F2 can be implemented. To this end, the model makes it possible to generate a proximity score according to the definition of a semantic distance between the first set ENSi and each ARTi article processed and to classify said articles according to a set of classes making it possible to identify the articles comprising a reference to a PROi health study.

[0081] Named entities

[0082] For example, named entity recognition of natural language text can be performed. Different techniques can be implemented.

[0083] In one example, a dictionary of named entities may be used. In another example, a remote server with a named entity recognition algorithm may be used.

[0084] In this example, access to a dictionary with many words, synonyms, and a vocabulary collection can be configured. The method then makes it possible to check whether a particular named entity is also available in an article. Using a string matching algorithm, a cross-entity check is performed.

[0085] In order to identify words to be checked, a software component may use, on the one hand, identified words from the first data set and words generated from a generative machine learning algorithm making it possible to multiply concepts and keywords close to the first data set.

[0086] According to one embodiment, a digital connector, also called an “application programming interface”, designated by the acronym API and designating in English literature “application programming interface” can be used to extract named entities from a free text in natural language and possibly to generate new ones and to compare them to a resource accessible from a data network.

[0087] Using a named entity recognition algorithm within the second function F2 allows for better performance in detecting the presence of named entities. Indeed, named entities characteristic of a health study can be more easily detected using such an algorithm.

[0088] In a second example, a system based on predefined rules may be implemented to recognize named entities. In one example, a set of pattern-based rules may be used to exploit a morphological pattern or a string of words used in the document. In another example, context-based rules such as contextual rules dependent on the meaning or context of the word in the document may be implemented.

[0089] Finally, in a third example, a machine learning algorithm can be implemented. In this case, statistical modeling is used to detect named entities. A representation based on the document's characteristics is used. One interest is to recognize types of named entities despite slight variations in their spelling.

[0090] In this case, the Fi function ensures the generation of a proximity score and the second function ensures the detection and extraction of named entities ENTi characterizing a citation to a PROi health study.

[0091] By running the first function Fi and the second function F2, we obtain a selection of ARTi articles having at least one reference to a PROi health study and having a similarity score IND S with the first data set ENS1. The set of articles thus selected makes it possible to define a first list of articles LIST1.

[0092] MLA1 Learning

[0093] According to one embodiment, the training of the first machine learning algorithm MLA1 is carried out in such a way that a similarity score IND Sbe generated between on the one hand each ARTi article or each extracted data set characterizing a PROi health study and on the other hand the first ENS1 data set. The learning can aim to configure the quantification of data from a considered article to calculate the proximity or similarity with a smaller ENS1 data set.

[0094] According to one embodiment, the training comprises correct dataset associations between, on the one hand, free texts describing a health study and, on the other hand, ARTi articles or each extracted dataset characterizing a PROi health study. The associations make it possible to train the MLA1 model so that it can train the machine learning model. According to one embodiment, incorrect associations are also made when it is desired to train the network so that it does not associate certain subjects with each other.

[0095] MLA2 Learning

[0096] The method of the invention comprises training the MLA2 machine learning model. This training takes into account a first set of training data comprising several tens or hundreds of article references which are known to contain references to studies.

[0097] When the learning model is configured to identify ARTi articles containing PROi health studies, the learning allows the model to be trained to recognize a study containing a reference and to extract the elements characterizing this PROi health study in the ARTi article in question. The learning also optionally takes into account a second training set of articles not containing a reference to a health study. This learning allows for improved detection of an ARTi article containing a reference to a PROi health study.

[0098] In the case of implementing a classifier, the latter may include several classes allowing to reference or categorize ARTi articles containing a PROi health study. A syntactic and / or semantic analysis allows to identify PROi health studies linked to the first data set.

[0099] Learning can therefore include training with input text types that can describe PROi health studies differently by considering different ways of evoking the problem of a health study. One advantage is to train the model with a wide variety of PROi health study descriptions.

[0100] Furthermore, the semantic field of a health study concerning a given theme can also be taken into account during learning.

[0101] According to one embodiment, the method of the invention comprises training on particular therapeutic fields. This training makes it possible to take into account, for example, the semantics of particular fields such as the field of geriatrics, the field of menopause, the field of oncology, the field of myasthenia, the field of cardiovascular diseases, etc.

[0102] One advantage of this training is that it allows you to quickly score an ARTi article according to a domain and to discard it or retain it depending on its proximity to the input data.

[0103] Expert system

[0104] According to one embodiment, an expert system type algorithm can be used. Such an expert system type algorithm is based on the development of rules, such as business rules and a knowledge base. One advantage is to inject knowledge into the data processing to exploit a corpus of documents when they have a given, well-identified structure or when certain elements can be identified in a recurring manner within said articles. For example, when specific numbers are assigned to PROi health study types, such a number and / or its type can be searched directly within the articles.

[0105] In a first case, an expert system type algorithm can be used to replace the MLA2 machine learning algorithm. In this case, the expert system is configured to first identify ARTi articles containing at least one reference to a PROi health study, and extract the information characterizing these PROi health studies. A dictionary and a thesaurus can be used, for example, to generate a proximity score.

[0106] According to one embodiment, a second semantic analysis can be carried out more specifically on the data extracted from the article corresponding to at least one PROi health study with the objective of evaluating the proximity between, on the one hand, the relative extracted data characterizing the PROi health studies which are more specific than the set of data from the ARTi article and, on the other hand, the first set of data ENS1. A second similarity score IND S2 can then be generated.

[0107] In a second case, such an expert system type algorithm can be used in combination with the MLA2 machine learning algorithm. For example, the identification of data characterizing a health study or a reference to a health study is carried out from an expert system through the use of rules and a knowledge base. When the ARTi articles include a reference to a PROi health study or include data characterizing at least one PROi health study, an MLAi machine learning algorithm is configured to generate a proximity score between, on the one hand, the data characterizing the PROi health study or those of the ARTi article and, on the other hand, the first ENSi data set.

[0108] Extraction of articles and studies

[0109] The method of the invention comprises an EXT2 extraction of ARTi articles from the BDa database which comprise:

[0110] ■ on the one hand a similarity score among the highest with the first set of ENS1 data and;

[0111] ■ on the other hand a citation, a reference to a PROi health study or the data describing a PROi health study.

[0112] The extraction of articles containing health studies can be implemented by two successive extraction operations, including the first extraction EXT1 and the second extraction EXT2.

[0113] The data characterizing the PROi health studies are in turn extracted from the ARTi articles and the method of the invention makes it possible to generate a list, noted LIST2, possibly ordered according to the IND similarity score. Sof PROi health studies. For this purpose, a designation of each PROi health study is generated from the data characterizing said PROi health studies.

[0114] The function F1 to generate the list LIST1 executes the trained machine learning model MLA1, then the function F2 to generate the list LIST2 executes the trained machine learning model MLA2. When the ARTi articles in the list LIST2 are identified, the data characterizing the PROi health studies of the ARTi articles having the strongest proximity to the first dataset ENS1 are extracted and stored in a memory.

[0115] Descriptors

[0116] The EXT2 extraction step comprises the identification and extraction of all the descriptors qualifying or describing a PROi health study within each processed ARTi article. The descriptors may relate to a title, a designation, an identifier, a reference, a subtitle, an abstract, a note or a comment. According to other embodiments, the descriptors may comprise an extract from the study, a conclusion of the study, one or more authors of the study. A descriptor may also be the language of the PROi health study or that of the ARTi article or both. The descriptors according to the articles may not be homogeneous. One advantage of the invention is to standardize and homogenize the extracted data forming descriptors of the study.

[0117] All extracted data including the descriptors of each identified PROi health study are possibly sent via a data network to a remote server for processing of this data, in particular their recording in a new BD2 database.

[0118] According to one embodiment, a descriptor corresponds to at least one piece of data characterizing at least one pathology, or at least one sign of a disease or at least one psychological or physiological state of a cohort of individuals, called a sign hereinafter.

[0119] For example, a sign within the scope of the present invention may be a quantification of a level of depression, a quantification of a level of anxiety, a quantification of a level of quality of life. According to other examples, it may be a quantification of rehabilitation, a quantification of re-education of a function of the human body, a quantification of functional recovery, a quantification of a level of healing or remission. According to another example, it may be a quantification of side effects of a medical or therapeutic treatment.

[0120] According to one embodiment, a first data item characterizes a sign. This first data item corresponds to a description of the sign, a wording, keywords, symptoms or diseases and / or a description of risk factors, etc. One or more of these attributes can be used according to different configurations of the invention.

[0121] According to one embodiment, a second piece of data characterizes a scale for normalizing the quantification of the sign within a cohort. The scales may correspond to a score from 1 to 10 or a score from 1 to 5 or any other reference. The scale may correspond to a quantification of the type "weakly", "moderately", "significantly" or "abundant". According to one embodiment, a third piece of data characterizes an overall value of the presence of the sign within a cohort questioned in a health study.

[0122] One advantage of the first, second and third data characterizing a sign within a cohort is that it allows studies to be classified according to finer criteria in order to identify appropriate health studies.

[0123] For example, a query might be to obtain health studies representing a higher-than-normal level of anxiety in a population of adolescents in Europe. Thus, the extracted health studies can be classified according to different criteria including the publication date, the similarity between the query and the content of an article and according to a quantification of a sign. In this example, we are interested in the sign "anxiety" with data characterizing a level "higher than normal".

[0124] The method of the invention saves valuable time for the medical profession or biostatistician and makes it possible to obtain qualified health studies that can be reused in a new study framework.

[0125] New BD2 database

[0126] According to one embodiment, the method comprises creating and updating a database BD2 of PROi health studies extracted from a plurality of ARTi articles, possibly a large number of articles, in order to enrich the data forming descriptors of these PROi health studies. This new database is denoted BD2. It comprises all the PROi health studies extracted from the different articles processed by the function F2. The database BD2 records the link or association between each PROi health study and each ARTi article from which it was extracted.

[0127] Thus, the method of the invention allows the creation of a BD2 database for storing the information extracted from the ARTi articles, including the identifiers of the PROi health studies, the associated data and the original sources of the studies such as the authors, the context of the study, the description of the cohort, etc. An advantage of the invention is to structure this BD2 database in such a way as to allow a subsequent efficient search and a rapid retrieval of the information when a query is carried out against a BD2 database comprising ARTi articles and PROi health studies.

[0128] The creation of the BD2 database includes the creation of specific data sets describing health studies. These groups can take the form of columns or rows or tables or files of a database. The BD2 database includes association keys with other data including the name of the study, the publication date of the health study, characteristics of the cohort such as the number of patients interviewed or surveyed. In addition, the data describing health studies are associated with the reference of at least one scientific article if applicable, indicators, keywords. In addition, this database can be updated each time a health study is used, for example each time it is downloaded so as to quantify the use of a health study and qualify its relevance. In addition, the number of citations of this PRO1 health study can be updated regularly in the created BD2 database.For this purpose, a query can be configured in the NET1 data network to check for new references or citations of a given health study occurring in a given period of time. Finally, the created BD2 database can save the names of the authors and laboratories, the language of the article. According to one embodiment, translations into different languages ​​of the abstracts, summaries or titles of the health studies are saved in the BD2 database.

[0129] Figure 2 represents a scenario in which the BD2 database is updated following reprocessing of new articles extracted from one or more BDa, BDb, BDc article databases. This updated BD2 study database includes data describing an association between each PROi health study and each ARTi article extracted from one of these BDa, BDb, BDc databases.

[0130] Comparison

[0131] The method comprises a step aimed at comparing COMP1 each PROi health study reference or each PROi health study designation or each PROi health study identification with a set of reference health studies {PROk} stored in a remote memory such as a database BD1 accessible from a remote server. This latter database BDi is called in the present description reference database and is denoted BDi. The comparison step COMPi aims to generate an availability label Do when a study from the list LIST2 is stored in the database BD1. When a PROi health study is present in the reference database BD1, a label Do makes it possible to inform that the given PROi health study is indeed valid and can be used.

[0132] For example, different levels of availability are managed in the BD1 reference database. For example, a first level of availability is that a study is available because it has been validated in a country other than France, a second level of availability is that a study is available because it has been validated in France. Other parameters can be taken into account to generate a Do availability label.

[0133] The first database BD1 allows for the organized storage of PROk health studies that can be categorized in different ways. The notation "PROk" is used to denote reference health studies. For example, a PROk health study may include an identifier, a title, an author, a short description, a date, or a list of keywords. A PROk health study in the reference database BD1 includes a set of data quantifying the health status of a given population, said data being organized according to different questions and answers. In one example, the PROk health study may include study results, conclusions, or references.

[0134] The COMPi comparison step can be performed so that each data extracted from an ARTi article qualifying or describing the PROi health study can be used to result in an identification or correspondence of the present PROi health study in the BD1 reference database.

[0135] An advantage of the method of the invention is to establish a correspondence that is as reliable as possible between a study in the BD2 database and a study in the BD1 reference database. This correspondence makes it possible to check whether a health study identified in the ARTi scientific articles has already been referenced in another BD1 database. When such a correspondence is identified, a component makes it possible to manage the correspondences and to update the data in the BD2 database with the information extracted from the BDi reference database. Thus, the BDi database makes it possible to aggregate data relating to ARTi articles and PROi health studies and thus allows better accessibility to health information.

[0136] According to one embodiment, the method of the invention comprises a duplicate management component. According to one embodiment, when the same study is referenced in different ways in different ARTi scientific articles, a processing aimed at associating lines with each other from the second database BD2 is implemented. One interest is to enrich the descriptors of each PROi health study and to associate them with different scientific articles.

[0137] Extracting a new list

[0138] The method of the invention therefore comprises a step EXT3 of extracting a subset of PROi health studies and / or associated ARTi articles having an availability label Do. Preferably, the third list comprises a reference to a PROi health study and a reference to the article from which it was extracted.

[0139] The method of the invention makes it possible to generate a third list LIST3 when several PROi health studies are labeled with the availability indicator following the comparison step COMP1.

[0140] Recommendation

[0141] When different PROi health studies are present in the third list LIST3, a recommendation algorithm is configured to calculate an INDR recommendation index of each PROi health study in the third list LIST3.

[0142] The recommendation algorithm is configured to take into account different criteria. According to a first embodiment, the recommendation algorithm takes into account a first criterion Ci of publication date of each PROi health study to calculate an INDR recommendation index. The more recent the date, the higher the INDR recommendation index is likely to be.

[0143] According to one embodiment, the recommendation algorithm takes into account a second criterion C2 which corresponds to the frequency of use of a given PROi health study. The more a study is used, the more its reliability is approved, thus the higher the frequency of use of a study, the higher the recommendation index INDR. In this embodiment, the reference database BDi of health studies comprises a usage indicator for each study which is updated with each new use of a PROi health study. The criteria Ci and C2 can also be designated as recommendation criteria.

[0144] According to another embodiment, the recommendation algorithm takes into account other criteria, also called recommendation criteria. For example, a third criterion corresponds to a number of citations of the ARTi article or the PROi health study or an author of the article, or an index calculated from the number of citations of the ARTi article. According to another example, the recommendation algorithm takes into account a fourth criterion corresponding to the presence of a right associated with the ARTi article or the PROi health study. According to another example, the recommendation algorithm takes into account a fifth criterion corresponding to the language of the PROi health study or that of the ARTi article.

[0145] According to an exemplary embodiment, the recommendation algorithm takes into account the similarity index IND S previously calculated using the machine learning algorithm.

[0146] The method of the invention comprises a step GEN1 aimed at generating a fourth list LIST4 of health studies PROi ordered according to the recommendation index INDR. This list can also be noted LIST4(PROi) in Figure 1. According to an example, this recommendation index INDR corresponds to a weighted value of all the criteria taken into account. Figure 1 represents the notation INDR(CI, C2, INDs) illustrating that according to this embodiment, the recommendation index INDR can depend on the first criterion Ci, the second criterion C2 and the similarity index INDs.

[0147] System

[0148] Figure 3 represents an example of architecture of a system of the invention comprising at least one server SERV2 hosting a database BD2 made accessible remotely via an administration console CON1 and a data network denoted NET1 and which may for example be the internet network. This administration console CON1 makes it possible to edit the data of the study database BD2, process duplicates, enrich the fields, manage user rights. A terminal Ti is represented and can be used to define a new query describing the first set of data ENSi. It can correspond to a PC type computer, a smartphone, otherwise called a smart phone, or a digital tablet. The terminal Ti allows a user to check the existence of existing health study(s) on a given medical field close to the one that the user would like to carry out.Thanks to the BD2 database, which is built and updated regularly, and the BD1 reference database, it can check the availability of a study validated by an organization that has medically verified this study.

[0149] Figure 3 also represents a database BDa of articles, such as a database of scientific articles accessible from a user terminal T1.

[0150] The system is configured to implement the method of the invention. The system includes in particular a calculator implemented in each server to carry out the operations of each algorithm, in particular those implemented by the functions F1 and F2. In addition, a calculator can be used to carry out the operations of comparing the health studies with the reference studies. Furthermore, memories can be implemented to store and record extracted or calculated data.

[0151] It will be appreciated that the various aspects of the invention described here provide a concrete and specific technical solution to a technical problem, namely: identifying in a data network comprising heterogeneous resources existing health studies which are qualified, i.e. corresponding to at least one clinical case, for their reuse or their adaptations in other clinical cases. A heterogeneous database may comprise numerous indexed or non-indexed resources such as scientific article databases, health study databases, etc.

[0152] It should be noted that the solution aimed at carrying out a search directly in a memory already containing qualified studies does not allow for the precise extraction of reusable studies for an updated need to carry out a new similar study. Indeed, it is difficult to search directly in these studies which are poorly referenced, poorly documented and sometimes cannot be interpreted by a machine such as a computer. Indeed, many studies are documents which are not easily readable by a computer because the formats are image formats or other formats which cannot be interpreted by a computer.Furthermore, almost no metadata is present in these health / clinical study databases, in particular such as information on clinical cases, the objectives of the study, the perspective of the latter in relation to the context in which it was carried out. This memory containing qualified health studies generally only contains data characterizing certain data from the study. However, in order to reuse a health study, it is necessary to obtain medical and scientific data characterizing the context of the study such as the year, the characterization of the cohort, the context of the study, etc.

[0153] No document discloses a method exploiting both a scientific article database and a qualified clinical health study database to construct a new clinical health study database.

[0154] It will also be appreciated that the disclosed method is not a mere abstract idea; it is implemented through a tangible process involving specific steps and components. These include the automatic identification and selection of documents from predefined sources, processing of machine learning algorithm configuration to generate similarity factors between documents, to perform comparisons or to calculate semantic distances to identify documents to be processed. It further involves the automatic extraction of particular data in the selected documents defining health studies in order to prioritize and index them according to input data defining a query.Finally, an operation of generating a specific database recorded on one or more memories makes it possible to produce a source of exploitable data to use health studies in order to generate new ones for close, related or similar updated contexts.

[0155] In one implementation, the method uses a system architecture comprising multiple servers (SERVI, SERV2, etc.) that are explicitly configured to perform distinct computational and evaluation functions, demonstrating a concrete and specific technological framework for achieving the desired results. In addition, equipment such as computers for configuring administration consoles (CON1). Various aspects of the invention provide a clear improvement in the technical field of medical data exploitation for the purpose of developing appropriate clinical health studies.

[0156] For example :

[0157] The method automates the generation of a database and its updating for exploitation in order to produce new organized data. This allows for easier exploitation of health studies in different contexts, at different times, in different geographical areas.

[0158] By using a first machine learning model to identify and select documents from different data resources to implement a second machine learning model to extract data that better characterizes the studies in which similarities are sought, certain aspects of the invention significantly reduce the instances of errors, improve the completeness of health data identification, and improve the reliability of new studies to be developed.

[0159] The aspects of the invention are not abstract but closely related to the technological implementation involving specific hardware and software integrations. In one implementation, the system uses interconnected data servers for data identification, selection, extraction, processing and storage, combined with computer models operating in separate environments. The use of predefined configurations to generate document lists, extract study lists, evaluate results and prioritize them according to predefined criteria demonstrates a specific and non-generic application of artificial intelligence technology.

[0160] In one or more embodiments, the system includes dedicated computing units, such as GPUs or TPUs, optimized for executing machine learning models. These units may be hosted in a server and execute the MLA1 or MLA2 machine learning models. Each computing unit may apply model weights trained on specific training datasets, generate outputs through a sequence of matrix computations and predictions, and optimize response generation by applying predefined constraints and error control or cost function algorithms during execution.

[0161] The machine learning model execution device may include software modules designed to load pre-trained language models into memory, dynamically configure model prompts, allowing results to be customized to linguistic, cultural, or domain-specific requirements.

[0162] In one or more embodiments, the device for executing the machine learning models may use a scalable cloud-based infrastructure, such as, for example, servers hosted in the cloud dynamically allocate computing resources to execute instances of the computer programs implementing the machine learning models MLA1, MLA2, clusters of virtual machines provide redundancy and scalability to process large volumes of test data, and containerized environments ensure reproducibility and isolation of the different models during execution.

[0163] Expressions such as "comprise," "include," "incorporate," "contain," "is," and "have" should be interpreted in a non-exclusive manner when construing the description and associated claims, i.e., interpreted to allow for the presence of other elements or components that are not explicitly defined. Reference to the singular should also be interpreted as a reference to the plural and vice versa.

[0164] The articles "a" and "an" may be used in connection with various elements and components of the compositions, methods or structures described herein. This is merely for convenience and general meaning of the compositions, methods or structures. Such description includes "one or at least one" of the elements or components. Furthermore, in this document, articles in the singular also include a description of a plurality of elements or components, unless it appears from a specific context that the plural is excluded.

[0165] As used in the specification and in the claims, the expression "at least one", with reference to a list of one or more elements, is to be understood to mean at least one element selected from one or more elements in the list of elements, but not necessarily including at least one of each element specifically listed in the list of elements and not excluding every combination of elements in the list of elements. This definition also allows for the optional presence of elements other than the specifically identified elements in the list of elements to which the expression "at least one" refers, whether or not related to the specifically identified elements.

[0166] The expression "and / or", as used in the specification and in the claims, is to be understood to mean "either or both" of the elements so joined, i.e., elements which are present conjunctively in some cases and disjunctively in other cases. Multiple elements listed with "and / or" are to be interpreted in the same way, i.e., "one or more" of the elements so joined. Other elements may optionally be present in addition to the elements specifically identified by the "and / or" clause, whether or not they are related to the specifically identified elements.

[0167] A person skilled in the art will readily appreciate that various elements, features, and parameters disclosed in the description may be modified and that various disclosed embodiments may be combined without departing from the scope of the invention. For example, various aspects of the present disclosure may be used alone, in combination, or in a variety of arrangements not specifically described in the embodiments described in the foregoing and is therefore not limited in its application to the details and arrangement of components set forth in the foregoing description or illustrated in the drawings. For example, aspects described in one embodiment may be combined in any manner with aspects described in other embodiments.

[0168] Having described above several aspects of at least one embodiment, it is appropriate to appreciate the various alterations, modifications and improvements that persons skilled in the art can readily make to this embodiment. These alterations, modifications and improvements are intended to constitute aspects of the present disclosure. Accordingly, the foregoing description and drawings are given by way of example only.

Claims

CLAIMS 1. Process for recommending a qualified health study comprising: ■ Receipt (RECi) of a first set of data (ENSi) encoding free text describing the objective of a study characterizing the health status of a set of patients; ■ Extraction (EXTi) of a first list (LISTi) of articles (ARTi) from at least one database (BD2, BD a , BDb), from the execution of a first function (F1) implementing a first machine learning algorithm (MLA1) trained to generate a set of semantic similarity indices (INDs) between the first data set (ENS1) and each article (ARTi) identified from said at least one database, said first list of articles (LISTi) corresponding to a set of articles (LISTi) having a similarity index (INDs) greater than a predefined threshold, said first data set (ENS1) defining an input of the first function (F1); ■ Extraction (EXT2) of a second list of health studies (LIST2) from the execution of a second function (F2) implementing a second machine learning algorithm (MLA2) trained to classify the articles (ARTi) of the first list (LISTi), a first class (CL1) corresponding to the articles (ARTi) comprising first data describing at least one health study (PROi), a health study characterizing the health status of a cohort of patients and being called “health study” (PROi), said first data being automatically extracted to define a second list of health studies (LIST2); ■ Comparison (COMP1) of the first data of each health study (PROi) from the second list (LIST2) extracted with a set of studies (PROk) present within a first memory of a data server (SERV1), the result of said comparison (COMP1) making it possible to generate a label of availability (Do) when a study from the second list (LIST2) is present in said first memory; ■ Extraction (EXT3) of a third list of health studies (LIST3) having an availability label (Do); ■ Generation (GEN1) of a fourth list (LIST4) of health studies (PROi) according to a recommendation index (INDR), said recommendation index (INDR) being calculated from a first criterion (Ci) of date of each study (PRO) of the third list (LIST3), a second criterion (C2) of frequency of use of each study (PRO) of the third list (LIST3) and the similarity index (INDs) of each study (PRO) of the third list (LIST3) ■ Extraction of health studies (PRO1) from each article (ARTi) of the fourth list (LIST4) defining new entries of a first database (BD2); ■ Addition of the said new entries in the database (BD2).

2. Method according to claim 1 characterized in that the new database (BD2) comprises for each health study (PRO1) indexed in said database is associated with an updated indicator of use of said health study calculated on the number of downloads of said study.

3. Method according to claim 1 characterized in that the new database (BD2) comprises for each health study (PRO1) indexed in said database (BD2) is associated with data from said database encoding the date of publication of the study, data from said database encoding a reference to at least one scientific article, and / or data encoding a language.

4. Method according to claim 1 characterized in that the first machine learning algorithm (MLA1) comprises a software component making it possible to automatically generate a translation of the first data set (ENS1) into a given language.

5. Method according to any one of the preceding claims, characterized in that the semantic similarity index (INDs) is calculated from a definition of a semantic distance.

6. Method according to the preceding claim, characterized in that the calculation of the semantic distance between two sets of data is implemented by the first machine learning algorithm (MLAi), two input vectors being constructed from the first set of data (ENSi) and the data characterizing an article (ARTi), said semantic distance making it possible to calculate a distance between said two vectors encoding data specific to two distinct sets of data.

7. Method according to any one of the preceding claims, characterized in that the first machine learning algorithm (MLAi) comprises the implementation of a generative autoregressive language model of the Transformers type trained on the one hand from a set of data describing in natural language a health study and on the other hand from data quantifying the state of health of a given population, said data being organized according to different questions and answers.

8. Method according to any one of the preceding claims, characterized in that the second function (F2) implements a syntactic analysis comprising the execution of the second machine learning algorithm (MLA2) trained to classify a set of articles (ARTi), a class of articles (CL1) being generated for which a citation to a health study has been identified in said articles of the class, said second machine learning algorithm (MLA2) comprising the identification and extraction of named entities (ENTi) characterizing a citation to a health study (PROi).

9. Method according to any one of the preceding claims, characterized in that the second machine learning algorithm (MLA2) involves the calculation of a confidence score associated with each article (ARTi) classified in the first class (CL1).

10. Method according to any one of the preceding claims, characterized in that the data extracted from the articles (ARTi) corresponding to the health studies (PROi) comprise metadata comprising titles, identifiers, author names, themes, author names.

11. Method according to any one of the preceding claims, characterized in that the second machine learning algorithm (MLA2) comprises a second class comprising the articles (ARTi) not comprising health studies (PROi), a selection of said articles (ARTi) of the second class being made accessible to a user by means of a display.

12. Method according to any one of claims 8 to 11, characterized in that it comprises the acquisition of a classification indicator making it possible to reinforce the training of the first algorithm (MLA1) and / or the second algorithm (MLA2), said classification indicator making it possible to indicate that a classified health study (PROi) does not correspond to the first data set (ENS1) or that an article does not include a health study (PROi) while said article has been classified in the first class (CLi).

13. Method according to any one of the preceding claims, characterized in that it comprises: ■ Extraction of health studies (PROi) from each article (ARTi) of the fourth list (LIST4) defining new entries of a first database (BD2); ■ Addition of the said new entries in the database (BD2).

14. Method according to claim 13 characterized in that a study database (BD2) is updated, said update comprising: ■ Recording of the fourth list (LIST4) within said study database (BD2); said recording comprising a first recording of all the articles (ARTi) of the fourth list (LIST4) defining a first entry of the database (BD2) and a second recording of all the health studies (PROi) of each article (ARTi) of the fourth list (LIST4) defining a second entry of the first database (BD1) and recording of the associations between the health studies (PROi) extracted from the articles and the articles themselves.

15. Method according to any one of the preceding claims, characterized in that the recommendation index (INDR) is calculated from at least one other recommendation criterion among which we find: the publication date, the similarity index, the presence of a right associated with a study, the citation of an article or an author, a citation index of the article.

16. System comprising a server hosting a database (BD2) and a terminal (T1) configured to implement the method of any one of claims 1 to 15.

Citation Information

Patent Citations

  • Systems and methods of clinical trial evaluation

    US20200381087A1

  • Medical Literature Recommender Based on Patient Health Information and User Feedback

    US20210391075A1