PROCESS FOR RECOMMENDING A QUALIFIED HEALTH STUDY

The method employs machine learning algorithms to analyze free text data and recommend qualified health studies, addressing the challenges of scattered and poorly referenced health studies by improving data processing and retrieval efficiency.

FR3156944A1Pending Publication Date: 2025-06-20SKEZI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
FR2023014520
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing health studies are scattered, poorly referenced, and often cited in scientific studies without proper validation, making it difficult to identify and reuse qualified studies for new research.

Method used

A method using machine learning algorithms to extract and recommend qualified health studies by analyzing free text data, generating semantic similarity indices, and classifying articles to identify relevant health studies, while also considering availability and recommendation indices.

Benefits of technology

The method effectively identifies and recommends qualified health studies by improving data processing and retrieval efficiency, enhancing the relevance of results, and maintaining an up-to-date database of reusable health studies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

METHOD FOR RECOMMENDING A QUALIFIED HEALTH STUDY Method for recommending a qualified health study comprising: Reception (REC1) of a first data set (ENS1) encoding a free text describing the objective of a study characterizing the health status of a set of patients; Extraction (EXT1) from at least one database (BDa) of a first list (LIST1) of articles (ARTi); Extraction (EXT2) of a second list of health studies (LIST2) from the execution of a second function (F2) to classify the articles (ARTi) of the first list (LIST1); Comparison (COMP1) of each health study (PROi) of the second list (LIST2) extracted with a set of studies (PROk) present within a first memory making it possible to generate an availability label (D0); Generation (GEN1) of a fourth list (LIST4) according to a recommendation index (INDR), said recommendation index (INDR) being calculated from a date criterion (C1). Figure for the abstract: Fig.1.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: METHOD FOR RECOMMENDING A QUALIFIED HEALTH STUDY Field of invention

[0001] The field of the invention relates to that of computer-implemented methods for generating an ordered list of documents corresponding to health studies with the aim of recommending at least one of them to a user. The field of the invention relates to solutions for extracting, identifying and exploiting data in natural languages ​​to generate recommendations for exploiting health studies. State of the art

[0002] Currently, there is a need to create health studies allowing the monitoring of patients suffering from a given pathology or disease. More and more health studies are based on questionnaires to be sent to patients. Some of these questionnaires have a particularly validated metrology and can be used as tools to be implemented within other larger questionnaires. These studies are difficult to carry out because it is necessary to establish a study with patients over a given period and requiring a certain homogeneity in the types of responses to the questions asked, and therefore satisfying a metrology, although the responses themselves may differ. In other words, the design of a monitoring form comprising a set of questions is a particularly sensitive exercise and difficult to implement.Indeed, the questions must be unambiguous, cover all possible cases and be organized and structured in an optimal way to collect the maximum amount of useful information that can be used for a given patient. Furthermore, the outline of a study with regard to a pathology or disease of a group of patients is not always easy to establish because it is necessary to have a sufficient, heterogeneous and statistically representative panel of a phenomenon.

[0003] These monitoring forms therefore involve a lot of design work and often require validation by a health organization or other. Then, these studies are generally referenced and stored in a database.

[0004] One problem is that these studies are associated with medical data that cannot be made accessible without data processing. References to these studies and other studies not referenced in such a database are most often found in all available resources. These resources can be, for example, scientific articles. Indeed, these studies are generally scattered, poorly referenced and often cited in scientific studies in support of a publication. scientist.

[0005] However, there are many resources of scientific documents such as articles, theses, scientific abstracts, patent documents, etc. which sometimes mention health studies which can be more or less serious. The seriousness of a study can depend on the medical body which carried out the study, the number of patients subject to the study, their representativeness of a pathology or not, the methodologies used, the conclusions of the study and many other criteria. It is therefore important that the studies identified can be validated or certified by a third party organization.

[0006] There is a need to define a solution for identifying existing studies that can be used or reused in the design of a new health study. Summary of the invention

[0007] Method for recommending a qualified health study comprising: • Receipt of a first set of data encoding free text describing the objective of a study characterizing the health status of a group of patients; • Extracting a first list of articles from at least one database, from the execution of a first function implementing a first machine learning algorithm trained to generate a set of semantic similarity indices between the first data set and each identified article from said at least one database, said first list of articles corresponding to a set of articles having a similarity index greater than a predefined threshold, said first data set defining an input of the first function; • Extraction of a second list of health studies from the execution of a second function implementing a second machine learning algorithm trained to classify the articles of the first list, a first class corresponding to the articles comprising first data describing at least one health study, a health study characterizing the health status of a cohort of patients and being called “health study”, said first data being automatically extracted to define a second list of health studies; • Comparison of the first data of each health study from the second extracted list with a set of studies present within a first memory of a data server, the result of said comparison making it possible to generate an availability label when a study from the second list is present in said first memory; • Extraction of a third list of health studies with an availability label; • Generation of a fourth list of health studies according to a recommendation index, said recommendation index being calculated from a first criterion of date of each study in the third list, a second criterion of frequency of use of each study in the third list and the similarity index of each study in the third list.

[0008] According to one embodiment, the first machine learning algorithm comprises a software component for automatically generating a translation of the first data set into a given language. One advantage is that it makes it possible to obtain a greater number of references of interest.

[0009] According to one embodiment, the semantic similarity index is calculated from a definition of a semantic distance.

[0010] According to one embodiment, the calculation of the semantic distance between two sets of data is implemented by the first machine learning algorithm, two input vectors being constructed from the first set of data and the data characterizing an article, said semantic distance making it possible to calculate a distance between said two vectors encoding data specific to two distinct sets of data.

[0011] According to one embodiment, the first machine learning algorithm comprises the implementation of a generative autoregressive language model of the Transformers type trained on the one hand from a set of data describing in natural language a health study and on the other hand from data quantifying the state of health of a given population, said data being organized according to different questions and answers.

[0012] According to one embodiment, the second function implements a syntactic analysis comprising the execution of the second machine learning algorithm trained to classify a set of articles, a class of articles being generated for which a citation to a health study has been identified in said articles of the class, said second machine learning algorithm comprising the identification and extraction of named entities characterizing a citation to a health study. An advantage is to make it possible to select health studies from a set of articles being preselected in the field of interest of a user. Thus, the invention makes it possible to reduce calculation times and to obtain more relevant results.

[0013] According to one embodiment, the second machine learning algorithm comprises calculating a confidence score associated with each classified article of the first class.

[0014] According to one embodiment, the second machine learning algorithm is of the convolutional neural network type.

[0015] According to one embodiment, the data extracted from the articles corresponding to the health studies include metadata comprising titles, identifiers, author names, themes, author names.

[0016] According to one embodiment, the second machine learning algorithm comprises a second class comprising articles not comprising health studies, a selection of said articles of the second class being made accessible to a user by means of a display. An advantage is to make it possible to provide articles of interest to a user even if they do not comprise a health study.

[0017] According to one embodiment, the method comprises the acquisition of a classification indicator making it possible to reinforce the training of the first algorithm and / or the second algorithm, said classification indicator making it possible to indicate that a classified health study does not correspond to the first data set or that an article does not include a health study while said article has been classified in the first class.

[0018] One advantage is to strengthen machine learning by taking into account user feedback on the relevance of a health study that has been classified according to the method of the invention.

[0019] According to one embodiment, the method comprises: • Extraction of health studies from each article in the fourth list defining new entries from a first database; • Adding the said new entries to the database.

[0020] One advantage is maintaining an up-to-date database of health studies that can be reused by other services.

[0021] According to one embodiment, a study database is updated, said update comprising: • Recording of the fourth list within said database of studies; said recording comprising a first recording of all the articles in the fourth list defining a first entry in the database and a second recording of all the health studies of each article in the fourth list defining a second entry in the first database and recording of the associations between the health studies extracted from the articles and the articles themselves.

[0022] One advantage is to organize a knowledge base comprising a set of health studies that can be reused for medical purposes, for example.

[0023] According to one embodiment, each article comprises a set of symbols discrete encoding data in a natural language.

[0024] According to one embodiment, the recommendation index is calculated from at least one other recommendation criterion among which we find: the publication date, the similarity index, the presence of a right associated with a study, the citation of an article or an author, a citation index of the article. An advantage is to provide data of interest in an organized and ordered manner taking into account a sorting key of the data in order to improve the relevance of the results displayed.

[0025] According to one embodiment, the fourth list is ordered according to the value of the recommendation index. One advantage is to facilitate the choice of a study of interest for a user.

[0026] According to another aspect, the invention relates to a system comprising a server hosting a database and a terminal configured to implement the method of the invention. Brief description of the figures

[0027] Other characteristics and advantages of the invention will emerge on reading the detailed description which follows, with reference to the appended figures, which illustrate: • [Fig.l]: a diagram of the main steps of an embodiment of the method of the invention; • [Fig.2]: an example of implementation of the method of the invention when dif various article databases are used; • [Fig.3]: an example of system architecture allowing the implementation of the method of the invention. First ENS1 dataset

[0028] [Fig.l] represents an exemplary embodiment of the method of the invention. It comprises a first step of receiving, denoted RECi, a first set of data ENSi encoding a free text describing the objective of a study characterizing the state of health of a set of patients. This free text may correspond to a few lines expressed by a language of a doctor or a scientist or any individual expressing what study he seeks to carry out. For example, such a text may be: "I am seeking to carry out a study on the quality of life of women aged between 20 and 40 years suffering from Myasthenia Gravis".

[0029] A health study may advantageously comprise a questionnaire intended for patients and patient responses associated with the questions in the questionnaire. A health study is noted PROi in the present application. Some of these questionnaires have a particularly validated metrology and can be used as tools to be implemented within other questionnaires concerning other health studies or can be used for statistical purposes. Finally, a validated metrology allows also more reliable exploitation by computer systems allowing calculations to be carried out or using the data for machine learning model training purposes.

[0030] One of the interests of PRO validated health studies is the quality of the quantified and qualified information which makes it possible to obtain, for example, numerous indicators for different operations. By way of example, the following treatments are cited in a non-exhaustive manner: quantifications of side effects of a disease in a cohort, statistics on the prevalence of a disease, the specificity of a symptom of a given disease, the sensitivity of a symptom present in a disease, or even indicators on remissions, cures or recurrences of a cohort or data characterizing a treatment of a disease.

[0031] Finally, PROi health studies make it possible to obtain numerous indicators of a disease during all the phases during which it is addressed from its detection to the different stages of treatment.

[0032] According to one embodiment, the first ENSi data set comprises a predefined number of sequences of discrete symbols in a natural language, each sequence defining a word, the predefined number being between 10 and 50 sequences.

[0033] According to one embodiment, the predefined number being between 10 and 400 sequences.

[0034] The method of the invention is capable of processing a text in any language and with any grammar. Indeed, the invention implements a pre-trained machine learning algorithm making it possible to transpose a first matrix of statistical scores of a first set of words in a given language to a second matrix of statistical scores of a second set of words in another given language, the second set of words corresponding to an automatic translation or not of the first set of words.

[0035] According to one example, the free text is translated into a given language from a software component allowing to generate an automatic translation. Such a translation can be carried out using a dictionary, a thesaurus and a machine learning algorithm.

[0036] The first ENSi data set may also include queries reformulated from a “Transformer” type algorithm for regenerating a text in the form of a question or a query from a natural language text. A model of such a machine learning algorithm may be of the large language model type, autoregressive model, generative model or pre-trained or pre-trained generative Transformer.

[0037] The first ENSi dataset can also encode a list of words according to their index or reference in a dictionary.

[0038] First function Fl for the identification of similar articles

[0039] According to a first example, the first data set ENSi defines an input of a function denoted Fi implementing a first machine learning algorithm MLAi trained to calculate a similarity score between the first data set ENSi and each article of an article database, for example denoted BDa. According to an example, different article databases are consulted BDa, BDb, BDc.

[0040] Such a score can also be called a proximity score between an article on the one hand and the first set of ENSi data on the other hand.

[0041] The execution of the first function Fi makes it possible to extract ART articles; from at least one database BDa to generate a first list LISTi whose similarity score is greater than a given threshold. The list of articles LISTi can be ordered so as to index the articles according to their similarity score. The threshold can be fixed or considered so as to retain a given number of ART articles; similar to the set ENSi.

[0042] The method of the invention therefore comprises the extraction, noted EXTi in [Fig.l], of a set of articles ARTi or more generally of digital documents corresponding to scientific articles, abstracts, patent documents, theses and any other document accessible within a first database of articles BDa.

[0043] Each item comprises a set of discrete symbols encoding data in a natural language that can be analyzed by a machine learning algorithm, denoted MLAb

[0044] For this purpose, a uniform resource locator, commonly referred to as URL, makes it possible to execute a search in at least one remote BDa database. The method of the invention can allow the configuration of a plurality of URLs; in the case where a search for articles in several databases would be configured.

[0045] The first ENSi data set may also include concepts extracted from the text received in natural language. The extraction of concepts may be carried out for example using a semantic graph and a calculation of the proximity of concepts in the graph from the definition of a distance in the graph. Such a distance is generally used to calculate the distance between different nodes of the semantic graph and therefore their proximity due to the fact that a given number of edges in the graph connects two concepts in the graph to each other. Article databases

[0046] According to one embodiment, the method of the invention comprises a configuration making it possible to query a plurality of article bases BDa, BDb, BDc. [Fig.2] re presents an example in which a plurality of article bases are queried in order to process a large number of scientific articles from different sources and to extract a maximum number of PROi health studies.

[0047] It is understood that this phase can be carried out independently of the reception of a new request comprising a first set of ENSp data. Thus, a database BD2, described below, grouping the references to the extracted PROi health studies and the associated ART; articles can be created and this database BD 2 can then be updated when the article bases are enriched over time with new articles.

[0048] However, upon a new request and the receipt of a new ENSi set, an update of the extraction of new articles comprising data characterizing health studies can be carried out so that the enrichment of the BD2 database is continuously carried out. According to one embodiment, the first MLAi algorithm is executed to identify a set of articles regardless of whether or not they may contain data describing studies. Thus, in this case, the first MLAi algorithm makes it possible to identify articles close to the first ENSi set by means of the implementation of a predefined semantic distance. In a second step, the articles comprising data characterizing health studies will be selected from a second MLA2 algorithm applied to all of the selected articles.

[0049] The EXTi extraction can be carried out on only part of the BDa base, or any other BDb or BDc base according to a date criterion in order to select the most recent articles and to apply the MLAi algorithm or the Fi function on these new articles from the BDa base or another. First machine learning model

[0050] A pre-trained MLAi machine learning model is then implemented from the ENSi input data organized in the form of an input vector. According to one embodiment, the MLAi machine learning model is configured to calculate a similarity score between each ART article; of the BDa database and the first data set.

[0051] According to another example, an algorithm implementing a technique other than a machine learning model can be implemented to calculate the similarity according to a predefined distance between the first ENSI set and an article. According to different examples, a Cosine similarity type function, the Jaccard distance, or even the N-gram technique.

[0052] In order to calculate an INDS similarity score, the method of the invention makes it possible, for example, to use a semantic distance based on a vector product. A semantic distance can then be defined to calculate a distance between two vectors encoding data specific to two distinct data sets, such as the data from the first ENS1 set and all or part of the data from an ART article;.

[0053] According to one embodiment, the MLAi machine learning model implements a regression which makes it possible to generate a proximity score between each ART article; and the first ENSp data set. The articles having a score above a given threshold can be retained and ordered in an order going from closest to most distant with respect to the first ENSp data set. The proximity is defined by an INDS similarity score which can be calculated from a regression of the MLAb learning model. The MLAi machine learning model can be configured to define a semantic distance thanks to the training of said model.

[0054] According to an exemplary embodiment, before the processing of each article ART;, a language detection of the article ART; is carried out. When an identified language does not correspond to the language of the data of the first set of data ENSi, then a preprocessing aims to automatically translate the first set of data ENS i into the language of the article. A new set of data ENSi' is then considered to identify the presence of relevant studies in the article considered. In the remainder of the description, the notation ENSi is retained, regardless of the language considered.

[0055] An advantage of the invention is to integrate a preprocessing capable of processing natural language and identifying not only PRO relevant health studies; but also their language. Second function F2 for study extraction

[0056] When the first list LISTi is generated, a second function F2 can be implemented to extract the ART articles; comprising data characterizing scientific studies noted PRO;.

[0057] According to one embodiment, a machine learning algorithm MLA2 is configured to classify the articles according to classes. A first class CLi corresponds to the articles comprising data characterizing health studies PRO;. To this end, the algorithm is trained from training data comprising articles. A first set of articles comprises data characterizing health studies and a second set of articles does not comprise data characterizing health studies. Thus, supervised training implementing a cost function makes it possible to train the machine learning algorithm in order to classify the articles defining the input data for example of a neural network.A CNN type neural network meaning in Anglo-Saxon terminology "convolutional neural networks" and called convolutional neural network can be implemented to implement a classifier and extract data characterizing health studies.

[0058] The data characterizing the health studies may include, for example, a reference to a health study, test results, questionnaires and associated responses for a given cohort of patients, study authors, titles, summaries or abstracts, keywords, and conclusions.

[0059] The health study may be named by a given designation, it may in particular be identified thanks to the recognition of terms specific to health trials on a given population. Finally, elements of context, references to patients, to a test protocol, to dates or types of trials may be used to identify the presence of a reference to a health study.

[0060] According to one embodiment, the function F2 implements a syntactic analysis of the articles. This syntactic analysis can be carried out automatically by means of a machine learning algorithm, such as a neural network,

[0061] A neural network may be trained to classify a set of ART articles;. According to one example, a class of CLi articles is generated for which a citation to a health study has been identified in said articles of the class.

[0062] Other characteristic data may characterize these health studies. It is recalled that these studies are rated PRO;.

[0063] Alternatively, this classification can be carried out from a function implementing an expert system comprising a knowledge base and rules.

[0064] According to another embodiment which may be complementary, the machine learning model MLA2 is configured to calculate a proximity score of each article or part of the ART; articles of the BDa database comprising a reference to a health study, noted PRO;. According to this example, a single machine learning model grouping the functions Fi and F2 can be implemented. To this end, the model makes it possible to generate a proximity score according to the definition of a semantic distance between the first set ENSi and each ART; article processed and to classify said articles according to a set of classes making it possible to identify the articles comprising a reference to a PRO; health study. Named entities

[0065] According to one example, recognition of named entities in natural language text may be performed. Different techniques may be implemented.

[0066] According to a first example, a dictionary of named entities may be used. According to one example, a remote server having a named entity recognition algorithm may be used.

[0067] According to this example, access to a dictionary with many words, synonyms and a vocabulary collection can be configured. The method then makes it possible to check whether a particular named entity is also available in an article. In Using a string matching algorithm, cross-entity verification is performed.

[0068] In order to identify words to be checked, a software component can use on the one hand identified words from the first data set and words generated from a generative machine learning algorithm making it possible to multiply concepts and keywords close to the first data set.

[0069] According to one embodiment, a digital connector, also called an “application programming interface”, designated by the acronym API and designating in English literature “application programming interface” can be used in order to extract named entities from a free text in natural language and possibly to generate new ones and to compare them to a resource accessible from a data network.

[0070] The use of a named entity recognition algorithm within the second function F2 makes it possible to obtain better performance in detecting the presence of named entities. Indeed, the named entities characteristic of a health study can be more easily detected using such an algorithm.

[0071] According to a second example, a system based on predefined rules can be implemented to recognize named entities. According to one example, a set of model-based rules can exploit a morphological pattern or a string of words used in the document. According to another example, context-based rules such as contextual rules dependent on the meaning or context of the word in the document can be implemented.

[0072] Finally, according to a third example, a machine learning algorithm, also called a machine learning algorithm, can be implemented. In this case, statistical modeling is used to detect named entities. A representation based on the characteristics of the document is used. One interest is to recognize types of named entities despite slight variations in their spelling.

[0073] In this case, the Fi function ensures the generation of a proximity score and the second function ensures the detection and extraction of named entities ENT; characterizing a citation to a PROi health study.

[0074] By executing the first function Fi and the second function F2, a selection of ART articles is obtained; having at least one reference to a PRO health study; and having an INDS similarity score with the first data set ENSp. The set of articles thus selected makes it possible to define a first list of LISTp articles. ML Al Learning

[0075] According to one embodiment, the learning of the first learning algorithm MLA1 machine is realized in such a way that an INDS similarity score is generated between on the one hand each ART article; or each extracted dataset characterizing a PRO health study; and on the other hand the first ENSi dataset. The learning can aim to configure the data quantification of a considered article to calculate the proximity or similarity with a smaller ENS i dataset. MLA2 Learning

[0076] The method of the invention comprises training the MLA2 machine learning model. This training takes into account a first set of training data comprising several tens or hundreds of article references which are known to contain references to studies.

[0077] When the learning model is configured to identify ART articles; comprising PRO health studies;, the learning makes it possible to train the model to recognize a study comprising a reference and to extract the elements characterizing this PRO health study; in the ART article; in question. The learning also optionally takes into account a second training set of articles not comprising a reference to a health study. This learning makes it possible to improve the detection of an ART article; comprising a reference to a PRO health study;.

[0078] In the case of the implementation of a classifier, the latter may comprise several classes making it possible to reference or categorize ART articles; comprising a PRO health study;. A syntactic and / or semantic analysis makes it possible to identify PRO health studies; in connection with the first set of data.

[0079] Learning can therefore include training with input text types that can describe PRO health studies differently; considering different ways of evoking the problem of a health study. One advantage is to train the model with a wide variety of PRO health study descriptions.

[0080] Furthermore, the semantic field of a health study concerning a given theme can also be taken into account during learning.

[0081] According to one embodiment, the method of the invention comprises training on particular therapeutic fields. This training makes it possible to take into account, for example, the semantics of particular fields such as the field of geriatrics, the field of menopause, the field of oncology, the field of myasthenia, the field of cardiovascular diseases, etc.

[0082] One advantage of this training is that it allows you to quickly score an ART article according to a domain and to discard it or retain it depending on its proximity to the input data. Expert system

[0083] According to one embodiment, an expert system type algorithm can be used. Such an expert system type algorithm is based on the development of rules, such as business rules and a knowledge base. One advantage is to inject knowledge into the data processing to exploit a corpus of documents when they have a well-identified given structure or when certain elements can be identified in a recurring manner within said articles. For example, when specific numbers are assigned to types of PRO health study, such a number and / or its type can be searched directly within the articles.

[0084] According to a first case, an expert system type algorithm can be used to replace the MLA2 machine learning algorithm. In this case, the expert system is configured to firstly identify the ART articles; comprising at least one reference to a PRO health study;, and extract the information characterizing these PRO health studies;. A dictionary and a thesaurus can for example be used to generate a proximity score.

[0085] According to one embodiment, a second semantic analysis can be carried out more specifically on the data extracted from the article corresponding to at least one PRO health study; with the objective of evaluating the proximity between on the one hand the relative extracted data characterizing the PROi health studies which are more specific than the set of data from the ARTi article and on the other hand the first set of ENSi data. A second similarity score INDs2 can then be generated.

[0086] In a second case, such an expert system type algorithm can be used in combination with the MLA2 machine learning algorithm. For example, the identification of data characterizing a health study or a reference to a health study is carried out from an expert system using rules and a knowledge base. When the ART articles; include a reference to a PRO health study; or include data characterizing at least one PROi health study, an MLAi machine learning algorithm is configured to generate a proximity score between, on the one hand, the data characterizing the PRO health study; or those of the ART article; and, on the other hand, the first set of ENSp data Extraction of articles and studies

[0087] The method of the invention comprises an EXT2 extraction of the ART articles; from the BDa database which comprise: • on the one hand a similarity score among the highest with the first set of ENSi data and; • on the other hand a citation, a reference to a PRO health study; or the data describing a PRO health study;.

[0088] The extraction of articles containing health studies can be implemented by two successive extraction operations, including the first extraction EXTi and the second extraction EXT2.

[0089] The data characterizing the PRO health studies; are in turn extracted from the ART articles; and the method of the invention makes it possible to generate a list, noted LIST2, possibly ordered according to the INDS similarity score of PRO health studies;. To this end, a designation of each PRO health study; is generated from the data characterizing said PRO health studies;.

[0090] The function Fi for generating the list LISTi executes the trained machine learning model MLAb then the function F2 for generating the list LIST2 executes the trained machine learning model MLA2When the ART; articles of the list LIST2 are identified, the data characterizing the PROi health studies of the ART; articles having the strongest proximity with the first set of ENSi data are extracted and stored in a memory. Descriptors

[0091] The extraction step EXT2 comprises the identification and extraction of all the descriptors qualifying or describing a PRO health study; within each ART article; processed. The descriptors may relate to a title, a designation, an identifier, a reference, a subtitle, a summary, a note or a comment. According to other embodiments, the descriptors may comprise an extract from the study, a conclusion of the study, one or more authors of the study. A descriptor may also be the language of the PROi health study or that of the ART article; or both. The descriptors according to the articles may not be homogeneous. An advantage of the invention is to standardize and homogenize the extracted data forming descriptors of the study.

[0092] All of the extracted data comprising the descriptors of each identified PROi health study are possibly sent via a data network to a remote server for processing of this data, in particular their recording in a new BD2 database. New BD2 database

[0093] According to one embodiment, the method comprises creating and updating a database BD2 of PRO health studies; extracted from a plurality of ART articles;, possibly a large number of articles in order to enrich the data forming descriptors of these PRO health studies;. This new database is noted BD2. It comprises all the PRO health studies; extracted from the different articles processed by the function F2. The database BD2 records the link or association between each PRO health study; and each ART article; from which it was extracted.

[0094] Thus, the method of the invention allows the creation of a BD2 database for storing the information extracted from the ART; articles, including the identifiers of the PROi health studies, the associated data and the original sources of the studies such as the authors, the context of the study, the description of the cohort, etc. An advantage of the invention is to structure this BD2 database in such a way as to allow a subsequent efficient search and a rapid retrieval of the information when a query is carried out against a BD2 database comprising ART; articles and PRO; health studies.

[0095] [Fig.2] represents a scenario in which the BD2 database is updated following reprocessing of new articles extracted from one or more BDa, BDb, BDc article databases. This updated BD2 study database includes data describing an association between each PRO health study; and each ART article; extracted from one of these BDa, BDb, BDc databases. Comparison

[0096] The method comprises a step aimed at comparing COMPi each PRO health study reference; or each PRO health study designation; or each PRO health study identification; with a set of reference health studies {PROk} stored in a remote memory such as a BDi database accessible from a remote server. This latter BDi database is called in the present description reference database and is noted BDi. The COMPi comparison step aims to generate an availability label Do when a study from the list LIST2 is stored in the BDi database. When a PROi health study is present in the reference database BDb a label Do makes it possible to inform that the given PROi health study is indeed valid and can be used.

[0097] According to an example, different levels of availability are managed in the reference database BDi. For example, a first level of availability is that a study is available because it has been validated in a country other than France, a second level of availability is that a study is available because it has been validated in France. Other parameters can be taken into account in order to generate an availability label Do.

[0098] The first BDi database allows PROk health studies to be stored in an organized manner, which can be classified in different ways. The notation "PROk" is used to designate reference health studies. For example, a PROk health study may include an identifier, a title, an author, a short description, a date, or a list of keywords. A PROk health study in the BDi reference database includes a set of data quantifying the health status of a given population, said data being organized according to different questions and answers. According to an example, the PROk health study may include the results of the study, conclusions or references.

[0099] The COMPi comparison step can be carried out so that each data extracted from an ART article; qualifying or describing the PRO health study; can be used to result in an identification or a correspondence of the present PRO health study; in the BDi reference database.

[0100] An advantage of the method of the invention is to establish a correspondence that is as reliable as possible between a study in the BD2 database and a study in the BDi reference database. This correspondence makes it possible to check whether a health study identified in the ART scientific articles; has already been referenced in another BDb database. When such a correspondence is identified, a component makes it possible to manage the correspondences and to update the data in the BD2 database with the information extracted from the BDb reference database. Thus, the BDi database makes it possible to aggregate data relating to the ART articles; and to the PRO health studies; and thus allows better accessibility to health information.

[0101] According to one embodiment, the method of the invention comprises a duplicate management component. According to one embodiment, when the same study is referenced in different ways in different scientific articles ART;, a processing aimed at associating lines with each other from the second database BD2 is implemented. One interest is to enrich the descriptors of each PRO health study; and to associate them with different scientific articles. Extracting a new list

[0102] The method of the invention therefore comprises a step EXT? of extracting a subset of PRO health studies; and / or associated ART articles; having an availability label Do. Preferably, the third list comprises a reference to a PROi health study and a reference to the article from which it was extracted.

[0103] The method of the invention makes it possible to generate a third list LIST3 when several PRO health studies are labeled with the availability indicator following the COMPi comparison step. Recommendation

[0104] When different PRO health studies are present in the third list LIST3, a recommendation algorithm is configured to calculate a recommendation index INDr of each PRO health study of the third list LIST3.

[0105] The recommendation algorithm is configured to take into account different criteria. According to a first embodiment, the recommendation algorithm takes into account a first criterion Ci of publication date of each PRO health study; to calculate an INDR recommendation index. The more recent the date, the higher the INDR recommendation index is likely to be.

[0106] According to one embodiment, the recommendation algorithm takes into account a second criterion C2 which corresponds to the frequency of use of a given PRO health study. The more a study is used, the more its reliability is approved, thus the higher the frequency of use of a study, the higher the recommendation index INDR. In this embodiment, the reference database BDi of health studies comprises a usage indicator for each study which is updated with each new use of a PRO health study. The criteria Ci and C2 can also be designated in the form of recommendation criteria.

[0107] According to another embodiment, the recommendation algorithm takes into account other criteria, also called recommendation criteria. For example, a third criterion corresponds to a number of citations of the ART article; or of the PRO health study; or of an author of the article, or even of an index calculated from the number of citations of the ART article;. According to another example, the recommendation algorithm takes into account a fourth criterion corresponding to the presence of a right associated with the ART article; or with the PRO health study;. According to another example, the recommendation algorithm takes into account a fifth criterion corresponding to the language of the PRO health study; or that of the ART article;.

[0108] According to an exemplary embodiment, the recommendation algorithm takes into account the INDS similarity index calculated previously using the machine learning algorithm.

[0109] The method of the invention comprises a step GENi aimed at generating a fourth list LIST4 of PRO health studies; ordered according to the INDR recommendation index. This list can also be noted LIST4(PRO;) in [Fig.l]. According to an example, this INDR recommendation index corresponds to a weighted value of all the criteria taken into account. [Fig.l] represents the notation INDr(Ci, C2, INDs) illustrating that according to this embodiment, the INDR recommendation index can depend on the first criterion Ci, the second criterion C2 and the INDS similarity index. System

[0110] [Fig.3] represents an example of architecture of a system of the invention comprising at least one SERV2 server hosting a BD2 database made accessible remotely via a CONi administration console and a data network denoted NET 1 and which may for example be the Internet network. This CONi administration console makes it possible to edit the data of the BD2 study database, process duplicates, enrich fields, and manage user rights.

[0111] A terminal T1 is represented and can be used to define a new request describing the first ENSp II data set can correspond to a PC-type computer, a smartphone, otherwise known as a smart phone, or a digital tablet. The Ti terminal allows a user to check the existence of study(s) of existing health in a given medical field close to that which the user would like to carry out. Thanks to the BD2 database built and updated regularly and the BDi reference database, he can check the availability of a study validated by an organization having medically verified this study.

[0112] [Fig.3] also represents a database BDa of articles, such as a database of scientific articles accessible from a user terminal Th

[0113] The system is configured to implement the method of the invention. The system comprises in particular a calculator implemented in each server to carry out the operations of each algorithm, in particular those implemented by the functions Fi and F2. In addition, a calculator can be used to carry out the operations of comparing the health studies with the reference studies. Furthermore, memories can be implemented to store and record extracted or calculated data.

Claims

Claims

1. A method for recommending a qualified health study comprising: • Receipt (RECi) of a first set of data (ENSi) encoding free text describing the objective of a study characterizing the health status of a set of patients; • Extraction (EXTi) of a first list (LISTi) of articles (ART;) from at least one database (BD2, BDa, BDb), from the execution of a first function (Fi) implementing a first machine learning algorithm (MLAi) trained to generate a set of semantic similarity indices (INDs) between the first data set (ENSi) and each article (ART;) identified from said at least one database, said first list of articles (LISTi) corresponding to a set of articles (LISTi) having a similarity index (INDs) greater than a predefined threshold, said first data set (ENSi) defining an input of the first function (Fi); • Extraction (EXT2) of a second list of health studies (LIST 2) from the execution of a second function (F2) implementing a second machine learning algorithm (MLA2) trained to classify the articles (ART;) of the first list (LISTi), a first class (CLi) corresponding to the articles (ART;) comprising first data describing at least one health study (PRO;), a health study characterizing the state of health of a cohort of patients and being called “health study” (PROi), said first data being automatically extracted to define a second list of health studies (LIST2); • Comparison (COMPi) of the first data of each health study (PROi) from the second list (LIST2) extracted with a set of studies (PROk) present within a first memory of a data server (SERVi), the result of said comparison (COMPi) making it possible to generate an availability label (Do) when a study from the second list (LIST2) is present in said first memory; • Extraction (EXT3) of a third list of health studies (LIST3) having an availability label (Do); • Generation (GENi) of a fourth list (LIST4) of health studies (PROi) according to a recommendation index (INDR), said recommendation index (INDR) being calculated from a first criterion (Ci) of date of each study (PRO) of the third list (LIST3), a second criterion (C2) of frequency of use of each study (PRO) of the third list (LIST 3) and the similarity index (INDs) of each study (PRO) of the third list (LIST3).

2. Method according to claim 1 characterized in that the first machine learning algorithm (MLAJ) comprises a software component making it possible to automatically generate a translation of the first data set (ENSi) into a given language.

3. Method according to any one of the preceding claims, characterized in that the semantic similarity index (INDs) is calculated from a definition of a semantic distance.

4. Method according to the preceding claim characterized in that the calculation of the semantic distance between two sets of data is implemented by the first machine learning algorithm (MLAi), two input vectors being constructed from the first set of data (ENSi) and the data characterizing an article (ARTi), said semantic distance making it possible to calculate a distance between said two vectors encoding data specific to two distinct sets of data.

5. Method according to any one of the preceding claims, characterized in that the first machine learning algorithm (MLAi) comprises the implementation of a generative autoregressive language model of the Transformers type trained on the one hand from a set of data describing in natural language a health study and on the other hand from data quantifying the state of health of a given population, said data being organized according to different questions and answers.

6. Method according to any one of the preceding claims, characterized in that the second function (F2) implements a syntactic analysis comprising the execution of the second machine learning algorithm (MLA2) trained to classify a set of articles (ART;), a class of articles (CLi) being generated for which a citation to a health study was identified in said articles of the class, said second machine learning algorithm (MLA2) comprising the identification and extraction of named entities (ENTi) characterizing a citation to a health study (PROi).

7. Method according to any one of the preceding claims, characterized in that the second machine learning algorithm (MLA 2) comprises the calculation of a confidence score associated with each article (ART;) classified in the first class (CLi).

8. Method according to any one of the preceding claims, characterized in that the data extracted from the articles (ART;) corresponding to the health studies (PRO;) comprise metadata comprising titles, identifiers, author names, themes, author names.

9. Method according to any one of the preceding claims, characterized in that the second machine learning algorithm (MLA 2) comprises a second class comprising the articles (ART;) not comprising health studies (PROi), a selection of said articles (ART;) of the second class being made accessible to a user by means of a display.

10. Method according to any one of claims 6 to 9, characterized in that it comprises the acquisition of a classification indicator making it possible to reinforce the training of the first algorithm (MLAi) and / or the second algorithm (MLA2), said classification indicator making it possible to indicate that a classified health study (PROi) does not correspond to the first data set (ENSi) or that an article does not include a health study (PROi) while said article has been classified in the first class (CLi).

11. Method according to any one of the preceding claims, characterized in that it comprises: • Extraction of the health studies (PROi) from each article (ART; ) of the fourth list (LIST4) defining new entries of a first database (BD2); • Addition of said new entries in the database (BD2 )•

12. Method according to claim 11 characterized in that a study database (BD2) is updated, said update comprising: Recording of the fourth list (LIST4) within said studies database (BD2); said recording comprising a first recording of all the articles (ART;) of the fourth list (LIST4) defining a first entry of the database (BD2) and a second recording of all the health studies (PRO;) of each article (ART;) of the fourth list (LIST4) defining a second entry of the first database (BDi) and recording of the associations between the health studies (PRO;) extracted from the articles and the articles themselves.

13. Method according to any one of the preceding claims, characterized in that the recommendation index (INDR) is calculated from at least one other recommendation criterion among which we find: the publication date, the similarity index, the presence of a right associated with a study, the citation of an article or an author, a citation index of the article.

14. System comprising a server hosting a database (BD2) and a terminal (Ti) configured to implement the method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Systems and methods of clinical trial evaluation

    US20200381087A1

  • Medical Literature Recommender Based on Patient Health Information and User Feedback

    US20210391075A1