METHOD FOR ASSOCIATING DATA WITH A DIGITAL DOCUMENT, ASSOCIATED SYSTEM
The method uses GPT-3 transformers to analyze and structure data subsets, generating keywords and queries, enhancing data collection efficiency and organization for digital documents.
Patent Information
- Application Number
- FR2022005349
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-02
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-06-02
AI Technical Summary
Existing data collection methods for digital documents are inefficient in handling heterogeneous data, failing to structure and prioritize data collection, and struggle with organizing and qualifying data relevance during the document preparation process.
A method involving multiple learning algorithms (GPT-3 transformers) to analyze and structure data subsets, generate keywords, and create text queries, allowing data to be associated with a digital document while preserving navigation, using a superimposed graphical window for data organization and management.
Enables efficient, structured data collection and organization, reducing noise and optimizing computing resources, allowing quick decision-making and in-depth analysis of relevant data for document creation.
Smart Images

Figure 00000018_0000 
Figure 00000019_0000 
Figure 00000019_0001
Abstract
Description
Title of the invention: METHOD FOR ASSOCIATING DATA WITH A DIGITAL DOCUMENT, ASSOCIATED SYSTEM Field of invention
[0001] The invention relates to a method for associating a set of data with a digital document such as a document being prepared. The field of the invention relates to methods for assisting and helping a user of a data network in collecting data for the preparation of a digital document. The field of the invention relates more particularly to methods for generating non-homogeneous data containers for their subsequent exploitation. State of the art
[0002] There are data containers that allow data to be gathered or collected from a data network for processing. We know the "shopping cart", a function generally offered on online shopping sites on the internet that allows data to be gathered during a consultation of a WEB page without losing the thread of navigation on the network. This function has the advantage of parallelizing tasks carried out on a data server while retaining navigation histories in order to facilitate the operations of a user consulting different data resources. This function is useful when we wish to collect data for a single operation, that is to say an online payment of a sum. However, it is not suitable for carrying out different operations of different natures handling heterogeneous data.
[0003] There are also options in available browsers that allow you to display resources from a data network to pin content and add it to a favorites list. However, these solutions do not allow you to prioritize the collected content and differentiate it according to a document to be developed later.
[0004] There are also data containers defining electronic tools for taking notes for writing a dissertation, however most of the time these solutions do not allow the document to be structured based on the data collected in the notes.
[0005] However, some digital documents require the collection of large amounts of data of a heterogeneous nature. The user of word processing software wastes a significant amount of time gathering data in an unstructured manner. Finally, when collecting From this data, it is difficult to judge at the same time the relevance of the data collected and their arrangement in a document.
[0006] For this purpose, a "shopping cart" type function, or a function for feeding browser favorites or even a software function for producing a digital notebook do not allow data to be collected while archiving a data structure and qualifying the resources of this data in the same process.
[0007] There is therefore a need to define a data container making it possible to provide assistance in structuring a digital document according to the nature of the data collected. Summary of the invention
[0008] According to a first aspect, the invention relates to a method for associating data produced with a first digital document from a data resource displayed on a display of computer equipment, said method comprising: • Displaying a first set of data accessible from at least a first uniform resource locator in a first window of a browser of a data network; • Actuating a digital control to extract a first subset of data displayed in the first window and to record said first subset of data in a memory; • Executing a first learning algorithm comprising a language model and being pre-trained with a first training data set, said language model comprising a statistical model that models the distribution of discrete symbol sequences in a natural language, to generate a first output data set from the extracted at least one subset of data and from a first data model defining a first specific training domain, said first specific training domain comprising a data set defining inputs of the first algorithm and sets of desired outputs of the first algorithm; • Executing a second learning algorithm comprising a second language model to generate a set of keywords defining a second set of outputs from at least one extracted subset of data; • Execution of a third learning algorithm comprising a third language model to generate a set of text queries forming entries of a search engine comprising a set of resources indexed on a data network; • Generation of said queries within a first search engine, retrieval of a set of uniform resource locators returned by the search engine and processing of said resources to filter them according to a predefined criterion, said filtered resources defining a third output set; • Generation of a second graphical window superimposed on the first window, said second window displaying at least one data item from each set of outputs and a first digital actuator making it possible to record the first uniform resource locator and said output data produced in a memory, the actuation of the first actuator causing the creation of an association of the output data with a user identifier.
[0009] One advantage is to allow navigation on pages of a network of data of interest while allowing analysis of the consulted data providing relevant indications quickly and a means to collect this data without losing the thread of navigation.
[0010] According to one embodiment, the actuation of the first actuator causes the creation of an association of the output data with an identifier of a memory space and / or with a first predefined digital document.
[0011] One advantage is that it allows a document to be dynamically constructed by annotating it during a discussion with information of interest which can be organized and ordered according to a predefined plan or structure or which can be modified according to the data collected.
[0012] According to one embodiment, the first output data set is a natural language summary or opinion of the extracted data subset when the latter is natural language text.
[0013] One advantage is that it allows a user to make a decision quickly by consulting numerous data resources.
[0014] According to one embodiment, the method comprises processing the first subset of data to homogenize said data.
[0015] One advantage is to reduce noise and optimize the output of the learning function(s).
[0016] According to one embodiment, a second digital actuator allows access to all of the data of at least one set of output data.
[0017] One advantage is that it allows for in-depth analysis on demand depending on the context of the data produced. The second window provides decision support on whether to retain relevant data, consult it at the time of analysis, or consult it later.
[0018] According to one embodiment, the method comprises producing a set of data as input to the third learning algorithm defining a specific training domain, said third learning algorithm being pre-trained from a generic domain, said specific training domain comprising a second data model defining a second specific training domain, said second specific training domain comprising natural language texts as input and natural language queries defining desired outputs of the third learning algorithm.
[0019] According to one embodiment, the first learning algorithm, the second learning algorithm and the third learning algorithm are GPT-3 type transformers designating “generative Pre-Training Transformer”.
[0020] One advantage is that it allows queries to be defined in natural languages with a syntax or grammar specific to a specific domain.
[0021] According to one embodiment, the first learning algorithm has a dimension of 12288, the second learning algorithm has a dimension of 1024 and the third learning algorithm has a dimension of 4096.
[0022] One advantage is to separate the correct computing capacity for each function used. The separation of dimensions according to the different algorithms makes it possible to produce the most optimized output depending on the result that one is seeking to obtain. Thus, such a configuration offers a good compromise between the relevance of the results returned by the three algorithms and reduced computing time.
[0023] According to one embodiment, the method comprises a configuration step aimed at defining a size of the output data of the first algorithm, a maximum number of keywords output from the second algorithm and a maximum number of queries output from the third algorithm and a maximum number of resource locators for each query produced by the third algorithm.
[0024] An advantage is that it allows the creation of a second window containing rich and varied information that can be analyzed by a user. The user is not overwhelmed by a multitude of information; the invention allows the production of sufficient data to help the user make a decision.
[0025] According to one embodiment, the method comprises a step aimed at associating a plurality of output data with the same first digital document.
[0026] One advantage is to provide assistance in designing a document from several resources collected during one or more navigations on a data network.
[0027] According to one embodiment, the association of all the output data with the first digital document comprises an indexing of said data according to an ordered sequence of a plurality of output data corresponding to other sub- datasets from the same dataset or from another dataset.
[0028] According to a second aspect, the invention relates to a system comprising at least one data server comprising a memory and calculation means making it possible to define a workspace in which at least one digital document being developed is recorded and user profile data, the system further comprising an electronic terminal provided with a display, a memory and a calculator, said system being configured to carry out the steps of the method of the invention. Brief description of the figures
[0029] Other characteristics and advantages of the invention will emerge on reading the detailed description which follows, with reference to the appended figures, which illustrate:
[0030] [Fig-1]: an example of steps implemented according to an example of realization of the process;
[0031] [Fig.2]: an example of representation of a browser window and of representation of a window produced by the method of the invention;
[0032] [Fig.3]: an example of a system comprising a data server to execute the learning functions. Description of the invention
[0033] [Fig.l] represents an exemplary embodiment of the different steps of the method of the invention. According to one embodiment, when browsing on a data network, such as the Internet, the user accesses content that is displayed on a display of an electronic terminal such as a personal computer, a mobile telephone such as a Smartphone, a virtual reality headset, a display device with an augmented reality function or even a digital tablet. The content can be stored on one or more remote memories of data servers. A URL, designating a uniform resource locator, makes it possible to retrieve data from a data request sent to at least one data server.
[0034] The method of the invention therefore comprises a first step of displaying AFFi digital content originating from at least one remote server. This display is preferably carried out from a browser, such as a WEB browser. It is more generally software allowing the consultation and display of structured content originating from a data network such as the Word Wide Web. A browser is an http client.
[0035] The method comprises a step aimed at selecting all or part of the content. This step is denoted ACTi in [Fig.l]. It can be activated from a button on a browser or from a menu accessible by the browser and offering commands for accessing and interpreting the displayed data. The button may, for example, be a browser extension module, called a "plug-in", to provide access to a new software function such as that which can be executed by the method of the invention.
[0036] According to one embodiment, in order to select data included in a portion of the displayed page, denoted subset of SSENSi data, a plurality of buttons arranged in different areas of the browser makes it possible to select said subset of SSENSi data. The method of the invention relates more particularly to the case where the SSENSi data are data representing symbols of a natural language.
[0037] [Fig.2] shows an example of a browser with different zones each containing data. The different zones define digital contents of different nature, there can be text zones, titles, menus, images, footers or headers, etc. [Fig.2] illustrates some subsets of SSENSi, SSENS2, SSENS3, SSENS4 data having different characteristics according to their arrangement, their font, etc. In the case of [Fig.2], the user selects the SSENSi subset which includes different paragraphs represented by rectangles in the figure. The selection of all the data of the SSENSI subset results in the production of the F2 window thanks to the method of the invention. An advantage of the selection of a subset of digital content is to obtain a processing of the data by the algorithms described below more relevant.In fact, only the data of interest are processed by the artificial intelligence algorithms, which allows for optimized processing and a reduction in noise in the production of output data.
[0038] According to one example, in order to select the digital content defining the data of the SSENSi subset that is displayed within the browser, a selection tool such as a mouse or a touch screen makes it possible to select a given region of the browser. According to one embodiment, the zones are defined by tags or a machine-interpretable language and can be automatically delimited from a function interpreting the language encoding the data in the browser. Thus, according to this option, a plurality of zones can be segmented automatically. One advantage is to propose an analysis of portions of interest of the page having a structural coherence considered together, such as a succession of paragraphs under a title. Another advantage is not to consider data from another zone that is of no interest to the user or has a distant relationship with other content from another zone.
[0039] When digital content is digital text possibly including metadata, the selected data is stored in a memory. The metadata may correspond to links, URLs, a language, tags of a language interpreted by the browser or other data relating to the displayed data. The memory may be a memory of the electronic terminal or a memory of a remote server.
[0040] The method of the invention comprises a pre-processing step aimed at processing the selected data. According to one embodiment, these processes include automatic actions aimed at homogenizing the data, in particular removing formatting such as underlining, capital letters, font, text style or even removing line breaks. The text thus processed is recorded so as to define an input for an artificial intelligence algorithm.
[0041] First algorithm: generation of summary or opinion
[0042] According to an example, a first learning function ALGOi is implemented to process the selected text. The first learning function is also called a first learning algorithm since the algorithm is instantiated with physical variables and a parameterization to define a function. The instantiation is for example defined by desired sizes of output data and the parameterization can be defined by a specific learning of a generic model carried out from a specific learning domain.
[0043] The first learning function can be executed for example by means of a service offered by a website for example in the form of an API, designating an application programming interface. The first learning function preferably comprises a deep learning model. According to one example, the first learning function is learned with a language model MOD_LANG i which comprises a statistical model which models the distribution of sequences of discrete symbols in a natural language. According to one example, it is pre-trained with a first set of training data TRb In the latter case, the pre-trained model is a BERT system, designating in the English terminology "Bidirectional Encoder Representations from Transformers" or a GPT system, designating in the English terminology "Generative Pre-training Transformer".According to an exemplary embodiment, the GPT-3 language model is implemented in the invention. The DA VINCI model of GPT-3 of dimensions 12288 can be used to obtain the production of a summary or an opinion from a text provided as input to the model. The training data allowing the pre-training can for example come from a Wikipedia corpus. According to another example, the Common Crawl corpus comprising a large number of textual units. sublexicals encoded by the BPE algorithm can be used and / or the WebText2 corpus and / or the Booksl or Books2 corpus.
[0044] According to one embodiment, the method of the invention makes it possible to define two inputs of this algorithm. The first input corresponds to the selected and possibly preprocessed data. This first input can therefore be an English text of a plurality of lines comprising a plurality of sentences.
[0045] The second input comprises a set of models defining a specific training domain TR2 comprising inputs of the first algorithm and desired outputs of said algorithm. According to one embodiment, the set of models comprises input and output pairs of the learning function, said output corresponding to said input. Each input comprises a text, preferably in English, of a given length. Each output comprises a text, preferably in English, the length of which may depend on the desired result. According to one embodiment, the length of the output text may be a parameter of the first learning algorithm ALGOi.
[0046] The output of the second learning algorithm, denoted ENSA, may correspond to a summary of the text delivered as input. According to another example, the output may correspond to an opinion of the text delivered as input. The desired outputs of the specific training domain TR2 may be revised texts obtained with the first pre-trained ALGOi algorithm. According to another case, the desired outputs of the specific training domain TR2 are data produced “by hand”, that is to say generated by one or more individuals. According to another example, the desired outputs comprise a first set of outputs generated by one or more individuals and a second set of outputs generated by the first ALGOi algorithm and modified by an individual.
[0047] The advantage of using a specific training domain TR2 is to obtain more relevant results depending on the desired case. The specific training domain can relate to a technical domain or a semantic domain or a mixture of the two. Second algorithm#: concept generation
[0048] According to an embodiment of the invention, a second learning algorithm ALGO2 is executed from the same input, i.e. the selected data SSENS i and possibly preprocessed. The second learning algorithm is also called second learning function insofar as the algorithm is instantiated with physical variables and a parameterization. The instantiation is for example defined by a number of concepts or keywords desired as output and the parameterization of the algorithm can be defined by a specific learning of a generic model carried out from of a specific learning domain. This second learning algorithm may be of the RNN type designating a recurrent neural network or an LSTM designating “Long short-term memory which is a neural network. According to another embodiment, the second learning algorithm is a transformer of the BERT or GPT type, such as GPT-3. The AD A model of GPT-3 of dimensions 1024 can be used to obtain a set of keywords from a text provided as input to the model.
[0049] The second learning algorithm ALGO2 is trained so as to produce as output a list of concepts, i.e. keywords, relating to the selected text SSENSi. According to one example, the second learning function is learned with a language model MOD_LANG2. Preferably, the second language model MOD_LANG2 is identical to the first language model MOD_LANGi. The output of the second learning algorithm ALGO2 is noted ENSB. For example, the concepts can be produced according to their occurrence in the text and / or from a semantics taken together of the text and therefore of the domain(s) of the text. According to one example, the concepts can be produced by considering a dictionary and / or a thesaurus and / or an ontology. The dictionary makes it possible in particular to extract the different roots of a term or even synonyms in the case of a dictionary of synonyms.The thesaurus allows the extraction of terms from a given domain and the ontology allows the hierarchization of notions between them according to a connected graph of entities and links allowing the structuring of notions between them.
[0050] The second learning algorithm ALGO2 corresponds to the function for producing a list of concepts as output. The training of the second algorithm can be carried out by taking into consideration in the learning the dictionaries, thesauri or ontology relating to the set of sequences of symbols defining words of a given language.
[0051] According to another example, the learning is carried out in a supervised manner by labeling each text used in the learning with given concepts. According to another example, the learning is carried out in an unsupervised manner, that is to say that the texts used for learning are classified at the output of the algorithm into groups which are then annotated.
[0052] The concepts produced by the second algorithm ALGO2 from the text selected SSENSi in the browser are generated in the same window F2 as the output of the first algorithm ALGOp Third algorithm#: URL generation
[0053] Finally, a third learning algorithm ALGO3 is executed from the selected text SSENSi in the browser. The third algorithm ALGO3 includes an implementation of artificial intelligence to produce natural language queries produced from the SSENSp input text
[0054] The third learning algorithm is also called third learning function since the algorithm is instantiated with physical variables and a given parameterization. The instantiation is for example defined by a desired number of queries produced as output and the parameterization can be defined by a specific learning of a generic model carried out from a specific learning domain.
[0055] According to an example, the third learning function is learned with a third language model MOD_LANG3. Preferably, the third language model MOD_LANG3 is identical to the first language model MOD_LANGi or to the second language model MOD_LANG2.
[0056] Such a learning algorithm is for example implemented from a neural network of the RNN type designating a recurrent neural network or an LSTM designating “Long short-term memory which is a neural network. According to another embodiment, the third learning algorithm ALGO3 is a transformer of the BERT or GPT type, such as GPT-3. This third algorithm ALGO3 can be pre-trained so as to produce queries in natural language from an input text. The pre-training can be carried out with a corpus of data representing a plurality of texts in natural language in a generic domain noted TR3. Training carried out from a specific domain TR4 can be carried out from the selected text SSENSi according to the context of the text, metadata of the selected text SSENSi or data present in the displayed and unselected page comprising the text SSENSi.The specific TR4 training domain can be defined from the actions produced by a user with respect to the first results returned from the SSENSi selected text or another SSENSi selected text.
[0057] According to one embodiment, the queries produced each comprise a sequence of symbols defining sentences in natural language. The method comprises a step aimed at generating these queries with an MRb search engine. The MRi search engine is for example an application allowing a user to carry out a search on a data network. The results returned are generally data resources extracted and returned from a query composed of terms in natural language. The resources may in particular be web pages, forum articles, images, videos, files, works, educational sites, applications, open source software. The search engine preferably comprises an index of the content produced prior to the search carried out from at least one query produced by the third algorithm ALGO3.
[0058] An MRi search engine generally returns a list of locators of the identified resources in the form of a results page ordered according to a relevance criterion. The relevance criterion is generally calculated from a score representing a probability calculated by comparing the query and the content of a resource.
[0059] According to one embodiment, a plurality of search engines MRI, MR2, etc. are used to produce results to the different queries generated by the third algorithm ALGO3.
[0060] The method of the invention makes it possible to select a part of the resource locators produced in order to gather them within a list generated in the second window F2.
[0061] In order to select only a portion of the locators returned by a search engine obtained by means of one or more queries, filters can be configured. The filters make it possible to delete, for example, types of locators according to their category or according to the address of the locator. The method of the invention also makes it possible to configure the assignment of a priority to the locator, for example to sort and select them by importance. The filtering can be carried out from a predefined criterion and / or a dictionary containing a list of locators of interest.
[0062] According to one embodiment, it is possible to retain a limited number of locators per query according to a relevance criterion and according to the number of queries generated.
[0063] According to one embodiment, the method of the invention comprises a step which aims to count the links activated or consulted by a user. This step makes it possible in particular to generate a specific training domain Window F2
[0064] The method of the invention comprises a step GEN2 aimed at generating a window F2. Thus, the window F2 produced by superimposing the window Fi comprises a set of data gathered from the outputs ENSA, ENSB and ENSC of the three algorithms ALGOi, ALGO2 and ALGO3 and from the generation of the queries GENi. The second window F2 is advantageously generated by superimposing the first window Fb; such a window is referred to in the English technical literature as a “pop-up”. From the data synthesized in the window F2, a user is able to quickly take action regarding the selected content; in fact, the first set ENSi allows him to quickly understand the subject or opinion of the text, the keywords allow to represent a semantic field relative to the selected text and resource locators allow to quickly access content related to the selected text. In other words, a user of a data network having a browser and consulting data resources can very quickly qualify the data which is consulted.
[0065] The method of the invention not only makes it possible to display this data in a window F2 superimposed on the window Fi but also to propose actions with regard to the collected data.
[0066] A first action consists of recording the data of the sets ENSA, ENSB and ENSC in a memory. According to this embodiment, the memory can be that of a remote server or that of a local electronic terminal. A first digital button Bi making it possible to generate a command to record the content gathered in the second window F2 can be arranged within the second window F2. Preferably, the data thus recorded from the three sets are associated with a user profile Ub. The user profile can itself be associated with a user identifier and a password of a workspace offering functions making it possible to organize one or more digital documents Db D2, etc.According to one example, a user's workspace may include a plurality of digital documents being developed such as a thesis, a technical report, a publication project, a follow-up document for an intern or a person completing a thesis, etc. According to one example embodiment, the data collected within the window F2 may be associated with a digital document Db.
[0067] A digital document in the broadest sense means a data container such as a digital file, for example comprising a predefined format, such as the .doc, .txt, .pdf format or even a database allowing data to be arranged in an organized manner within a memory.
[0068] A second action allows you to consult the DATAi data and to deepen an exploration on the data network of certain data sets, in particular data from the ENSi and ENS3 sets. When the text of the ENSi set is greater than a given size, consulting the data from the ENSi set allows you either to enlarge the F2 window or to generate a new window for viewing the entire data of the ENSi set. When a certain number of URL locators are present in the F2 window but only a limited number are displayed in the F2 window, a command allows you to enlarge the F2 window or to display the list of locators in another window. According to an example, a simple activation of a locator, for example if it includes a hyperlink opens a new browser tab or browser window to display the resource accessible from the locator.
[0069] Association of data with a user and / or a document
[0070] According to one embodiment, the method of the invention comprises a step of associating the collected content DATAi available in the window F2 with a user identifier IDui. According to another example, the method of the invention comprises a step of associating the collected content DATAi available in the window F2. According to one exemplary embodiment, a user can associate the content collected in the window F2 with a user identifier Ui and then, in a second step, it is possible to associate this content DATAi with a specific digital document Di in his workspace. According to one exemplary embodiment, the method comprises a step making it possible to select all or part of the content of the window F2 and to save it in the workspace of the user Ui.Thus, thanks to the method of the invention, a user Ui can only decide to save the ENSA data and associate them with a digital document Di without having to save all of the DATAi data. According to another example, the method of the invention offers a function making it possible to save the DATAi data of each set ENSA, ENSB and ENSC in a first step and then to select a part of it, for example the data of the set ENSC in order to associate them with a digital document Db Multi-association with the same digital document.
[0071] According to one embodiment, the method comprises a plurality of associations of DATAi data to the same digital document Db
[0072] The method comprises a step of automatically arranging the different DATAi contents. The automatic arrangement may comprise, for example, an ordering of the different DATAi contents according to a given plan or structure. According to another example, a proposal for ordering the different DATAi contents makes it possible to organize the document Di with the data extracted from the different DATA contents.
[0073] [Fig. 3] illustrates an example of a system of the invention comprising an electronic terminal Ti and a remote data server SERVi for storing the data produced by the different algorithms implemented by the method of the invention. The data network NETi is for example the internet network. The server SERV2 is for example a server for executing the learning algorithms with predefined configurations. According to one example, the algorithms executed on the server SERV2 deliver data by means of an API accessible via the data network NETi.
Claims
1. Claims Computer-implemented method for associating produced data (DATAi) with a user identifier (IDm), said user being identified within a workspace comprising at least one memory allowing the recording of a first digital document (DJ) from a data resource displayed on a display of computer equipment, said method comprising: • Display (AFFi) of a first data set (ENSi) accessible from at least a first uniform resource locator (URLi) in a first window (Fi) of a browser of a data network (NETJ; • Actuation (ACTi) of a digital control to extract a first subset of data (SSENSi) displayed in the first window (FJ and to record said first subset of data (SSENSi) in a memory; • Execution (EXECi) of a first learning algorithm (ALGOi) comprising a language model (MOD_LANG i) and being pre-trained with a first training data set (TRJ, said language model (MOD_LANGi) comprising a statistical model which models the distribution of discrete symbol sequences in a natural language, to generate a first output data set (ENSA) from the at least one extracted data subset (SSENSi) and from a first data model (MODo defining a first specific training domain (TR2), said first specific training domain (TR2) comprising a data set defining inputs of the first algorithm (ALGOi) and sets of desired outputs of the first algorithm (ALGOi); • Execution (EXEC2) of a second learning algorithm (ALGO2) comprising a second language model (MOD_LANG2) to generate a set of keywords defining a second set of outputs (ENSB) from of at least one extracted data subset (SSENSi, SSENS2); • Execution (EXEC3) of a third learning algorithm (ALGO3) comprising a third language model (MOD_LANG3) to generate a set of textual queries forming inputs of a search engine (MRi) comprising a set of resources indexed on a data network (NETi); • Generation (GENi) of said queries within a first search engine (MRi), recovery of a set of uniform resource locators (URL;) returned by the search engine (MRi) and processing of said resources to filter them according to a predefined criterion, said filtered resources defining a third output set (ENSC);• Generation (GEN2) of a second graphic window (F2) superimposed on the first window (Fi), said second window (F2) displaying at least one data item from each set of outputs (ENSA, ENSB, ENSc, DATAi) and a first digital actuator (Bi), the actuation of the first actuator (BJ initiating the recording of the first uniform resource locator (URLi) and said output data (ENSA, ENSB, ENSC, DATAJ produced in a memory, the actuation of the first actuator (Bi) causing the creation of an association of the output data (DATAi) with a user identifier (IDui).;
2. Method according to claim 1 characterized in that the actuation of the first actuator (Bi) causes the creation of an association of the output data (DATAi) with an identifier of a memory space and / or with a first predefined digital document (DJ.
3. Method according to claim 1 characterized in that the first output data set (ENSA) is a summary or opinion in natural language of the extracted data subset (SSENSi) when the latter is a text in natural language.
4. Method according to claim 1 characterized in that it comprises processing of the first subset of data (SSENSi) to homogenize said data.
5. Method according to claim 1 characterized in that a second digital actuator (B2) allows access to all the data of at least one set of output data (ENSA, ENSB, ENSC).
6. Method according to claim 1 characterized in that it comprises the production of a data set (M0D2) as input to the third learning algorithm (ALGO3) defining a specific training domain (TR4), said third learning algorithm (ALGO3) being pre-trained from a generic domain (TR2), said specific training domain (TR4) comprising a second data model (M0D2) defining a second specific training domain (TR4), said second specific training domain (TR4) comprising texts in natural languages as input and queries in natural languages defining desired outputs of the third learning algorithm (ALGO3).
7. Method according to claim 1 characterized in that the first learning algorithm (ALGOi), the second learning algorithm (ALGO2) and the third learning algorithm (ALGO3), are transformers of the GPT-3 type designating “generative Pre-Training Transformer”.
8. Method according to claim 7 characterized in that the first learning algorithm (ALGOi) has a dimension of 12288, the second learning algorithm (ALGO2) has a dimension of 1024 and the third learning algorithm (ALGO3) has a dimension of 4096.
9. Method according to claim 1 characterized in that the method comprises a configuration step aimed at defining a size of the output data of the first algorithm (ALGOi), a maximum number of keywords at the output of the second algorithm (ALGO2) and a maximum number of queries at the output of the third algorithm (ALGO3) and a maximum number of resource locators for each query produced by the third algorithm (ALGO3).
10. Method according to claim 1 characterized in that it comprises a step aimed at associating a plurality of output data (DATAi) with the same first digital document (Di).
11. Method according to claim 10 characterized in that the association of the set of output data (DATAi) with the first document digital (DJ) comprises an indexing of said data (DATAi) according to an ordered sequence of a plurality of output data (DATA;) corresponding to other subsets of data (SSENSi) of the same data set (ENSi) or of another data set (ENSi).
12. System comprising at least one data server comprising a memory and calculation means making it possible to define a workspace in which at least one digital document being developed is recorded and user profile data, the system further comprising an electronic terminal equipped with a display, a memory and a calculator, said system being configured to carry out the steps of the method of any one of claims 1 to 11.