METHOD AND DEVICE FOR AUTOMATICALLY GENERATING TEXTUAL CONTENT GUIDED BY AN EXPERT SYSTEM FOR APPLYING BIBLIOMETRIC TECHNIQUES

The method and device use bibliometric techniques to generate contextually relevant search results by forming hierarchical file groupings, addressing the limitations of existing search engines and large language models, ensuring accuracy and efficiency.

FR3158814A3Pending Publication Date: 2025-08-01SCANLITT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2024000756
Authority / Receiving Office
FR · FR
Patent Type
Utility models
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2025-08-01
Estimated Expiration
2034-01-26

AI Technical Summary

Technical Problem

Existing search engines struggle to provide meaningful, contextually relevant search results due to limitations in quantifying query adequacy, lack of intelligent filtering, and high computational resource consumption, while large language models suffer from hallucination effects, leading to inaccurate results.

Method used

A method and device using an expert system that applies bibliometric techniques to generate textual content by forming hierarchical groupings of cited and citing files, leveraging a large language model for filtering and sorting without user intervention, thus avoiding hallucination effects and reducing resource consumption.

Benefits of technology

This approach effectively extracts and processes a corpus to identify influential files and main themes, providing accurate, contextually relevant results with minimal computational and energy resources, while respecting intellectual property rights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000039_0000
    Figure 00000039_0000
  • Figure 00000040_0000
    Figure 00000040_0000
  • Figure 00000041_0000
    Figure 00000041_0000
Patent Text Reader

Abstract

TITLE: METHOD AND DEVICE FOR AUTOMATICALLY GENERATING TEXTUAL CONTENT GUIDED BY AN EXPERT SYSTEM FOR APPLYING BIBLIOMETRIC TECHNIQUES The method (700) for generating content, which comprises:- a step (705) of entering a set of search keywords,- a step (710) of extracting a corpus extracted from relational files from the single corpus of files recorded in the database, according to the searched keywords,- a step (715) of statistically forming two complementary groupings of distinct files to form:- a first partitioned grouping of identifiers of classified and hierarchical cited files,- a second partitioned grouping of identifiers of classified and hierarchical citing files,- a step (745) of extracting the metadata of each grouping,- a step (720) of determining two queries for a large language model, according of the first and second groups formed,each query being representative of a content generation query,- a step (725) of providing the two determined queries to a large trained language model,- a step (730) of receiving the first and second filtered groupings and- a step (735) of providing the first and second filtered groupings. Figure for the abstract: Figure 6,
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: METHOD AND DEVICE FOR AUTOMATICALLY GENERATING TEXTUAL CONTENT GUIDED BY AN EXPERT SYSTEM FOR APPLYING BIBLIOMETRIC TECHNIQUES Technical field of the invention

[0001] The present invention relates to a method for automatically generating textual content guided by an expert system for applying bibliometric techniques. It applies to any set of files including a bibliography containing the references cited in each file, and for example, to the field of search engines for scientific publications, research articles, theses, books and patents. State of the art

[0002] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been conceived or pursued previously. Therefore, unless otherwise indicated, it should not be assumed that any of the approaches described in this section constitute prior art solely because of its inclusion in this section.

[0003] The number of files accessible to individuals increases significantly annually, so much so that it results in difficulty for these individuals to make sense of all the information available in these files.

[0004] Essentially, these accessible files refer to other files, in a more or less methodical manner. For example, scientific publications and patents cite the state of the prior art considered relevant, while reference works refer to other works or other scientific references. Such files are said to be relational, that is to say, they are included through their bibliography in families of files whose analysis, currently expert and complex, reveals unexplained units of meaning.

[0005] The main difficulty, for an individual wishing to understand a subject dealt with in a large number of files, is to be able to identify units of meaning to which the targeted files belong, but also to identify which are the important file(s) and / or themes of the field, that is to say these files and / or themes necessary for a good understanding of the subject.

[0006] On the one hand, it is known that the implementation of ordinary search engines makes it possible to identify, on the basis of a query, files relevant to keywords or the semantic meaning of the query and allows these files to be sorted based on relevance to the query only. Occasionally, these search engines have sorting functions, allowing sorting alphabetically, by author name or by publication date for example.

[0007] However, such search engines are limited to quantifying the query-results adequacy, without worrying about the general meaning of interpretation of the results. In other words, these search engines do not make it possible to provide individuals with a set of files prioritized both according to the query formulated and the specific characteristics of these files making it possible to improve the search results to bring out the main themes linked to the query as well as the founding files linked to this query.

[0008] Furthermore, such search engines do not carry out any intelligent, statistical and contextual filtering and / or grouping of files based on the search carried out.

[0009] On the other hand, we know of search engines based on the implementation of large language models (abbreviated LLM). Such large language models are very effective for thematically grouping files in a given corpus of files. However, these large language models consume considerable computing and energy resources to be trained and implemented, on the one hand, and these large language models suffer from so-called hallucination effects on the other hand. These hallucination effects originate from the probabilistic approach, from one step to the next, implemented during the generation of a response to a given query. The main consequence of these hallucination effects is to make search engines based on large language models inaccurate or even false and therefore unusable. Summary of the invention

[0010] The present invention aims to remedy all or part of these drawbacks.

[0011] To this end, according to a first aspect, the present invention aims at a method for automatically generating textual content guided by an expert system for applying bibliometric techniques from a single corpus of relational files recorded in a database, each file of the single corpus being associated with at least one keyword, which comprises: - a step of entering, via an entry interface associated with a calculation device, a set of search keywords, - a step of extraction, by a calculation device, of a corpus extracted from relational files from the unique corpus of files recorded in the database, according to the keywords searched and keywords associated with each file of the unique corpus, and a set of metadata associated with each file of the unique corpus, - a step of statistical formation, by the calculation device, of two complementary groupings of distinct files, according to the extracted corpus and the metadata associated with each file of the extracted corpus, to form: - a first partitioned grouping of identifiers of classified and hierarchical cited files, - a second partitioned grouping of classified and hierarchical citing file identifiers, - a step of extracting the metadata of each grouping, - a step of determining, by the calculation device, two queries for a large language model, as a function of the first and second partitioned groupings and the extracted metadata associated with said groupings, each query being representative of a query for generating textual content, - a step of providing, by the computing device, the two determined requests to a large trained language model, - a step of receiving, by the computing device, the first and second textual content generated and - a step of providing, via a computer interface, the first and second textual content generated.

[0012] Thanks to these provisions, a corpus relevant to a query is extracted and processed so as to bring out two hierarchical, distinct and complementary groupings, representative of the influential files in the cognitive constitution of the extracted corpus from the extracted corpus and of the main themes from the extracted corpus. These corpora can, due to the distinct filtration steps, present a number of different files and, due to the distinct hierarchization steps, order the relevance of the files in a distinct manner within each corpus.

[0013] Thus, each grouping is provided to a user or recorded in a computer memory.

[0014] In addition, the groupings are filtered and sorted statistically, by implementing a large language model configured to perform this filtering and sorting without intervention from a user.

[0015] Thus implemented, such a large language model cannot undergo any hallucination effect with regard to the identified files, because these are identified beforehand by the implementation of an ordinary search engine coupled with the implementation of a statistical and bibliometric approach to the corpus obtained in response to the implementation of the search report.

[0016] Furthermore, the mechanism implemented by the present invention can be said to be a "guided generative transformer" (GGT). That is to say, the mechanism implemented by the present invention allows to transform an unordered set of text metadata with a bibliography, into two partitioned sets from which to generate labeled textual content summarizing the texts, with their supporting sources.

[0017] Furthermore, implemented in this way, such a large language model requires few computing and energy resources to carry out the filtering and sorting step.

[0018] Furthermore, such an invention may be based only on metadata, thus avoiding any infringement of copyright and other intellectual property rights. related.

[0019] In particular embodiments, at least one query for a large language model is representative of a query for labeling at least one partition of at least one grouping.

[0020] In particular embodiments, at least one query for a large language model is representative of a query for generating a summary of at least one partition of at least one grouping.

[0021] In particular embodiments, at least one query for a large language model is representative of a subpartition identification query for at least one partition of at least one cluster.

[0022] In particular embodiments, the method which is the subject of the present invention comprises a step of determining a sub-partition of at least one partition of at least one grouping.

[0023] In particular embodiments, at least one query for a large language model is representative of: - a request for labeling at least one sub-partition of at least one grouping and / or - a request to generate a summary of at least one sub-partition of at least one grouping.

[0024] In particular embodiments, the method which is the subject of the present invention comprises: - a selection step, via an input interface, of an indicator representative of an application discipline and - upstream of the step of providing the determined requests to a large trained language model, a step of selecting a large trained language model, from among a plurality of large trained language models, as a function of the indicator representative of a selected application discipline.

[0025] These embodiments significantly improve the relevance of the results of large language models, trained specifically for a plurality of disciplines.

[0026] In particular embodiments, the statistical training step includes a step of forming, by a calculation device, a first grouping partitioned with classified and hierarchical file identifiers, which includes: - a step of initializing a first threshold value, by a calculation device, representative of a minimum level of relevance of a cited file associated with the defined computer query, - a step of automatic adjustment of the first threshold value, by a calculation device, as a function of numerical characteristics representative of the extracted corpus, - a step of filtration, by a calculation device, of the first set as a function of the first adjusted threshold value, - a step of determining, by a calculation device, a co-citation index for at least one pair of identifiers of cited files from the first set, filtered, as a function of a number of co-referencings of said pair of identifiers of cited files in the second set, - a step of reducing the first filtered set, - a step of partitioning, by a calculation device, the first set according to at least one determined co-citation index, - a step of determining, by a calculation device, at least one commonality index for at least one cited file identifier of the first partitioned set and - a step of hierarchization, by a calculation device, of identifiers of files cited from the first partitioned set, to form the first grouping.

[0027] In particular embodiments, the statistical training step comprises a step of training, by a calculation device, a second partitioned grouping of classified and hierarchical citing file identifiers, which comprises: - a step of initializing a second threshold value, by a calculation device, representative of a minimum level of relevance of a citing file associated with the defined computer query, - a step of dynamic adjustment of the second threshold value, by a calculation device, according to numerical characteristics representative of the extracted corpus, - a filtration step, by a calculation device, of the second set as a function of the second adjusted threshold value, - a step of determining, by a calculation device, an index of files cited in common representative of a number of files cited in common for at least one pair of identifiers of citing files of the second filtered set, - a step of reduction of the second filtered set, - a step of partitioning, by a calculation device, the second reduced set according to at least one index of files cited in common determined, - a step of determining, by a calculation device, at least one index of communalist for at least one citing file identifier from the second partitioned set and - a step of hierarchization, by a calculation device, of identifiers of citing files of the second set, to form the second grouping.

[0028] In particular embodiments, the initialization step of the step of forming the second grouping comprises a step of defining, by a calculation device, a default value for at least one criterion among: - a publication date of the citing file, - a standardized number of citations from the citing file, - a minimum number of file identifiers, - a maximum number of file identifiers and / or - a total number of citing file identifiers, said number being greater than or equal to a determined threshold value and less than or equal to a determined ceiling value, at least one said criterion being implemented during the adjustment step and / or the filtration step of the step of forming the second reduced set.

[0029] These embodiments make it possible to achieve optimal filtration of the results within the second reduced set, with regard to the processing carried out during the step of forming the second reduced set.

[0030] In particular embodiments, the adjustment step of the second grouping formation step is configured to: - keep in the analysis only the citing files published after a specific date, - increase the minimum required normalized citation value to reduce the number of citing files to be retained if the total number of retained citing file identifiers exceeds a maximum number of files to be retained, and - reduce the minimum required normalized citation value to reduce the number of citing files to retain if the total number of citing file identifiers retained is less than the minimum number of files to retain.

[0031] These embodiments make it possible to carry out dynamic filtering of the results obtained according to characteristics specific to the second reduced set extracted.

[0032] According to a second aspect, the present invention relates to a device for automatically generating textual content guided by an expert system for applying bibliometric techniques from a single corpus of relational files recorded in a database, each file of the single corpus being associated with at least one keyword, which comprises a computer memory, storing computer instructions, and a calculation device, which, when this calculation device executes the instructions stored in the computer memory, executes the following steps: - a step of entering, via an entry interface associated with a calculation device, a set of search keywords, - a step of extraction, by a calculation device, of a corpus extracted from relational files from the unique corpus of files recorded in the database, according to the keywords searched and keywords associated with each file of the unique corpus, and a set of metadata associated with each file of the unique corpus, - a step of statistical formation, by the calculation device, of two complementary groupings of distinct files, according to the extracted corpus and the metadata associated with each file of the extracted corpus, to form: - a first partitioned grouping of identifiers of classified and hierarchical cited files, - a second partitioned grouping of classified and hierarchical citing file identifiers, - a step of extracting the metadata of each grouping, - a step of determining, by the calculation device, two queries for a large language model, as a function of the first and second partitioned groupings and the extracted metadata associated with said groupings, each query being representative of a query for generating textual content, - a step of providing, by the computing device, the two determined requests to a large trained language model, - a step of receiving, by the computing device, the first and second textual content generated and - a step of providing, via a computer interface, the first and second textual content generated.

[0033] The advantages obtained by implementing the device which is the subject of the present invention are similar to the advantages obtained by implementing the method which is the subject of the present invention. Brief description of the figures

[0034] Other advantages, aims and particular characteristics of the invention will emerge from the following non-limiting description of at least one particular embodiment of the method and device which are the subject of the present invention, with reference to the appended drawings, in which:

[0035] [Fig-1] schematically represents a particular succession of steps of the process object of the present invention,

[0036] [Fig.2] schematically represents a particular embodiment of a device which is the subject of the present invention,

[0037] [Fig.3] schematically represents a graphical representation of a first partial result of the method which is the subject of the present invention,

[0038] [Fig.4] schematically represents a graphic representation of a second partial result of the method which is the subject of the present invention,

[0039] [Fig.5] represents, schematically, a succession of states of a corpus of documents from an initial extracted corpus,

[0040] [Fig.6] schematically represents a particular succession of steps of the process object of the present invention and

[0041] [Fig.7] schematically represents a particular embodiment of an ar computer architecture capable of carrying out the method which is the subject of the present invention. Description of the embodiments

[0042] The present description is given without limitation, each characteristic of an embodiment being able to be combined with any other characteristic of any other embodiment in an advantageous manner.

[0043] It should be noted from now on that the figures are not to scale.

[0044] As understood from reading this description, various concepts The inventive methods may be implemented by one or more methods or devices described below, several examples of which are provided herein. The actions or steps performed in carrying out the method or device may be ordered in any suitable manner. Accordingly, it is possible to construct embodiments in which the actions or steps are performed in a different order than illustrated, which may include performing certain acts simultaneously, even if they are shown as sequential acts in the illustrated embodiments.

[0045] The expression "and / or", as used herein, is to be understood to mean "either or both" of the elements so conjoined, i.e., elements which are present conjunctively in some cases and disjunctively in other cases. Multiple elements listed with "and / or" are to be interpreted in the same way, i.e., "one or more" of the elements so conjoined. Other elements may optionally be present, other than the elements specifically identified by the "and / or" clause, whether or not they are related to these specifically identified elements.Thus, by way of non-limiting example, a reference to "A and / or B", when used in conjunction with open language such as "comprising" may refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.

[0046] As used herein in the description, "or" is to be understood inclusively.

[0047] As used herein, the term "at least one," with reference to a list of one or more elements, is to be understood to mean at least one element selected from one or more elements in the list of elements, but not necessarily including at least one of each element specifically listed in the list of elements and not excluding any combination of elements in the list of elements. This definition also allows for the optional presence of elements other than the specifically identified elements in the list of elements to which the term "at least one" refers, whether or not related to those specifically identified elements.Thus, by way of non-limiting example, "at least one of A and B" (or, equivalently, "at least one of A or B", or, equivalently, "at least one of A and / or B") may refer, in one embodiment, to at least one, optionally including more than one, A, without B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, without A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.

[0048] In the description below, all transitional expressions such as "comprising", "including", "carrying", "having", "containing", "involving", "holding", "composed of", and the like, are to be understood as open, i.e., as meaning including but not limited to. Only the transitional expressions "consisting of" and "consisting essentially of" are to be understood as closed or semi-closed transitional expressions, respectively.

[0049] It is noted that the terms “computing device” designate any electronic assembly capable of executing instructions representative of a step of an algorithm or of the method which is the subject of the present invention. Such a computing device is exemplified with regard to [Fig.2].

[0050] It is noted that the terms “input means” or “input interface” designate any means of interaction from an external system or user and a computing device. Such an input means may correspond, for example, to an API (for “Application Programming Interface”) or to a human-machine interface (such as a keyboard and / or a mouse) which may be associated with a graphical user interface (translated from “graphical user interface”, or “GUI”).

[0051] It is noted that the terms “computer interfaces” designate any means of interaction from a computing device to an external system or user. Such a computer interface may correspond, for example, to an API or a human-machine interface (such as a screen) which can be associated with an interface user graphic.

[0052] It is noted that the terms “document database” designate any database or computer memory, structured or not, allowing the execution of a search query for documents (or document identifiers) and the extraction of such documents.

[0053] It is noted that the term "expert system" refers to any computer program designed to simulate the expertise and decision-making of a human specialist. These systems are a specific application of artificial intelligence and rely primarily on domain-specific knowledge to solve complex problems at a level comparable to that of a human expert.

[0054] It is noted that the term "bibliometrics" designates a technical discipline which focuses on the quantitative analysis of scientific and technical literature. Bibliometrics uses statistical and mathematical methods to evaluate and quantify the production and dissemination of research documents, such as journal articles, books, and patents.

[0055] The documents recorded in the document database include, for at least some of the documents, bibliographic metadata such as: - at least one publication number, - at least one author, - at least one title, - at least one abstract, and / or - at least one cited document identifier.

[0056] It is noted that the terms “documents” and “files” are interchangeable and that they both designate content stored in a computer memory comprising at least one element among text, an image, a video or a sound and at least one reference to another document or to at least one other file stored in the computer memory. Such a reference may correspond to an address or to a bijective unit reference.

[0057] An "identifier" is any bijective digital representation of an object, such as a document for example. A document identifier corresponds, for example, to a publication number of the document or to a character string representing, for example, at least the authors and the year of publication.

[0058] The result of any automatic classification is called “partitioning” or “clustering”.

[0059] A "communality index" is a measure of the relative intensity of connectivity of a member of a partition with all the other members of the same partition.

[0060] This commonality index corresponds to the intensity with which each reference cited or each citing document contributes to the unity of meaning of its cluster of belonging.

[0061] The commonality index implemented during training step 120 measures the relative intensity with which each reference of a partition is co-cited with all the other references classified in the same partition.

[0062] The commonality index implemented during training step 160 measures the relative intensity with which each document in a partition shares references cited in its bibliography with all other documents classified in the same partition.

[0063] Such indices of commonality represent a radically different approach from contemporary approaches, based on distances or links between documents.

[0064] Such indices allow, at a cognitive level, users to better understand and absorb information than when implementing contemporary approaches.

[0065] In this regard, the present invention implements a logic of quantification of thought and cerebral pathways, in the sense of cognitive sciences, contributing to an improvement in the cerebral activity of users who implement it, this improvement resulting in better memorization and understanding of the information contained in a corpus of documents resulting from a search by keywords.

[0066] [Fig.l], which is not to scale, shows a schematic view of a particular embodiment of the method 100 which is the subject of the present invention. This automated method 100 for collecting, classifying and hierarchizing a corpus of relational documents into two complementary hierarchical groupings comprises: - a step 105 of defining, by a calculation device, a computer query, - a step 110 of extracting, by a calculation device and from a database computer data, of a corpus represented by at least one identifier representative of a document, called “citing document” and of at least one set of metadata associated with at least one said citing document identifier, at least one said metadata being representative of an identifier of another document, called “cited document”, - a step 111 of forming, by a calculation device, a first set of identifiers of cited documents and a second set of identifiers of citing documents corresponding to the extracted corpus, - a step 115 of cleaning, by a calculation device, at least one identifier of the same document cited in the first set, by the implementation of a character string matching algorithm, to harmonize the identifiers of the same cited document, - a step 116 of harmonization, by a calculation device, of identifiers of documents cited in the second set based on the result of cleaning step 115, then, concomitantly: - a step 120 of forming, by a calculation device, a first partitioned grouping of identifiers of classified and hierarchical cited documents, comprising: - a step 125 of initializing a first threshold value, by a calculation device, representative of a minimum level of relevance of a cited document associated with the defined computer query, - a step 130 of automatic adjustment of the first threshold value, by a calculation device, as a function of numerical characteristics representative of the extracted corpus, - a step 135 of filtration, by a calculation device, of the first set as a function of the first adjusted threshold value, - a step 140 of determining, by a calculation device, a co-citation index for at least one pair of identifiers of cited documents from the first set, filtered, as a function of a number of co-referencings of said pair of identifiers of cited documents in the second set, - a step 141 of reduction of the first filtered set, - a step 145 of partitioning, by a calculation device, the first set according to at least one determined co-citation index, - a step 150 of determining, by a calculation device, at least one commonality index for at least one cited document identifier of the first partitioned set and - a step 155 of prioritizing, by a calculation device, identifiers of cited documents from the first partitioned set, to form the first grouping, - a step 160 of forming, by a calculation device, a second partitioned grouping of identifiers of citing documents, classified and prioritized, comprising: - a step 165 of initializing a second threshold value, by a calculation device, representative of a minimum level of relevance of a citing document associated with the defined computer query, - a step 170 of dynamic adjustment of the second threshold value, by a calculation device, as a function of numerical characteristics representative of the extracted corpus, - a step 175 of filtration, by a calculation device, of the second set as a function of the second adjusted threshold value, - a step 180 of determining, by a calculation device, an index of documents cited in common representative of a number of documents cited in common for at least one pair of identifiers of documents citing the second filtered set, - a step 181 of reducing the second filtered set, - a step 185 of partitioning, by a calculation device, the second set reduced according to at least one index of documents cited in common determined, - a step 190 of determining, by a calculation device, at least one index of commonality for at least one identifier of a citing document of the second partitioned set and - a step 195 of hierarchization, by a calculation device, of identifiers of documents citing the second set, to form the second grouping and - a step 200 of supplying, on a digital interface, the first and second hierarchical groupings.

[0067] During the definition step 105, an input means may be implemented to formulate a query consisting of Boolean expressions and / or keywords to be searched in a document database. Such a query may be formulated, for example, using a keyboard interacting with an input field of a graphical interface, validation of the query thus entered resulting in the execution of the search in the document database. Such a definition step 105 is carried out, for example, on a personal computer connected, by a means of communication (such as the Internet, for example) to a computer server responsible for the execution of search queries in the document database. Such a definition step 105 may be carried out in a heavy client of a computer program organized according to a client-server architecture.

[0068] During the extraction step 110, the defined query is executed by a computing device interacting with a document database. Such an extraction step 110 is common in the field of search engines and is not repeated here.

[0069] At the end of this extraction step 110, a documentary corpus is obtained. Such a documentary corpus is formed of at least one citing document identifier associated with at least one set of metadata, at least one such set of metadata comprising at least one document identifier cited by the citing document associated with the set of metadata. Such a corpus may comprise the citing and / or cited documents or only identifiers of citing and / or cited documents, as well as metadata associated with these citing documents, such metadata being able to include a link representative of an address on a computer network of said document.

[0070] For example, a document identifier "A" may be associated with a metadata set that includes, as a cited document identifier, a document identifier "B". We then say that A cites B, that is, B is a document cited for A and A is a document citing for B. The terms "cited document" and "cited reference" are, for all practical purposes, interchangeable.

[0071] During the training step 111, for example, computer processing is carried out to collect at least one document identifier cited in at least one set of metadata associated with at least one citing document identifier from the extracted corpus. Preferably, each cited document identifier from the metadata set associated with each citing document is collected.

[0072] This list of cited document identifiers forms the first set.

[0073] During training step 111, for example, computer processing is carried out to collect at least one citing document identifier from the extracted corpus.

[0074] This list of citing document identifiers forms the second set.

[0075] Alternatively, no collection is carried out, the second set being formed from the result of the extraction step 110.

[0076] During the cleaning step 115, a word processing algorithm may, for example, be implemented. Such an algorithm may have the purpose of removing unnecessary characters from document identifiers, or of reorganizing elements of a character string representative of a document identifier, with a view to bringing such document strings together. By bringing together, here is meant the implementation of a character string matching algorithm, configured to produce a proximity index between several character strings. Such an algorithm makes it possible, when a proximity index is greater than a determined threshold value, to determine that two distinct character strings are representative of the same cited document. This makes it possible, in particular, to constitute correspondence tables between character strings considered as synonyms or even to replace one character string with another.

[0077] Upstream of the cleaning step 115, the method 100 may comprise a pre-cleaning step configured to remove from the extracted corpus the excess occurrences of identifiers of the same cited document. During this pre-cleaning step, the input errors are corrected, in particular by eliminating excess space characters, special characters, unnecessary or redundant characters for example.

[0078] During the harmonization step 116, for example, computer processing is carried out in order to correct, in the second set, the identifiers of erroneous cited documents, that is to say the identifiers of cited documents incorrectly representing these cited documents.

[0079] During such a harmonization step 116, a cited document identifier associated with a citing document identifier may be replaced by a cited document identifier identified, during step 115, as correct (or majority).

[0080] In variants, an ad-hoc cited document identifier is defined for a plurality of character strings considered to be close and implemented to replace each iteration of the character strings belonging to this plurality.

[0081] In variants, an identifier of an entry in a lookup table is defined for a plurality of character strings considered to be close and implemented to replace each iteration of the character strings belonging to this plurality.

[0082] Once the first and second sets have been cleaned, two steps are carried out concomitantly, or in parallel: - step 120 of forming a first partitioned grouping of identifiers of classified and hierarchical cited documents and - step 160 of forming a second partitioned grouping of classified and hierarchical citing document identifiers.

[0083] During step 120 of forming a first partitioned grouping of identifiers of classified and hierarchical cited documents, the objective is to hierarchize the methodologies and theoretical foundations linked to the subject of the query and therefore to the extracted corpus.

[0084] During the initialization step 125, a first threshold value is set, either automatically based on a default initialization instruction, or by a user via a human-machine interface, or by a third-party computer program via an API for example.

[0085] This initialization step 125 may optionally include a step 225 of definition by a calculation device of a default value for at least one criterion among: - a minimum percentage of the number of document identifiers citing the extracted corpus to be retained, - a minimum number of identifiers of cited documents, - a maximum number of cited document identifiers and / or - a total number of identifiers of cited documents associated with a number of citing documents, said number being greater than or equal to a determined threshold value, at least one said criterion being implemented during the adjustment step 130 and / or the filtration step 135 of the first grouping formation step 120.

[0086] The particular values for these criteria depend on the particular implementation use case.

[0087] During the automatic adjustment step 130, the first threshold value is adjusted so as to provide a sample of the extracted corpus, corresponding to the first set, sufficient to allow efficient implementation of the step 140 of determining a co-citation index.

[0088] For example, this automatic adjustment step 130 is configured to modify the value of the minimum percentage of the number of document identifiers to be retained if no result validates a rule associated with at least one criterion implemented during the initialization step 125.

[0089] For example, at least one of the following rules may be implemented: - the total number of documents must be greater than the minimum percentage of the number of document identifiers in the extracted corpus to be retained multiplied by the number of documents in the extracted corpus, - the number of documents must be greater than the minimum number of documents defined and / or - the number of cited documents must be greater than the total number of document identifiers associated with a number of cited documents greater than or equal to a determined threshold value.

[0090] If no such rule can be applied, the minimum percentage value of the number of document identifiers is reduced incrementally until all rules can be applied.

[0091] The automatic adjustment step 130 may also not take place if the characteristics of the first grouping allow it.

[0092] Optionally, the method 100 which is the subject of the present invention may comprise a step (not shown) of readjusting the first initialized or adjusted threshold value. Such a readjustment step may implement a graphical input interface associated with an input means allowing a user to manually readjust the first threshold value.

[0093] The first adjusted threshold value is implemented during the filtration step 135, carried out for example by selecting a number of documents corresponding to the total number of documents obtained at the end of the extraction step 110, from the document identifier associated with the largest number of identifiers of cited documents, and this in a decreasing manner as a function of the number of identifiers of cited documents, and this until the total number of documents is reached.

[0094] During the determination step 140, a matrix associating each pair of cited document identifiers is formed and filled according to the number of times that this pair of cited document identifiers is identified in the metadata associated with the citing documents. This determination step 140 is therefore carried out by counting. Such an algorithm corresponds to a reference cited co-citation analysis, or “RCCA” (for “Referenced Co-Citation Analysis”).

[0095] Downstream of the determination step 140, a step (not shown) of normalizing the results obtained can be carried out. This determination step 140 can be carried out during the partitioning step 145.

[0096] The reduction step 141 reduces, for example, the reference co-citation matrix obtained in step 140 over all cited references into a sub-matrix whose rows and columns are composed of the references identified in the automatic adjustment step 130.

[0097]

[0098]

[0099]

[0100]

[0101]

[0102]

[0103] During partitioning step 145, partitions of identifiers of cited documents are formed and correspond to main themes in the first grouping. Any partitioning ("clustering") or classi fication can be implemented during this partitioning step 145. Such a algorithm can correspond, for example, to a k-means algorithm, fuzzy par- partitioning, pyramidal partitioning or hierarchical classification. During step 150 of determining at least one communality index, the following equation can be implemented, for example and without limitation: nk[Ck(R,) *Sk(Rj)] [Ck(R;) * Sk(RJ ] In which: - nk corresponds to the number of document identifiers cited in partition k, - Hk(Ri)k corresponds to the commonality index of a document identifier cited in partition k, - Ck(Ri) corresponds to the connectivity of the cited document identifier R; in partition k, i.e. the number of cited document identifiers in partition k with which the cited document identifier R; is connected and - Sk(Ri) corresponds to the connectivity strength of the cited document identifier R; in partition k, i.e. the average of the number of cited document identifiers in partition k with which the cited document identifier R; is connected. The average number of identifiers cited can be determined by the following equation, for example and without limitation: St(Ri) = In which: - nk corresponds to the number of document identifiers cited in partition k, - Sk(Ri) corresponds to the connectivity strength of the cited document identifier R; in partition k and - ay corresponds to the number of co-citations of the identifiers of cited documents R; and Rj by documents in the corpus. During the prioritization step 155, the cited documents forming the first grouping are classified among themselves globally (i.e. without taking into account the partitions) and / or within a given partition. In other words, this prioritization step 155 can be carried out according to a score associated with each cited document identifier, such a score being representative of the importance of the cited document in the first grouping. The result of the prioritization step 155 is implemented during the provisioning step 200 described below.

[0104] During initialization step 165, a second threshold value is initialized manually or automatically, similarly to initialization step 125.

[0105] In embodiments, step 165 of initializing step 160 of forming the second reduced set comprises a step 230 of defining, by a calculation device, a default value for at least one criterion among: - a publication date of the citing document, - a standardized number of citations of the citing document, - a minimum number of citing document identifiers, - a maximum number of citing document identifiers and / or - a total number of citing document identifiers, said number being greater than or equal to a determined threshold value and less than or equal to a determined ceiling value, at least one said criterion being implemented during the adjustment step 170 and / or the filtration step 175 of the step of forming the second reduced set.

[0106] These embodiments make it possible to achieve optimal filtration of the results within the second reduced set, with regard to the processing carried out during the step of forming the second reduced set.

[0107] In embodiments, the adjustment step 170 of the step 160 of forming the second reduced set is configured to: - keep in the analysis only citing documents published after a specific date, - increase the minimum standardized citation value required to reduce the number of citing documents to be retained if the total number of citing document identifiers retained exceeds a maximum number of documents to be retained, and - reduce the minimum standardized citation value required to reduce the number of citing documents to be retained if the total number of citing document identifiers retained is less than the minimum number of documents to be retained.

[0108] In particular embodiments, the adjustment step 170 of the step 160 of forming the second reduced set is configured to: - consider in the analysis only the citing documents published after a specific date which is set by default as A - 5 where A is the year of use of the solution by the user and, - increase the minimum standardized citation value required to reduce the number of citing documents to be retained if the total number of citing document identifiers retained by the application of the first criterion exceeds the maximum number of documents to be retained, the increase in the maximum value having no limit, and - reduce the minimum required normalized citation value to reduce the number of citing documents to be retained if the total number of citing document identifiers retained by the application of the first criterion is less than the minimum number of documents to be retained, the minimum value not being able to fall below zero.

[0109] These embodiments make it possible to perform dynamic filtering of the results obtained according to characteristics specific to the second reduced set extracted.

[0110] Definition step 230 is performed manually or automatically, similarly to definition step 225.

[0111] The dynamic adjustment step 170 is performed in a similar manner to the adjustment step 130, for example.

[0112] In particular embodiments, step 170 of adjusting step 160 of forming the second grouping is configured to: - reduce the dynamic value if the total number of document identifiers associated with a standardized number of cited documents is less than the minimum number of document identifiers and - increase the dynamic value if the total number of document IDs associated with a standardized number of cited documents is greater than the maximum number of document IDs.

[0113] Optionally, the method 100 which is the subject of the present invention may comprise a step (not shown) of readjusting the second initialized or adjusted threshold value. Such a readjustment step may implement a graphical input interface associated with an input means allowing a user to manually readjust the second threshold value.

[0114] The filtration step 175 is carried out, for example, in a similar manner to the filtration step 135, depending on the second threshold value adjusted during the adjustment step 170 or not.

[0115] At the end of this filtration step 175, a second filtered set is obtained.

[0116] The determination step 180 is carried out, for example, by establishing a matrix representative of each pair of citing document identifiers in the second set and by counting, for each pair of citing document identifiers, the number of document identifiers cited in common. Such an algorithm corresponds to a bibliographic coupling analysis of citing documents, or “DBCA” (for “Document Bibliography Coupling Analysis”).

[0117] The reduction step 181 reduces, for example, the bibliographic document coupling matrix obtained in step 180 over the set of citing documents into a sub-matrix whose rows and columns are composed of the citing documents identified in the automatic adjustment step 170.

[0118] The partitioning step 185 is carried out, for example, by implementing a suitable partitioning or classification algorithm. Such an algorithm may cor respond, for example, to a k-means, fuzzy partitioning, pyramid partitioning or hierarchical classification algorithm.

[0119] Step 190 of determining at least one communality index is carried out, for example, by implementing the following equation: h / dU 7 ^(D^CSADJ

[0120] In which: - nk corresponds to the number of document identifiers cited in partition k, - Hk(Di) corresponds to the commonality index of a document identifier in a partition k, - Ck(Di) corresponds to the connectivity of the document identifier D; in partition k, i.e. the number of document identifiers in partition k with which the cited document identifier D; is connected and - Sk(Di) corresponds to the connectivity strength of the document identifier D; in partition k, i.e. the average of the number of document identifiers in partition k with which the cited document identifier D; is connected.

[0121] The connectivity strength of the document identifier D; in the partition k can be determined by implementing the following equation: cjdd =

[0122] In which: - nk corresponds to the number of document identifiers cited in partition k, - Ck(Di) corresponds to the connectivity strength of the document identifier D; in partition k and - Rÿ corresponds to the number of cited document identifiers shared by document identifiers D; and Dj.

[0123] The prioritization step 195 is performed, for example, in a similar manner to the prioritization step 155.

[0124] The second hierarchical grouping is implemented during the provisioning step 200.

[0125] The provisioning step 200 is carried out, for example, by implementing a graphical user interface adapted to the display of the first and second hierarchical groupings.

[0126] Two results of this supply step 200 are visible in figures 3 to 5: - we observe, in [Fig.3], a graphic representation of partitions, 401 and 402, of the first set, each partition, 401 and 402, being associated with cited documents, 403 and 404, each cited document being associated with a particular highlighting, representative of the importance of the cited document in the first set (size of the circle) and / or in the score (intensity of the color of the link) and - we observe, in [Fig.4], a graphic representation of partitions, 501 and 502, of the first set, each partition, 501 and 502, being associated with citing documents, 503 and 504, each citing document being associated with a particular highlighting, representative of the importance of the citing document in the second set (size of the circle) and / or in the partition (intensity of the color of the circle).

[0127] We observe, in [Fig.5], a succession of processing states of an initial extracted corpus of 600, among which: - three citing documents, 605, 610 and 615 are associated with cited documents, 605', 610' and 615', in the extracted corpus 600, - several cited documents, 605', 610' and 615', are associated with a citing document, 605, 610 and 615, - if two or more cited documents are identical (for example those marked by a black square in the top left corner), only one such cited document is retained to avoid duplicates and - if two or more cited documents are different but representative of the same content (for example, the one marked by a hatched square), which results in a slight difference in the character strings identifying these documents, one of the two or more documents is corrected to correspond to a unique character string.

[0128] These operations result in the constitution of the first cleaned set, the result of which is used to obtain the second harmonized set, in which the citing documents, 605”, 610” and 615”, are associated with corrected cited documents, avoiding the risks of duplicates and confusion as to the uniqueness of these cited documents.

[0129] As understood, the steps of: - training 111, - cleaning 115, - harmonization 116, - formation of a first group 120 and - formation of a second group 160, can be arbitrarily grouped in a step 715 of statistical formation of two complementary groupings of distinct files to form: - a first partitioned grouping of identifiers of classified and hierarchical cited files and - a second partitioned grouping of classified and hierarchical citing file identifiers.

[0130] In [Fig.2], we observe schematically a particular embodiment of a device 300 or computer system capable of implementing the method 100. object of the present invention. This automated device 300 for collecting, classifying and hierarchizing a corpus of relational documents into two complementary hierarchical groupings, comprises a computer memory 325, storing computer instructions, and a calculation device 310, which, when this calculation device executes the instructions stored in the computer memory, executes the following steps: - a step of defining, by a calculation device, a computer request, - a step of extraction, by a calculation device and from a computer database, of a corpus represented by at least one identifier representative of a document, called "citing document" and of at least one set of metadata associated with at least one said citing document identifier, at least one said metadata being representative of an identifier of another document, called "cited document", - a step of forming, by a calculation device, a first set of identifiers of cited documents and a second set of identifiers of citing documents corresponding to the extracted corpus, - a cleaning step, by a calculation device, of at least one identifier of the same document cited in the first set, by the implementation of a character string matching algorithm, to harmonize the identifiers of the same cited document, - a step of harmonization, by a calculation device, of identifiers of documents cited in the second set according to the result of the cleaning step,

[0131] then, concomitantly: - a step of forming, by a calculation device, a first partitioned grouping of identifiers of classified and hierarchical cited documents, comprising: - a step of initializing a first threshold value, by a calculation device, representative of a minimum level of relevance of a cited document associated with the defined computer query, - a step of automatic adjustment of the first threshold value, by a calculation device, according to numerical characteristics representative of the extracted corpus, - a filtration step, by a calculation device, of the first set as a function of the first adjusted threshold value, - a step of determining, by a calculation device, a co-citation index for at least one pair of identifiers of cited documents from the first set, filtered, as a function of a number of co-referencings of said pair of identifiers of cited documents in the second set, - a step of reducing the co-citation matrix of references adjusted according to the identifiers of cited references making up the first filtered set, - a partitioning step, by a calculation device, of the first set into function of at least one determined co-citation index, - a step of determining, by a calculation device, at least one commonality index for at least one cited document identifier of the first partitioned set and - a step of prioritizing, by a calculation device, identifiers of cited documents from the first partitioned set, to form the first grouping, - a step of forming, by a calculation device, a second partitioned grouping of identifiers of classified and prioritized citing documents, comprising: - a step of initializing a second threshold value, by a calculation device, representative of a minimum level of relevance of a citing document associated with the defined computer query, - a step of dynamic adjustment of the second threshold value, by a calculation device, as a function of numerical characteristics representative of the extracted corpus, - a step of filtration, by a calculation device, of the second set as a function of the adjusted second threshold value, - a step of determining, by a calculation device, an index of documents cited in common representative of a number of documents cited in common for at least one pair of identifiers of documents citing the second filtered set, - a step of reducing the bibliographic coupling matrix of documents adjusted according to the identifiers of documents making up the second filtered set, - a step of partitioning, by a calculation device, the second reduced set according to at least one index of documents cited in common determined, - a step of determining, by a calculation device, at least one index of commonality for at least one identifier of a citing document of the second partitioned set and - a step of hierarchization, by a calculation device, of identifiers of documents citing the second set, to form the second grouping and - a step of providing, on a digital interface, the first and second hierarchical groupings.

[0132] Shown in [Fig. 2], which is not to scale, is a block diagram illustrating an exemplary computer system with which an embodiment of a method of the present invention may be implemented. In the example of [Fig. 2], a computer system 305 and instructions for implementing the disclosed technologies in hardware, software, or a combination of hardware and software are shown schematically, for example, as boxes and circles, at the same level of detail that is commonly used by those of ordinary skill in the art to which this disclosure relates to communicate on computer architecture and computer system implementations.

[0133] The computer system 305 includes an input / output (I / O) subsystem 320 that may include a bus and / or one or more other communication mechanisms for communicating information and / or instructions between components of the computer system 305 over electronic signal paths. The input / output subsystem 320 may include an input / output controller, a memory controller, and at least one input / output port. The electronic signal paths are shown schematically in the drawings, for example, as lines, one-way arrows, or two-way arrows.

[0134] At least one processor 310, or computing device, is coupled to the LO subsystem 320 to process information and instructions. The processor 310 may include, for example, a general-purpose microprocessor or microcontroller and / or a special-purpose microprocessor such as an embedded system or graphics processing unit (GPU) or a digital signal processor or an ARM processor. The processor 310 may include an integrated arithmetic logic unit (ALU) or may be coupled to a separate ALU.

[0135] The computer system 305 includes one or more memories 325, such as a main memory, which is coupled to the I / O subsystem 320 to electronically digitally store data and instructions to be executed by the processor 310. The memory 325 may include volatile memory such as various forms of random access memory (RAM) or any other dynamic storage device. The memory 325 may also be used to store temporary variables or other intermediate information during the execution of the instructions to be executed by the processor 310. Such instructions, when stored in a non-transitory computer-readable storage medium accessible to the processor 310, may transform the computer system 305 into a special purpose machine that is customized to perform the operations specified in the instructions.

[0136] The computer system 305 further includes non-volatile memory such as a read-only memory (ROM) 330 or other static storage device coupled to the RO subsystem 320 for storing information and instructions for the processor 310. The ROM 330 may include various forms of programmable ROM (PROM) such as erasable PROM (EPROM) or electrically erasable PROM (EEPROM). A persistent storage unit 315 may include various forms of non-volatile random access memory (NVRAM), such as FLASH memory, or solid state storage, a magnetic disk, or an optical disk such as a CD-ROM or DVD-ROM and may be coupled to the FO subsystem 320 for storing information and instructions for the processor 310. formations and instructions. Memory 315 is an example of a non-transitory computer-readable medium that can be used to store instructions and data that, when executed by processor 310, cause execution of computer-implemented methods for performing the techniques of this document.

[0137] The instructions in memory 325, ROM 330, or storage 315 may comprise one or more sets of instructions that are organized into modules, methods, objects, functions, routines, or calls. The instructions may be organized as one or more computer programs, operating system services, or application programs, including mobile applications. The instructions may comprise an operating system and / or system software; one or more libraries to support multimedia, programming, or other functions; data protocol instructions or stacks to implement TCP / IP, HTTP, or other communication protocols; file format processing instructions to parse or render files encoded using HTML, XML, JPEG, MPEG, or PNG;user interface instructions for rendering or interpreting commands for a graphical user interface (GUI), command-line interface, or text-based user interface;application software such as an office suite, Internet access applications, design and manufacturing applications, graphics applications, audio applications, software engineering applications, educational applications, games, or miscellaneous applications. The instructions may implement a web server, a web application server, or a web client. The instructions may be organized as a presentation layer, an application layer, and a data storage layer such as a relational database system using Structured Query Language (SQL) or no SQL, an object store, a graph database, a flat file system, or other data storage. ;

[0138] The computer system 305 may be coupled via the I / O subsystem 320 to at least one output device 335. In one embodiment, the output device 335 is a digital computer display. Examples of displays that may be used in various embodiments include a touch screen or a light-emitting diode (LED) display or a liquid crystal display (LCD) or an e-paper display. The computer system 305 may include one or more other types of output devices 335, either as a replacement for or in addition to a display device. Examples of other output devices 335 include printers, ticket printers, plotters, projectors, sound cards, or video cards, speakers, buzzers or piezoelectric or other audible devices, LED or LCD lamps or indicators, haptic devices, actuators or servos.

[0139] At least one input device 340 is coupled to the LO subsystem 320 to communicate signals, data, command selections, or gestures to the processor 310. Examples of input devices 340 include touchscreens, microphones, digital still and video cameras, alphanumeric and other keys, keyboards, graphics tablets, image scanners, joysticks, clocks, switches, buttons, dials, sliders.

[0140] Another type of input device is a control device 345, which may perform cursor control or other automated control functions such as navigating a graphical interface on a display screen, alternatively or in addition to input functions. The control device 345 may be a touchpad, mouse, trackball, or cursor direction keys to communicate direction information and control selections to the processor 310 and to control movement of the cursor on the display 335. The input device may have at least two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), which allows the device to specify positions in a plane.Another type of input device is a wired, wireless, or optical control device, such as a joystick, wand, console, steering wheel, pedal, gear shift mechanism, or other type of control device. An input device 340 may include a combination of several different input devices, such as a video camera and a depth sensor.

[0141] In another embodiment, the computer system 305 may include an Internet of Things (IoT) device in which one or more of the output device 335, the input device 340, and the control device 345 are omitted. Or, in such an embodiment, the input device 340 may include one or more cameras, motion detectors, thermometers, microphones, seismic detectors, other sensors or detectors, measuring devices, or encoders, and the output device 335 may include a special-purpose display such as a single-line LED or LCD display, one or more indicators, a display panel, a meter, a valve, a solenoid, an actuator, or a servomotor.

[0142] The output device 335 may include hardware, software, firmware, and interfaces to generate position report packets, notifications, pulse or heartbeat signals, or other recurring data transmissions that specify a position of the computer system 305, alone or in combination with other application-specific data, directed to host 350 or server 355.

[0143] The computer system 305 may implement the techniques described herein using custom hardwired logic, at least one ASIC (Application-Specific Integrated Circuit) or FPGA (Field-Programmable Gate Array), firmware, and / or program instructions or logic that, when loaded and used or executed in combination with the computer system, cause or program the computer system to operate as a special-purpose machine. In one embodiment, the techniques described herein are executed by the computer system 305 in response to the processor 310 executing at least one sequence of at least one instruction contained in the main memory 325.These instructions may be read into main memory 325 from another storage medium, such as memory 315. Execution of the instruction sequences contained in main memory 325 causes processor 310 to execute the process steps described herein. In other embodiments, hard-wired circuits may be used instead of or in combination with software instructions.

[0144] The term "storage medium," as used herein, means any non-transitory medium that stores data and / or instructions that enable a machine to operate in a specific manner. These storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as memory 315. Volatile media include dynamic memory, such as memory 325. Common forms of storage media include, for example, a hard disk drive, a solid-state drive, a flash drive, a magnetic data storage medium, any optical or physical data storage medium, a memory chip, etc.

[0145] Storage media are distinct from, but may be used in conjunction with, transmission media. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including the wires that constitute a bus of the I / O subsystem 320. Transmission media may also take the form of acoustic or light waves, such as those generated during radio and infrared data communications.

[0146] Various forms of media may be involved in the transport of at least one sequence of at least one instruction to the processor 310 for execution. For example, the instructions may initially be carried on a magnetic disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a communications link such as a fiber optic or coaxial cable or a telephone line using a modem. A modem or router local to the computer system 305 may receive the data over the communications link and convert the data into a format that can be read by the computer system 305. For example, a receiver such as a radio frequency antenna or an infrared detector may receive the data carried in a wireless or optical signal and appropriate circuitry may provide the data to the I / O subsystem 320, for example by placing the data on a bus.The I / O subsystem 320 transports data to memory 325, from which the processor 310 retrieves and executes instructions. Instructions received by memory 325 may optionally be stored on memory 315 before or after execution by the processor 310.

[0147] The computer system 305 also includes a communication interface 360 coupled to a bus 320. The communication interface 360 provides a bidirectional data communication coupling to the one or more network links 365 that are directly or indirectly connected to at least one communication network, such as a network 370 or a public or private cloud on the Internet. For example, the communication interface 360 may be an Ethernet network interface, an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for providing a data communication connection to a corresponding type of communication line, for example, an Ethernet cable or a metallic cable of any type or a fiber optic line or a telephone line.Network 370 broadly represents a local area network (LAN), a wide area network (WAN), a campus network, an Internet network, or any combination thereof. Communication interface 360 may include a LAN card to provide a data communication connection to a compatible LAN, or a cellular radiotelephone interface that is wired to send or receive cellular data according to cellular radiotelephone wireless network standards, or a satellite radio interface that is wired to send or receive digital data according to satellite wireless network standards. In any such implementation, communication interface 360 sends and receives electrical, electromagnetic, or optical signals over signal paths that carry digital data streams representing various types of information.

[0148] The network link 365 typically provides electrical data communication, electromagnetic or optical directly or via at least one network to other data devices, using, for example, satellite, cellular, Wi-Fi or BLUETOOTH technology. For example, the network link 365 may provide a connection through a network 370 to a host computer 350.

[0149] Further, network link 365 may provide a connection via network 370 or to other computing devices via interconnecting devices and / or computers that are operated by an Internet Service Provider (ISP) 375. ISP 375 provides data communication services via a global packet data communication network represented by Internet 380. A server computer 355 may be coupled to Internet 380. Server 355 broadly represents any computer, data center, virtual machine or virtual computing instance with or without a hypervisor, or computer running a containerized program system such as DOCKER or KUBERNETES.The server 355 may represent an electronic digital service that is implemented using more than one computer or instance and that is accessed and used by transmitting web service requests, Uniform Resource Locator (URL) strings with parameters in Hypertext Transfer Protocol (HTTP) payloads, application programming interface (API) calls, application service calls, or other service calls. The computer system 305 and the server 355 may form elements of a distributed computing system that includes other computers, a processing partition, a server farm, or other organization of computers that cooperate to perform tasks or run applications or services.The server 355 may include one or more sets of instructions that are organized as modules, methods, objects, functions, routines, or calls. The instructions may be organized as one or more computer programs, operating system services, or application programs, including mobile applications.The instructions may include an operating system and / or system software; one or more libraries to support multimedia, programming, or other functions; instructions or data protocol stacks to implement TCP / IP (for Transmission control protocol / Internet protocol), HTTP, or other communication protocols; file format processing instructions to parse or render files encoded using HTML (for Hypertext markup language), XML (for Extensible markup language), . JPEG (for Joint Photographie Experts Group), MPEG (for Moving Picture Experts Group), or PNG (for Portable Networks Graphics); user interface instructions for rendering or interpreting commands for a graphical user interface (GUI), command-line interface, or text-based user interface; application software such as an office suite, Internet access applications, design and manufacturing applications, graphics applications, audio applications, software engineering applications, educational applications, games, or miscellaneous applications.The server 355 may include a web application server that hosts a presentation layer, an application layer, and a data storage layer such as a relational database system using Structured Query Language (SQL) or no SQL, an object store, a graph database, a flat file system, or other data storage.

[0150] The computer system 305 may send messages and receive data and instructions, including program code, via the network(s), the network link 365, and the communications interface 360. In the Internet example, a server 355 may transmit requested code for an application program via the Internet 380, the ISP 375, the local area network 370, and the communications interface 360.The received code may be executed by the processor 310 as it is received, and / or stored in the memory 315, or in other non-volatile memory for later execution.

[0151] The execution of instructions as described in this section may implement a process as an instance of a currently executing computer program consisting of program code and its current activity. Depending on the operating system (OS), a process may consist of multiple threads that execute instructions concurrently. In this context, a computer program is a passive collection of instructions, while a process may be the actual execution of those instructions. Multiple processes may be associated with the same program; for example, having multiple instances of the same program open often means that more than one process is running. Multitasking may be implemented to allow multiple processes to share the processor 310.Although each processor 310 or processor core executes only one task at a time, the computer system 305 may be programmed to implement multitasking to allow each processor to switch between currently executing tasks without having to wait for each task to complete. In one embodiment, the switches may be performed when . Tasks perform input / output operations, when a task indicates it can be switched, or upon hardware interrupts. Time sharing may be implemented to enable rapid response to user-interactive applications by rapidly performing context switches to give the appearance of simultaneous execution of multiple processes. In one embodiment, for security and reliability reasons, an operating system may prevent direct communication between independent processes, by providing strictly mediated and controlled interprocess communication functionality.

[0152] In [Fig.6], we observe schematically a particular embodiment of the method 700 which is the subject of the present invention. This method 700 for automatic generation of textual content guided by an expert system for applying bibliometric techniques from a single corpus of relational files recorded in a database, each file of the single corpus being associated with at least one keyword, comprises: - a step 705 of entering, via an entry interface associated with a calculation device, a set of search keywords, - a step 710 of extraction, by a calculation device, of a corpus extracted from relational files from the unique corpus of files recorded in the database, according to the keywords searched and keywords associated with each file of the unique corpus, and a set of metadata associated with each file of the unique corpus, - a step 715 of statistical formation, by the calculation device, of two complementary groupings of distinct files, according to the extracted corpus and the metadata associated with each file of the extracted corpus, to form: - a first partitioned grouping of identifiers of classified and hierarchical cited files, - a second partitioned grouping of classified and hierarchical citing file identifiers, - a step 745 of extracting the metadata of each grouping, - a step 720 of determining, by the calculation device, two queries for a large language model, as a function of the first and second partitioned groupings and the extracted metadata associated with said groupings, each query being representative of a query for generating textual content, - a step 725 of supplying, by the calculation device, the two determined requests to a large trained language model, - a step 730 of reception, by the calculation device, of the first and second textual content generated and - a step 735 of supplying, via a computer interface, the first and the second generated text content.

[0153] During the input step 705, a set of keywords, defined by a user or a third-party system, are collected and stored in a temporary computer memory. Such an input step 705 may implement an API or a human-machine interface, such as a keyboard, a mouse and a graphical user interface for example.

[0154] Extraction step 710 may be performed in a similar manner to extraction step 110 as described with respect to [Fig.l].

[0155] The training step 715 is carried out, for example, as described with reference to [Fig.l].

[0156] Step 745 of extracting metadata is carried out, for example, by implementing a computer program executed by a computing device. Such a computer program is for example configured to extract the title, at least one name of an author, a publication journal, a year of publication, and / or at least one cited publication identifier.

[0157] The determination step 720 is carried out, for example, by implementing a computer program executed by a computing device. During this calculation step 720, two distinct queries in natural language are each formed by the concatenation of a standardized query text representative of a content generation instruction.

[0158] Such a query corresponds, for example, to a so-called “hole” query, i.e. a character string comprising variables, representative of the extracted metadata, partitions or sub-partitions and / or a grouping.

[0159] Such a query is then provided, via a graphical interface or a software interface (API type) to a large trained language model.

[0160] In particular embodiments, at least one query for a large language model is representative of a query for labeling at least one partition of at least one grouping.

[0161] In particular embodiments, at least one query for a large language model is representative of a query for generating a summary of at least one partition of at least one grouping.

[0162] In particular embodiments, at least one query for a large language model is representative of a subpartition identification query for at least one partition of at least one grouping.

[0163] In particular embodiments, at least one query for a large language model is representative: - a request for labeling at least one sub-partition of at least one grouping and / or - a request to generate a summary of at least one sub-partition of at least one grouping.

[0164] The provisioning step 725 is carried out, for example, by implementing a computer program executed by a computing device. During this provisioning step 725, an API or a robotic process automation system (RPA) may be implemented, for example. The large language model trained may correspond, for example, to ChatGPT (Registered Trademark) or to a large language model trained specifically on a corpus of a scientific field corresponding to the scientific field of the files sought.

[0165] The reception step 730 is carried out, for example, by implementing a computer program executed by a computing device. During this reception step 730, an API is for example implemented.

[0166] The provisioning step 735 is carried out, for example, by implementing a computer program executed by a computing device. During this provisioning step 735, an API is for example implemented.

[0167] In particular embodiments, the method 700 which is the subject of the present invention comprises a step 750 of determining a sub-partition of at least one partition of at least one grouping.

[0168] In particular embodiments, the method 700 which is the subject of the present invention comprises: - a step 740 of selection, via an input interface, of an indicator representative of an application discipline and - upstream of step 725 of providing the determined requests to a large trained language model, a step 745 of selecting a large trained language model, from among a plurality of large trained language models, as a function of the indicator representative of a selected application discipline.

[0169] The selection step 740 may implement an API or a human-machine interface, such as a keyboard, a mouse and a graphical user interface for example. This selection step 740 may constitute, for example, the selection of an item in a list, each item being representative of a discipline. A discipline may correspond, for example, to a scientific discipline such as medicine or economics.

[0170] The selection step 745 can be carried out manually, then optionally merging with the selection step 740, or automatically, via a set of instructions representative of a computer program for example.

[0171] In variants, each available large trained language model is associated with an attribute (“tag” in English) or numerical indicator representative of a discipline associated with this large trained language model. Thus, the selected large trained language model corresponds to the large trained language model associated with the digital discipline indicator which corresponds to the digital discipline indicator selected during the selection step 740.

[0172] In particular embodiments, the method 700 which is the subject of the present invention comprises at least one step of training a large language model, associated with a digital discipline identifier, on a corpus of relational documents filtered to correspond to documents of the determined discipline.

[0173] [Fig.7] schematically represents a particular embodiment of a computer architecture 800 capable of carrying out a particular embodiment of the method 700 which is the subject of the present invention. This computer architecture 800 comprises: - a set of 805 keywords entered via an input interface, - a text search engine 810 configured to, based on the keywords 805 entered, provide a corpus 815 of relational files or documents, - a means 820 for statistical preprocessing of the corpus 825 of files to configure two means, 830 and 860, for statistical formation of groupings, 835 and 865, the preprocessed corpus, 825 and 855, being supplied to the means, 830 and 860, for statistical formation of groupings, - a means 830 for statistically forming a first grouping 835 of files into a grouping by unit of meaning as a function of the result of the pre-processing means 820 and the pre-processed corpus 825, - a natural language query 840, combined with the first grouping 835 of files or with data representative of the first grouping of files, to form a content generation query to be addressed to a large language model 845, - a large language model 845 configured to produce text content that depends on the query 840, - a means 860 for statistically forming a second grouping 865 of files into a grouping by searching for meaning based on the result of the pre-processing means 820 and the pre-processed corpus 855, - a query 870 in natural language, combined with the second grouping 865 of files or with data representative of the second grouping of files, to form a filtration and / or sorting query to be addressed to a large language model 875, - a large language model 875 configured to produce text content that depends on the query 870.

Claims

Claims

1. Method (700) for automatically generating textual content guided by an expert system for applying bibliometric techniques from a single corpus of relational files recorded in a database, each file of the single corpus being associated with at least one keyword, characterized in that it comprises: - a step (705) of entering, via an entry interface associated with a calculation device, a set of search keywords, - a step (710) of extracting, by a calculation device, a corpus extracted from relational files from the single corpus of files recorded in the database, as a function of the keywords searched for and keywords associated with each file of the single corpus, and a set of metadata associated with each file of the single corpus, - a step (715) of statistical formation, by the calculation device, of two complementary groupings of distinct files,based on the extracted corpus and the metadata associated with each file of the extracted corpus, to form: - a first partitioned grouping of classified and hierarchical cited file identifiers, - a second partitioned grouping of classified and hierarchical citing file identifiers, - a step (745) of extracting the metadata of each grouping, - a step (720) of determining, by the computing device, two queries for a large language model, based on the first and second partitioned groupings and the extracted metadata associated with said groupings, each query being representative of a query for generating textual content, - a step (725) of providing, by the computing device, the two determined queries to a trained large language model, - a step (730) of receiving, by the computing device, the first and second generated textual content and - a step (735) of providing,via a computer interface, of the first and second generated textual content.,

2. The method (700) of claim 1, wherein at least one query for a large language model is representative of a query for labeling at least one partition of at least one cluster.

3. Method (700) according to one of claims 1 or 2, in which at at least one query for large language model is representative of a query to generate a summary of at least one partition of at least one grouping.

4. Method (700) according to one of claims 1 to 3, in which at least one query for large language model is representative of a sub-partition identification query for at least one partition of at least one grouping.

5. Method (700) according to one of claims 1 to 4, which comprises a step (750) of determining a sub-partition of at least one partition of at least one grouping.

6. Method (700) according to one of claims 5 or 6, in which at least one query for a large language model is representative of: - a query for labeling at least one sub-partition of at least one grouping and / or - a query for generating a summary of at least one sub-partition of at least one grouping.

7. Method (700) according to one of claims 1 to 6, which comprises: - a step (740) of selecting, via an input interface, an indicator representative of an application discipline and - upstream of the step (725) of providing the determined requests to a large trained language model, a step (745) of selecting a large trained language model, from among a plurality of large trained language models, as a function of the indicator representative of an application discipline selected.

8. Method (700) according to one of claims 1 to 7, in which the statistical training step comprises a step (120) of forming, by a calculation device, a first partitioned grouping of identifiers of classified and hierarchical cited files, which comprises: - a step (125) of initializing a first threshold value, by a calculation device, representative of a minimum relevance level of a cited file associated with the defined computer query, - a step (130) of automatically adjusting the first threshold value, by a calculation device, as a function of numerical characteristics representative of the extracted corpus, - a step (135) of filtering, by a calculation device, the first set as a function of the first adjusted threshold value, - a step (140) of determining, by a calculation device, a co-citation index for at least one pair of file identifiers

9. cited from the first set, filtered, based on a number of co-references of said pair of file identifiers cited in the second set, - a step (141) of reducing the first filtered set, - a step (145) of partitioning, by a calculation device, the first set according to at least one determined co-citation index, - a step (150) of determining, by a calculation device, at least one commonality index for at least one cited file identifier of the first partitioned set and - a step (155) of hierarchization, by a calculation device, of identifiers of files cited from the first partitioned set, to form the first grouping. Method (700) according to one of claims 1 to 8, in which the statistical training step comprises a step (160) of forming, by a calculation device, a second partitioned grouping of identifiers of classified and hierarchical citing files, which comprises: - a step (165) of initializing a second threshold value, by a calculation device, representative of a minimum relevance level of a citing file associated with the defined computer query, - a step (170) of dynamic adjustment of the second threshold value, by a calculation device, as a function of numerical characteristics representative of the extracted corpus, - a step (175) of filtering, by a calculation device, the second set as a function of the adjusted second threshold value, - a step (180) of determining, by a calculation device,of a commonly cited file index representative of a number of commonly cited files for at least one pair of citing file identifiers of the second filtered set, - a step (181) of reducing the second filtered set, - a step (185) of partitioning, by a calculation device, the second reduced set according to at least one determined commonly cited file index, - a step (190) of determining, by a calculation device, at least one commonality index for at least one citing file identifier of the second partitioned set and - a step (195) of hierarchizing, by a calculation device, citing file identifiers of the second set, to form the, second grouping.

10. Device (300) for automatically generating textual content guided by an expert system for applying bibliometric techniques from a single corpus of relational files recorded in a database, each file of the single corpus being associated with at least one keyword, characterized in that it comprises a computer memory (325), storing computer instructions, and a calculation device (310), which, when this calculation device executes the instructions stored in the computer memory, executes the following steps: - a step of entering, via an entry interface associated with a calculation device, a set of search keywords, - a step of extraction, by a calculation device, of a corpus extracted from relational files from the unique corpus of files recorded in the database, according to the keywords searched and keywords associated with each file of the unique corpus, and a set of metadata associated with each file of the unique corpus, - a step of statistical formation, by the calculation device, of two complementary groupings of distinct files, according to the extracted corpus and the metadata associated with each file of the extracted corpus, to form: - a first partitioned grouping of identifiers of classified and hierarchical cited files, - a second partitioned grouping of classified and hierarchical citing file identifiers, - a step of extracting the metadata from each grouping, - a step of determining, by the computing device, two queries for a large language model, based on the first and second partitioned groupings and the extracted metadata associated with said groupings, each query being representative of a query for generating textual content, - a step of providing, by the computing device, the two determined requests to a large trained language model, - a step of receiving, by the computing device, the first and second textual content generated and - a step of providing, via a computer interface, the first and second textual content generated.