SYSTEM AND METHOD FOR AUTOMATIC CONTENT GENERATION
An automated content generation system addresses the inefficiencies of manual simulation environment development by using a knowledge base and AI model customization to create cyber-influence and information warfare scenarios, improving efficiency and reusability.
Patent Information
- Application Number
- FR2023000442
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-01-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-01-18
AI Technical Summary
Current simulation environments for training in cyber-influence and information warfare are tedious to develop, require significant human intervention, and lack reusability, necessitating a more efficient and automated approach.
An automatic content generation system that utilizes a knowledge base, scheduler, and constraint selector to derive sequence and document-level constraints, customizing an artificial intelligence model to generate content based on scenario-level constraints, reducing the need for human operators.
Facilitates the rapid establishment of cyber-influence and information warfare simulation environments by automating content generation, enhancing efficiency and reusability while minimizing human intervention.
Smart Images

Figure 00000022_0000 
Figure 00000022_0001 
Figure 00000023_0000
Abstract
Description
Title of the invention: SYSTEM AND METHOD FOR AUTOMATIC CONTENT GENERATION Technical field
[0001] The present invention relates to the field of automatic generation of content for the purposes of simulating cyber-influence and information warfare, in particular for the purposes of training or coaching in understanding cyber-influence and information warfare. STATE OF PRIOR ART
[0002] Currently, cybersecurity specialists are creating simulation environments from scratch to train teams to understand cyber influence and information warfare.
[0003] These environments simulate realistic communication activity: actors, communication media, exchanges of messages of different types... with the inclusion of suspicious information so that participants in training sessions can detect them.
[0004] It is possible to create these simulation environments from real situations and data, but for obvious reasons of adaptation to educational needs, current practice involves human operators who manually design and set up the bulk of these simulation environments, lists of actors, communication media (online articles, tweets, etc.), and content.
[0005] The development of these simulation environments is tedious, involving several people over several weeks, and often lacks reusability.
[0006] It is then desirable to overcome at least this drawback of the prior art, and to do so in a simple, effective and inexpensive manner.
[0007] It is particularly desirable to provide a solution for setting up cyber-influence and information warfare simulation environments which limits the need for human operator intervention. Statement of the invention
[0008] For this purpose, an automatic content generation system is proposed, configured to automatically generate content from scenario-level constraints corresponding to an influence graph between actors in a cyber-influence and information warfare simulation scenario, the automatic content generation system comprising:
[0009] - a set of databases comprising a knowledge base and a database of constraint data;
[0010] - a configuration module comprising a scheduler and a selector of constraints, the scheduler being configured to derive sequence-level constraints from the scenario-level constraints and the knowledge base, the sequence-level constraints being representative of an action graph that defines a temporal sequence of documents to be automatically produced on behalf of said actors, and the constraint selector being configured to derive document-level constraints from the sequence-level constraints and linguistic constraints of the constraint database;
[0011] - and a content generation module.
[0012] Furthermore, said system is such that the configuration module is further capable of customizing an artificial intelligence model of the language model type with constraints, by training said artificial intelligence model from pre-existing content and said document-level constraints. Furthermore, said system is such that the content generation module is further capable of automatically generating content using the customized artificial intelligence model and said document-level constraints.
[0013] Thus, thanks to such a gradual approach by levels of constraints, a significant quantity of content can be automatically generated, which facilitates the establishment of cyber-influence and information warfare simulation environments by limiting the need for interventions by human operators.
[0014] In a particular embodiment, the set of databases further comprises a database of real contents which are or have been accessible in open source and which have been collected, and in the context of the training of said artificial intelligence, said pre-existing contents include real contents stored in the database of real contents.
[0015] In a particular embodiment, the set of databases further comprises a database of generated content which is content automatically generated by the automatic content generation system during cyber-influence and information warfare simulation sessions, and in the context of training said artificial intelligence, said pre-existing content includes generated content stored in the database of generated content.
[0016] In a particular embodiment, the set of databases further comprises a database of artificial intelligence models, the configuration module further comprises an artificial intelligence model selector capable of selecting, in the database of artificial intelligence models, a generic artificial intelligence model on which the training is carried out to obtain the customized artificial intelligence model.
[0017] In a particular embodiment, said system further comprises a ges user interface manager, as well as a content evaluation module capable of calculating a set of metrics on automatically generated content and comparing these metrics to predefined thresholds, in order to provide via the user interface manager a qualitative dashboard with regard to automatically generated content.
[0018] In a particular embodiment, the configuration module further comprises a constraint formatting module capable of converting document-level constraints from a machine-adapted language to a natural language.
[0019] In a particular embodiment, the configuration module is capable of analyzing input data to recover information making it possible to establish scenario-level constraints, to further recover keywords allowing the scheduler, by searching for similarities in the knowledge base, to define content contexts in the sequence-level constraints.
[0020] A method for automatically generating content from scenario-level constraints corresponding to an influence graph between actors in a cyber-influence and information warfare simulation scenario is also proposed, the method being implemented by an automatic content generation system comprising a set of databases comprising a knowledge base and a constraint database, the method comprising the following steps:
[0021] - derive sequence level constraints from scenario level constraints and the knowledge base, the sequence level constraints being representative of an action graph which defines a temporal sequence of documents to be produced automatically on behalf of said actors,
[0022] - derive document-level constraints from document-level constraints sequence and linguistic constraints of the constraint database;
[0023] - customize an artificial intelligence model of the language model type with constraints, by training said artificial intelligence model from pre-existing content and said document-level constraints; and
[0024] - automatically generate content using the customized template artificial intelligence and said document-level constraints.
[0025] Also provided is a computer program product comprising instructions for implementing the method presented above, when said instructions are executed by a processor. Also provided is a (non-transitory) information storage medium storing a computer program comprising instructions for implementing the method presented above, when said instructions are read from the information storage medium and executed by a processor. Brief description of the drawings
[0026] The following description of at least one embodiment is set forth in relation to the accompanying drawings, among which:
[0027] [Fig-1] schematically illustrates an architecture of a generation system at automation of content for the purposes of simulating cyber-influence and information warfare;
[0028] [Fig.2] schematically illustrates an arrangement of a configuration module included in the automatic content generation system;
[0029] [Fig.3] schematically illustrates a flowchart of a generation process at content automation;
[0030] [Fig.4] schematically illustrates an example of an influence graph and a graph actions, which are used in the automatic content generation process;
[0031] [Fig.5] schematically illustrates a refinement of constraints within the framework of the present invention; and
[0032] [Fig.6] schematically illustrates an example of a hardware unit allowing an im implementation of the automatic content generation process by the automatic content generation system.
[0033] DETAILED DESCRIPTION OF EMBODIMENTS
[0034] [Fig.l] schematically illustrates an architecture of a SYS 100 system for automatic generation of content for the purposes of simulating cyber-influence and information warfare, in particular for the purposes of training or coaching in understanding cyber-influence and information warfare. The generated content is based on texts, which may possibly be supplemented by publications of other types, such as images and / or videos and / or sound signals.
[0035] In an illustrative and non-limiting manner, the remainder of the description is based on embodiments of automatic generation of content for the purposes of training in understanding cyber-influence and information warfare.
[0036] The SYS 100 system for automatic generation of content comprises a set of DBS databases (“DataBases Set” in English) 110.
[0037] The set of databases DBS 110 includes a constraint database CDB (Constraints Database) 112. The constraint database CDB 112 contains numerous linguistic constraints classified by type (type of emotion, type of behavior, type of theme, etc.) making it possible to restore constraints for automatic generation of text of various natures from which to choose to set up this or that training scenario.
[0038] The set of databases DBS 110 includes a database of real contents RCDB (“Real Content Database” in English) 113. The real contents are contents which are or have been accessible in open source and which have been collected, in particular collected on social networking platforms (eg Twitter (brand registered trademark), Facebook (registered trademark)...), websites (online press, blogs...), or for example digital content (newspapers, magazines...).
[0039] The set of databases DBS 110 preferably includes a database of generated content GCDB (“Generated Content Database” in English) 114. The generated content is content automatically generated by the system SYS 100 for automatic generation of content during simulation sessions (e.g., training or practice) in cyber-influence and information warfare.
[0040] The set of databases DBS 110 preferably includes an artificial intelligence model database AIMDB (“Artificial Intelligence Models Database” in English) 111. An artificial intelligence model can thus be selected from different artificial intelligence models stored in the artificial intelligence model database AIMDB 111, such as for example GPT-2 (“Generative Pre-trained Transformer 2” in English), GPT-3 (“Generative Pre-trained Transformer 3” in English) or LaMDA (“Language Model for Dialogue Applications” in English), to name only these from the multitude of existing language models. Several artificial intelligence models make it possible in particular to automatically generate content in different languages (German, Chinese, etc.).
[0041] The artificial intelligence models stored in the AIMDB 111 artificial intelligence model database are constrained language models. These artificial intelligence models respect a base of lexical constraints allowing the generation of coherent text, for example from a set of keywords or from a simple beginning of a sentence. As disclosed below, such constrained language models cannot in themselves constitute an adequate training environment, and must therefore be customized (or specialized) to meet the specific needs of the training scenario.
[0042] Preferably, the artificial intelligence models stored in the AIMDB artificial intelligence model database 111 are pre-trained language models.
[0043] The set of databases DBS 110 includes a knowledge base KB (“Knowledge Base” in English) 115. The knowledge defines relational behavior patterns between actors in a training scenario (influence relationship, conflict relationship, chain of responses to messages, etc.). This knowledge makes it possible to define an action graph to set up the training scenario from an influence graph between actors. An example of an influence graph and an action graph is presented below in relation to [Fig. 4]. In addition, the knowledge includes collections of themes, vo cabular, of emotions, in order to allow the development of content on such and such a subject or around such and such a keyword.
[0044] The SYS 100 system for automatic content generation further comprises a PM processing module (Processing Module) 120 adapted and configured to automatically generate content from ID input data (Input Data) 101, more particularly using a CGM content generation module 122. The ID input data 101 comprises information describing a training scenario: description of actors in the training scenario (names, social network accounts, etc.) according to a graph of social relationships (members of the same community, influence relationships, opposition relationships, etc.), description of mood states of the actors, description of communication support (press articles, tweets, etc.) associated with the actors, keywords relating to subjects or events addressed in the training scenario, volume of content to be created (quantity of messages, etc.), temporality, etc.The input data ID 101 defines a framework for the training scenario for which contents are to be automatically generated. The processing module PM 120 is further adapted and configured to provide these contents in the form of generated data 102.
[0045] The SYS 100 system for automatic generation of content further comprises a user interface manager UIM (“User Interface Manager” in English) 130. The user interface manager UIM 130 may include a display and input devices (keyboard, etc.) and pointing devices (mouse, etc.), or in general, devices for interaction with one or more users. The user interface manager UIM 130 may include an interface with a communication network to allow remote interaction. The user interface manager UIM 130 is adapted and configured to receive the input data ID 101 and forward them to the processing module PM 120. The user interface manager UIM 130 is further adapted and configured to receive the generated data 102 from the processing module PM 120 and make them available to a training platform (or more generally, a simulation platform).
[0046] In a particular embodiment, the user interface manager UIM 130 is further adapted and configured to allow one or more users to evaluate automatically generated content, to give instructions for refining this content (delete or correct messages, rectify redundancies, shorten messages, introduce nuances, add keywords, etc.) and to validate automatically generated content in order to make the content thus validated available to the training platform (or more generally, the simulation platform).
[0047] The PM 120 processing module comprises a CM configuration module ("Configuration Module") 121 adapted and configured to configure the PM processing module 120, and more particularly the CGM content generation module 122 according to the input data ID 101, and any refinement instructions. The CM configuration module 121 is adapted and configured to customize (or specialize) a language model with constraints, that is to say to add new constraints to it, in order to make it suitable for the specificities of the training to be prepared. A particular embodiment of the CM configuration module 121 is disclosed below in relation to [Fig.2].
[0048] The PM processing module 120 further comprises a content generation module CGM (“Content Generation Module”) 122 adapted and configured to automatically generate content while respecting the constraints defined by the CM configuration module 121. The CGM content generation module 122 uses an artificial intelligence model, as customized by the CM configuration module 121, to automatically generate the content while respecting said constraints.
[0049] The PM processing module 120 preferably further comprises a CAM content assessment module (“Content Assessment Module”) 123 adapted and configured to manage the assessment of the content automatically generated by the CGM content generation module 122 and to transmit any refinement instructions to the CM configuration module 121 in order to instruct the CGM content generation module 122 to update the automatically generated content accordingly. Preferably, the CAM content assessment module 123 is adapted and configured to calculate a set of metrics on the generated content and compare these metrics to predefined thresholds, in order to provide via the UIM user interface manager 130 a qualitative dashboard with respect to the automatically generated content and thus guide one or more users in the possible refinement and validation of the automatically generated content.
[0050] The PM processing module 120 further comprises a CEM content export module (“Content Export Module”) 124 configured and adapted to export the generated data 102, namely the content automatically generated by the CGM content generation module 122.
[0051] [Fig.2] schematically illustrates an arrangement of the configuration module CM 121.
[0052] The configuration module CM 121 is arranged and configured to provide a customized artificial intelligence model CAIM (Customized Artificial Intelligence Model in English) 152 from the input data ID 101.
[0053] The configuration module CM 121 comprises a scheduler SCH 200. The scheduler SCH 200 is adapted and configured to transform scenario level constraints representative of the influence graph into constraints of sequence level representative of the action graph. The SCH 200 scheduler is thus adapted and configured to develop a diagram of content interactions between actors in the training scenario.
[0054] The configuration module CM 121 comprises a constraint selector CS (Constraints Selector) 201 adapted and configured to transform the sequence-level constraints into document-level constraints. The constraint selector CS 201 is configured and adapted to select linguistic constraints (synonyms, antonyms, contextual meaning of keywords in accordance with the context of the scenario, style adapted to the document media (tweets, press articles, etc.)) from the constraint database CDB 112 to define document structures conforming to the action graph.
[0055] In a particular embodiment, the configuration module CM 121 further comprises a constraint formatting module CF (“Constraints Formator” in English) 202. The configuration module CM 121 is adapted and configured to switch from a machine-friendly language, easily manipulated by human operators such as a programming language, to a natural language, in order to facilitate processing by an artificial intelligence model for automatic text generation.
[0056] The configuration module CM 121 preferably further comprises an artificial intelligence model selector AIMS (“Artificial Intelligence Model Selector” in English) 203. The artificial intelligence model selector AIMS 203 is configured and arranged to select a generic artificial intelligence model in the artificial intelligence model database AIMDB 111 which is adapted to the training scenario. A generic artificial intelligence model is a pre-existing artificial intelligence model, which is not yet specialized for the simulation / training theme. For example, if the input data ID 101 defines a specific language, the artificial intelligence model selector AIMS 203 selects a language model of the BERT (“Bidirectional Encoder Representations from Transformers” in English) type particularly adapted to this specific language.Alternatively, the AIMS 203 artificial intelligence model selector selects the generic artificial intelligence model to be used by following a user configuration expressed for example directly in the input data ID 101. .
[0057] The configuration module CM 121 is arranged and configured to carry out additional training of the artificial intelligence model selected by the artificial intelligence model selector AIMS 203, namely the so-called “generic” artificial intelligence model. The training is carried out using pre-existing content. In a particular embodiment, these pre-existing contents existing are real contents that are obtained from the real contents database RCDB 113 and that are selected according to the document-level constraints established by the constraint selector CS 201, possibly formatted by the constraint formatting module CF 202. In a particular embodiment, these pre-existing contents are supplemented by generated contents that are obtained from the generated contents database GCDB 114 and that are selected in the same way. After training, a specialized artificial intelligence model for the desired scenario is thus obtained, namely the customized artificial intelligence model CAIM 152. This aspect is detailed below in relation to [Fig. 3].
[0058] [Fig.3] schematically illustrates a flowchart of an automatic content generation method as implemented by the SYS 100 automatic content generation system.
[0059] In a step 300, the automatic content generation system SYS 100, and more particularly the configuration module CM 121, obtains and analyzes the input data ID 101. Predefined tags can be used to delimit and identify information contained in the input data ID 101. Named-Entity Recognition (NER) can be used, as a variant or in addition, to recognize and identify information contained in the input data ID 101.
[0060] The input data ID 101 are analyzed to retrieve information making it possible to establish the scenario-level constraints. In particular, the input data ID 101 are analyzed to retrieve information making it possible to identify the actors of the training scenario (identification of factions, quantity of actors per faction, etc.) and their relationships with each other, so as to be able to establish the influence graph.
[0061] The input data ID 101 are analyzed to further recover keywords allowing the SYS 100 system for automatic content generation, and more particularly the SCH 200 scheduler, by searching for similarities in the knowledge base KB 115, to define content contexts in the sequence level constraints.
[0062] In a step 301, the SYS 100 system for automatic content generation defines an influence graph between actors of the training scenario and an action graph. The influence graph defines the relationships between the actors of the training scenario and the action graph defines a temporal sequence of documents (tweets, press articles, etc.) to be produced automatically on behalf of said actors. The scenario-level constraints are a machine language representation of the influence graph. The action graph is represented, in machine language, by sequence-level constraints. The sequence-level constraints are derived from the scenario level constraints.
[0063] A proposal for an influence graph, based on the input data ID 101 and the knowledge base KD 115, can be made to one or more users via the user interface module UIM 130, for possible adjustment (i.e. possible adjustment of the scenario level constraints) and validation.
[0064] A proposal for an action graph, based on the input data ID 101 (keywords, etc.) and the influence graph, can be made to one or more users via the user interface module UIM 130, for possible adjustment (i.e. possible adjustment of the sequence level constraints) and validation.
[0065] In a step 302, the automatic content generation system SYS 100 defines document-level constraints derived from the sequence-level constraints. This aspect is detailed below in relation to [Fig.5].
[0066] In a step 303, in a particular embodiment, the automatic content generation system SYS 100 formats the document-level constraints by converting them from a machine-friendly language to a natural language, in order to make them more easily interpretable by an artificial intelligence model for automatic text generation.
[0067] In a particular embodiment, in a step 304, the automatic content generation system SYS 100 selects an artificial intelligence model from the artificial intelligence model database AIMDB 111 based on the input data ID 101. Alternatively, the automatic content generation system SYS 100 uses a single predefined artificial intelligence model. This artificial intelligence model is a language model with constraints respecting a base of lexical constraints making it possible to generate coherent text. This artificial intelligence model is said to be “generic” and must still be customized.
[0068] In a step 305, the automatic content generation system SYS 100 generates a customized artificial intelligence model from the generic artificial intelligence model taking into account the document level constraints obtained in step 302, possibly formatted via step 303. The generic artificial intelligence model then undergoes additional training in order to be customized to respect the constraints specific to the training scenario.
[0069] The generic artificial intelligence model is customized from pre-existing contents, such as real contents, which are retrieved from the real contents database RCDB 113, and possibly from generated contents, which are retrieved from the generated contents database GCDB 114. The real contents stored in the real contents database RCDB 113, and the generated contents stored in the generated contents database GCDB 114, are associated with document-level constraint schemas that said contents respect. The contents retrieved from the real content database RCDB 113, and possibly from the generated content database GCDB 114, are contents associated with document-level constraint schemas that correspond to the document-level constraints obtained in step 302, possibly formatted via step 303. For example, the contents retrieved from the real content database RCDB 113, and possibly from the generated content database GCDB 114, are obtained by searching for similarities with keywords contained in the document-level constraints, by applying a named-entity recognition technique NER (Named-Entity Recognition). Thus, the additional training provides machine learning focusing on the constraints resulting from the training scenario.
[0070] In a step 306, the SYS 100 system for automatic content generation uses the customized artificial intelligence model to generate content that therefore respects the scenario level constraints expressed in the training scenario.
[0071] In a step 307, the automatic content generation system SYS 100 interacts with one or more users as part of an evaluation of the automatically generated content. In the case where the automatically generated content is validated, a step 309 is performed; otherwise, a step 308 is performed.
[0072] In step 308, the automatic content generation system SYS 100 obtains instructions for refining the automatically generated content. Then, the automatically generated content is modified accordingly. Depending on the desired refinement, the influence graph between actors in the training scenario and the action graph may have to be modified (looping back to step 301). Another possibility is that the selection of constraints has to be modified (looping back to step 302), which requires a corresponding modification of the customized artificial intelligence model (step 305) with possibly a new prior selection of a generic artificial intelligence model (step 304). Another possibility is that only the generated documents have to be modified (looping back to step 306).For example, by refining certain automatically generated documents, the automatic content generation system SYS 100 will make new proposals for other automatically generated documents (which, according to the action graph, depend on the refined documents). Finally (not shown in [Fig. 3]), it is possible that one or more human operators make minor modifications (adding a sentence, changing a word, etc.) in one or more automatically generated documents, and step 309 is then carried out without having to repeat an automatic content generation pass.
[0073] In step 309, the automatic content generation system SYS 100 exports the automatically generated and validated content (namely the aforementioned documents) for subsequent use on the training platform (or more generally, on the simulation platform). In a particular embodiment, the automatically generated and validated content is stored in the generated content database GCDB 114 so that it can in turn be used to train an artificial intelligence model during future training sessions.
[0074] [Fig.4] schematically illustrates an example of an influence graph GI (“Graph of Influence” in English) 400 and GA (“Graph of Actions” in English) 450, which are used in the automatic content generation process.
[0075] In the influence graph GI 400, four factions of actors are represented. A first faction F1 410 corresponds for example to a pro-“campl” faction, and a second faction F2 420 corresponds for example to a pro-“camp2” faction, in the context of a subject concerning a war between “campl” and “camp2”. A third faction F3 430 is neutral, and a fourth F4 440 is in favor of an influential ally of “campl”. Each faction contains one or more, and typically several dozen, actors in the training scenario.
[0076] The influence graph GI 400 establishes a first conflict relationship 401 between the first faction F1 410 and the second faction F2 420. The influence graph GI 400 establishes a second influence relationship 402 between the fourth faction F4 440 and the first faction F1 410, where the fourth faction F4 440 dominates (in terms of influence) the first faction F1 410.
[0077] The influence graph GI 400 can also establish relationships within the same faction. For example, in the first faction F1 410, a third influence relationship 403 between a first group of actors A1 411 and a second group of actors A2 412 is established, where the second group of actors A2 412 dominates (in terms of influence) the first group of actors A1 411.
[0078] Thus, the influence graph GI 400 establishes that:
[0079] - documents (messages, press articles, etc.) published by an actor of a faction are likely to result in a message of support from another actor of the same faction,
[0080] - documents (messages, press articles, etc.) published by an actor of a faction (e.g., of the first faction Fl 410) in conflict relationship with another faction (e.g., the second faction F2 420) is subject to resulting in a contradictory message from an actor of that other faction,
[0081] - documents (messages, press articles, etc.) published by an actor of a faction (e.g., of the fourth faction F4 440) in a relationship of downward influence (domination) with another faction (e.g., of the first faction Fl 440) are subject to resulting in a relay or amplification message from an actor of this other faction,
[0082] - documents (messages, press articles, etc.) published by an actor of a first group of a faction (e.g., of the second group A2 412) in a relationship of descending influence (domination) with another group of this faction (e.g., the first group Al 411) are subject to entail a relay or amplification message from an actor of this other group,
[0083] - etc.
[0084] Other types of relationships can be defined in the influence graph GI 400.
[0085] Once the influence graph GI 400 has been established between the different actors of the training scenario, a scheduling operation 470 is carried out. This scheduling operation produces the action graph GA 450.
[0086] The GA 450 action graph comprises a representation of a chronological sequence of document descriptors intended to constitute the automatically generated content. Each document descriptor exposes at least one document type information (tweet, email, article, etc.), a timestamp information in the chronological sequence, information on the actor who authored the document, and information on the faction of the author of the document.
[0087] As an illustration, in the action graph GA 450, the chronological sequence begins with a first document DI 451. The document DI 451 is, for example, a tweet published by an actor of the first faction F1 410. The publication of the document DI 451 leads to the publication of a document D2 452 preceded by the publication of a document D3 453. The document D2 452 is, for example, a tweet published by an actor of the second faction F2 420 in reaction to opposition to the document DI 451. The document D3 453 is, for example, a tweet published by an actor of the third faction F3 430. The publication of the document D2 452 leads to the publication of a document D4 454 and then a document D5 455. The document D4 454 is, for example, a tweet published by an actor of the second faction F2 420 in support of the document D2 452 and document D5 455 is for example a tweet published by an actor of the first faction Fl 410 in reaction of opposition to document D2 452.The publication of documents D4 454 and D5 455 results in the publication of a document D6 456. Document D6 456 is, for example, a blog post published by an actor of the third faction F3 430. A document D7 457 is then published. Document D7 457 is, for example, a tweet from an actor of group A2 412 of the first faction Fl 410. The publication of documents D6 456 and D7 457 results in the publication of a document D8 458. Document D8 458 is, for example, a tweet from an actor of group Al 411 of the first faction Fl 410 in support of document D6 456 and in critical reaction to document D7 457.
[0088] The GA 450 action graph may show a more complex chronological sequence of document descriptors, with more or less relationships of causality between document publications, including documents published in isolation (without reaction). The document descriptors mentioned above are expressed as constraints, here sequence-level constraints, in order to be easily derived into document-level constraints and thus be expressed in a form suitable for processing by an artificial intelligence model of the constrained language model type.
[0089] [Fig.5] schematically illustrates an arrangement of the SYS 100 system for automatic content generation, which makes it possible to derive scenario-level constraints defining the broad outlines of a training scenario into sequence-level constraints defining a sequence of documents, then into document-level constraints defining constraints that can be used to automatically generate document content that is adapted to the training scenario.
[0090] From the scenario-level constraints SceLC (Scenario-Level Constraints) 501, the scheduler SCH 200 derives the sequence-level constraints SeqLC (Sequence-Level Constraints) 502. To generate the sequence-level constraints SeqLC 502 from the scenario-level constraints SceLC 501, the scheduler SCH 200 uses the knowledge base KB 115.
[0091] The scenario level constraints SceLC 501 are defined by the influence graph GI 400, that is to say that the scenario level constraints SceLC 501 are a representation of the influence graph GI 400 in the form of constraints.
[0092] The sequence level constraints SeqLC 502 are defined by the action graph GA 450, i.e. the sequence level constraints SeqLC 502 are a representation of the action graph GA 450 in the form of constraints.
[0093] From the sequence-level constraints SeqLC 502, the constraint selector CS 201 derives the document-level constraints DocLC (“Document-Level Constraints” in English) 503. The document-level constraints DocLC 503 can be used by the configuration module CM 121 to carry out the additional training of the artificial intelligence model selected by the artificial intelligence model selector AIMS 203, after possible formatting by the constraint formatting module CF 202. To generate the document-level constraints DocLC 503 from the sequence-level constraints SeqLC 502, the scheduler SCH 200 uses the constraint database CDB 112 (linguistic constraints) and a subset of the knowledge base KB 115 which was selected by the scheduler SCH 200 according to the scenario-level constraints SceLC 501.
[0094] Let's take a simplified example in which the input data ID expresses scenario-level constraints, as follows:
[0095] Type: microblogs
[0096] Actors: [{factionl: “campl”, number: 50], {faction2: “camp2”, number: 30], {faction3: neutral, number: 50}]
[0097] Subjects: [Subject: {"responsibility for crimes: attributed to the "campl"", lexical field: [hatred, racism, organized crime, human rights, evidence, investigation], exact mentions: ["City", "Organization!", "Personnel"]}; Subject2: "reality discovered civilian deaths: complete lie", lexical field: [propaganda, editing, manipulation, evidence, investigation], exact mentions: ["City2", "Person2"]}]
[0098] Postures: [{campl: [{Subjectl, denial], {Subject2, support}], {camp2: [{Subjectl, support}, {Subject2, denial}], {neutral: [{Subjectl, questioning}, {Subject2, diffuse}]
[0099] where: “campl” identifies a first belligerent camp, “camp2” identifies a second belligerent camp, “Villel” identifies a first city name, “Ville2” identifies a second city name, “Organisation!” identifies a first organisation name, “Personnel” identifies a first person name, “Person2” identifies a second person name.
[0100] As it appears in the simplified example above, the scenario level constraints provide lexical field information (and which therefore make it possible to define similar words or expressions to be used in the documents which will be automatically generated), as well as information (exact_mentions) intended to be mentioned exactly (and which must therefore be found as such in a certain number of documents which will be automatically generated).
[0101] The SCH 200 scheduler then defines the GA 450 action graph, which defines a number of documents to be generated automatically, with for each, a type of document, a timestamp (temporality), an author, a behavior, a posture with respect to a subject, and potentially other specific fields (places, emotion, type of actor, etc.).
[0102] Sequence No. 1:
[0103] TS1, Document: faction, (Subject, posture: denial), exact_mention: “Ville”, tone: passionate, author: institutional
[0104] TS2, Document2: faction3, (Subject, posture: request), tone: factual, author: media
[0105] TS3, Documents: faction2, (Subject, posture: support), tone: vindictive, author: individual
[0106] TS4, Document4: (Document reaction, opposition), exact_mention: “Villel”, faction2, (Subject, posture: support), tone: vindictive, author: institutional
[0107] where: TS1, TS2, TS3 and TS4 are timestamp (chronology) information.
[0108] The sequence presented above is taken from a model ("template" in English), of the "position" type, obtained from the KB 115 knowledge base. Numbers of actors and documents composing the sequence can be randomly or pseudo-randomly chosen. Attributes (institutional / media / individual type, emotion, posture...) of these actors can also be drawn randomly or pseudo-randomly.
[0109] Other sequences are defined in addition to Sequence No. 1 above in order to form the action graph GA 450, in particular to deal with Subject 2, etc. This same sequence model or other sequence models can be used to define these other sequences, so as to respect all the scenario-level constraints, as regards subjects, factions, postures, etc.
[0110] As already indicated, the action graph GA 450 can define links between these documents, so as to generate content in the form of conversations or chain reactions. For example, the document labeled Document4 can be linked to the document labeled Documentl by means of a link (“Documentl reaction”) indicating that the actors respond to each other.
[0111] The constraints expressed above are intended to be supplemented by linguistic constraints, from the constraint database CDB 112, in order to form, for each document to be produced, the document-level constraints to be provided to the content generation module CGM 122 to automatically generate the content of said document. An example of such linguistic constraints is proposed below.
[0112] Document:
[0113] Type: microblog
[0114] Timestamp: [TS1]
[0115] Constraint s Text:
[0116] Lexical field: [hatred, racism, organized crime]; Keywords: [“Villel”, aggression, fabricated evidence]; Behavior: denial; Emotion: neutral
[0117] Author: idlOOl
[0118] Document2:
[0119] Type: microblog
[0120] Timestamp: [TS2]
[0121] Constraints Text: Lexical field: [hate, racism, organized crime, evidence]; Keywords: [reliable evidence, expectation]; Behavior: neutral; Emotion: doubt
[0122] Author: id3001
[0123] Documents:
[0124] Type: microblog
[0125] Timestamp: [TS3]
[0126] Constraints Text: Lexical field: [racism, organized crime, investigation]; Keywords: [criminal acts, justice]; Behavior: approval; Emotion: anger, disgust
[0127] Author: id2001
[0128] Document4:
[0129] Type: microblog
[0130] Timestamp: [TS4]
[0131] Constraints Text: Lexical field: [racism, organized crime, human rights]; Keywords: [disinformation, dictatorship, criminals]; Behavior: denial; Emotion: anger, disgust
[0132] Author: id2002
[0133] Link: response to [Documentl]
[0134] In the example above, the information entered in “Text Constraints” blocks is information transmitted as is to a text generator itself. The lexical field information for a document is a partial replication of the lexical field information of the subject (e.g., Subjectl) indicated in the sequence level constraints for said document (and which comes, upstream, from the scenario level constraints). This lexical field information can include around ten or fifteen words allowing a text generator to be oriented on the theme to be addressed in automatically generated text. The keyword information is words or expressions (e.g., 3 to 6 words or expressions) to be placed in the automatically generated text, these keywords being either taken as is from the subject (e.g.,, Subject) indicated in the sequence level constraints for said document (and which come, upstream, from the scenario level constraints), or derived by similarities using the CDB 112 constraint database.
[0135] In the example above, authors are identified by an identifier (eg, id2002) linked to the faction to which the author in question belongs.
[0136] [Fig.6] schematically illustrates an example of a hardware unit 600 adapted and configured to implement the SYS 100 system for automatic content generation.
[0137] The hardware unit 600 comprises, connected by a communication bus 610: a processor or CPU (for “Central Processing Unit” in English) 601, or a cluster of such processors, such as for example GPUs (“Graphics Processing Units” in English); a RAM (for “Random Access Memory” in English) 602; a ROM (for “Read Only Memory” in English) 603, or a rewritable memory of the EEPROM type (“Electrically Erasable Programmable ROM” in English), for example of the Flash type; a data storage device, such as a hard disk HDD (for “Hard Disk Drive” in English) 604, or a storage media reader, such as an SD (for “Secure Digital” in English) card reader;a set of input and / or output I / O interfaces (“Inputs / Outputs” in English) 605, such as human-machine interfaces and / or communication interfaces, enabling in particular the SYS 100 system for automatic content generation to receive the input data ID 101 (or information enabling; constitute the input data ID 101) and export the generated data GD 102.
[0138] The processor 601 is capable of executing instructions loaded into the RAM 602 from the ROM 603, from an external memory (not shown), from a storage medium, such as an SD card or the HDD 604, or from a communication network. When the hardware unit 600 is powered on, the processor 601 is capable of reading instructions from the RAM 402 and executing them. These instructions form a computer program causing the processor 601 to implement the architecture and methods described herein.
[0139] All or part of the steps, behaviors and algorithms described here can thus be implemented in software form by executing a set of instructions by a programmable machine, such as a DSP (Digital Signal Processor) or a processor, or be implemented in hardware form by a machine or a dedicated component (chip) or a dedicated set of components (chipset), such as an FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit).
[0140] Generally speaking, the SYS 100 system for automatic generation of contents therefore comprises electronic circuitry arranged and configured to implement the architecture, as well as the steps, behaviors and algorithms described here.
Claims
Claims
1. System (100) for automatic generation of content configured to automatically generate content from scenario-level constraints corresponding to an influence graph (400) between actors in a cyber-influence and information warfare simulation scenario, the system (100) for automatic generation of content comprising: - a set of databases (110) comprising a knowledge base (115) and a constraint database (112);- a configuration module (121) comprising a scheduler (200) and a constraint selector (201), the scheduler (200) being configured to derive sequence-level constraints from the scenario-level constraints and the knowledge base, the sequence-level constraints being representative of an action graph (450) which defines a temporal sequence of documents to be produced automatically on behalf of said actors, and the constraint selector (201) being configured to derive document-level constraints from the sequence-level constraints and linguistic constraints of the constraint database (112); - and a content generation module (122);wherein the configuration module (121) is further capable of customizing an artificial intelligence model of the language model type with constraints, by training said artificial intelligence model from pre-existing content and said document-level constraints; and wherein the content generation module (122) is further capable of automatically generating content using the customized artificial intelligence model and said document-level constraints.;
2. The system (100) of claim 1, wherein the set of databases (110) further comprises a database of real content (113) that is or has been open sourced and collected, and in training said artificial intelligence, said pre-existing content includes real content stored in the real content database (113).
3. The system (100) of claim 2, wherein the set of databases (110) further comprises a database of generated contents (114) which are contents automatically generated by the automatic content generation system (100) over the course of cyber-influence and information warfare simulation sessions, and as part of the training of said artificial intelligence, said pre-existing content includes generated content stored in the generated content database (114).
4. System (100) according to any one of claims 1 to 3, wherein the set of databases (110) further comprises a database of artificial intelligence models (111), the configuration module (121) further comprises an artificial intelligence model selector (203) capable of selecting, in the database of artificial intelligence models (111), a generic artificial intelligence model on which the training is carried out to obtain the customized artificial intelligence model.
5. System (100) according to any one of claims 1 to 4, further comprising a user interface manager (130), as well as a content evaluation module (124) capable of calculating a set of metrics on the automatically generated content and comparing these metrics to predefined thresholds, in order to provide via the user interface manager (130) a qualitative dashboard with respect to the automatically generated content.
6. The system (100) of any one of claims 1 to 5, wherein the configuration module (121) further comprises a constraint formatting module (202) capable of converting the document-level constraints from a machine-friendly language to a natural language.
7. System (100) according to any one of claims 1 to 6, in which the configuration module is capable of analyzing input data (101) to retrieve information making it possible to establish the scenario level constraints, to further retrieve keywords allowing the scheduler (200), by searching for similarities in the knowledge base (115), to define content contexts in the sequence level constraints.
8. Method for automatically generating content from scenario-level constraints corresponding to an influence graph (400) between actors in a cyber-influence and information warfare simulation scenario, the method being implemented by a system (100) for automatically generating content comprising a set of databases (110) comprising a knowledge base (115), and a database of constraint data (112), the method comprising: - deriving sequence-level constraints from scenario-level constraints and the knowledge base (115), the sequence-level constraints being representative of an action graph (450) which defines a temporal sequence of documents to be produced automatically on behalf of said actors, - deriving document-level constraints from sequence-level constraints and linguistic constraints in the constraint database (112); - customize an artificial intelligence model of the language model type with constraints, by training said artificial intelligence model from pre-existing content and said document-level constraints; and - automatically generate content using the customized artificial intelligence model and said document-level constraints.
9. A computer program product comprising instructions for implementing the method according to claim 8, when said instructions are executed by a processor.
10. An information storage medium storing a computer program comprising instructions for implementing the method according to claim 8, when said instructions are read from the information storage medium and executed by a processor.