Training device, generation device, and terminal device

The described system addresses the challenge of generating stories that match existing character settings by using embedded numerical representations from descriptive and appearance documents, resulting in efficient and memory-effective story generation.

WO2025110072A1PCT designated stage expired Publication Date: 2025-05-30SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/040285
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-11-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing large language models struggle to generate stories that accurately match the settings of existing characters, due to limitations in conveying character settings and the need for extensive character information, which increases memory costs and complexity.

Method used

A learning device, generation device, and terminal device that utilize a first network for embedding descriptive documents and a second network for embedding appearance documents, with a learning unit that closes the numerical representations from both networks, allowing for the generation of stories that conform to character settings without requiring extensive character information.

Benefits of technology

Enables the efficient generation of stories that adhere to character settings, reducing the need for lengthy character descriptions and improving memory efficiency, while allowing for flexible and homogeneous story generation across multiple authors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024040285_30052025_PF_FP_ABST
    Figure JP2024040285_30052025_PF_FP_ABST
Patent Text Reader

Abstract

A training device according to the present disclosure comprises a training unit that trains a first network for acquiring a first numerical expression by embedding based on an explanatory document for explaining a character and included in an existing document and a second network for acquiring a second numerical expression by embedding based on an appearance document concerning the character appearing in a story and included in the existing document, so that the first numerical expression and the second numerical expression are close to each other. The first numerical expression resulting from training by the training unit is outputted as a character numerical expression representing the character.
Need to check novelty before this filing date? Find Prior Art

Description

Learning device, generation device, and terminal device

[0001] The present disclosure relates to a learning device, a generation device, and a terminal device.

[0002] In recent years, with the emergence of large language models (LLMs) such as Generative Pre-trained Transformer (GPT)-3 and ChatGPT, story generation has become easier.

[0003] Japanese Patent Laid-Open No. 8-314890 Japanese Patent Laid-Open No. 2008-084348

[0004] However, in the past, even when a large-scale language model was used, it was difficult to generate a story that matched the settings of existing characters.

[0005] Therefore, an object of the present disclosure is to provide a learning device, a generation device, and a terminal device that can easily generate a story based on the settings of existing characters.

[0006] The learning device of the present disclosure includes a first network for obtaining a first numerical expression by embedding based on an explanatory document for explaining a character contained in an existing document, and a second network for obtaining a second numerical expression by embedding based on an appearance document in which the character appears in a story contained in the existing document, and a learning unit that trains the first numerical expression and the second numerical expression so that they become closer to each other, and outputs the first numerical expression learned by the learning unit as a character numerical expression representing the character.

[0007] The generation device according to the present disclosure includes a generation unit that generates a document that is a generated sentence using a third network based on a first numerical representation generated using a first network that performs embedding based on an explanatory document to explain a character and a second network that performs embedding based on a document in which the character appears in a story, and a target document that is the subject of editing, and the generation unit trains the third network using a loss function based on the generated sentence.

[0008] The terminal device according to the present disclosure comprises an input unit that accepts user operations and a display control unit that controls a display screen displayed on a display device, and the display control unit generates a display control signal for causing the display device to display a user interface screen including: a numerical representation representing the character generated using a first network that performs embedding based on an explanatory document to explain the character and a second network that performs embedding based on an appearance document in which the character appears in a story; a target document that is the document to be edited; a document area that displays a generated sentence based on the numerical representation; and a character selection area for selecting an appearance character that the user wants to have appear in the generated sentence displayed in the document area, and places identification information associated with the numerical representation in the character selection area.

[0009] 1 is a block diagram showing the configuration of an example of an information processing system applicable to an embodiment. FIG. 1 is a functional block diagram for explaining functions of a server and a user terminal according to an embodiment. FIG. 1 is a block diagram showing the hardware configuration of an example of a server applicable to an embodiment. FIG. 2 is a block diagram showing the hardware configuration of an example of a user terminal applicable to an embodiment. FIG. 3 is a schematic diagram for explaining explanation-based character encoding according to an embodiment. FIG. 4 is a flowchart of an example for explaining processing by explanation-based character encoding according to an embodiment. FIG. 5 is a schematic diagram for explaining evaluation of a character expression CharEmb using an SNN according to an embodiment. FIG. 6 is a flowchart of an example for explaining processing by characteristic-conditional decoding according to an embodiment. FIG. 7 is a schematic diagram showing an example of attribute labels indicating attributes according to an embodiment. FIG. 8 is a flowchart of an example for explaining the overall flow of document generation processing according to an embodiment. FIG. 9 is a schematic diagram showing an example of a UI screen for determining a location where a sentence is to be generated according to an embodiment. FIG. 10 is a schematic diagram showing an example of visualizing and presenting the relationship between multiple characters according to an embodiment. FIG. 11 is a schematic diagram showing another UI screen for determining a location where a sentence is to be generated according to an embodiment. FIG. 12 is a schematic diagram showing yet another UI screen for determining a location where a sentence is to be generated according to an embodiment. FIG. 10 is a schematic diagram showing an example of yet another UI screen for determining a portion for which a sentence is to be generated, according to an embodiment; FIG. 11 is a schematic diagram showing an example of a UI screen for inputting information for document generation, according to an embodiment; FIG. 12 is a schematic diagram showing an example of a UI screen that allows a user to review a generated sentence according to variations in expression, according to an embodiment; and FIG. 13 is a schematic diagram for explaining an example of generating a generated sentence using an image of a manga, according to a modified example of an embodiment.

[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are denoted by the same reference numerals, and redundant description will be omitted.

[0011] Hereinafter, embodiments of the present disclosure will be described in the following order: 1. Overview of embodiments of the present disclosure 2. Configurations applicable to embodiments of the present disclosure 3. Existing technologies 4. Embodiments of the present disclosure 4-1. Model training method according to embodiments 4-2. Story generation method according to embodiments 4-3. User interface according to embodiments 5. Effects of embodiments of the present disclosure 6. Modified embodiments of embodiments of the present disclosure

[0012] (1. Overview of Embodiments of the Present Disclosure) First, an outline of an embodiment of the present disclosure will be described. The embodiment of the present disclosure aims to generate a story that matches the character settings of existing content and to improve the efficiency of the writing process for scenario writers.

[0013] Here, a story that matches the character's settings refers to a story (text) that does not contradict the character's physical characteristics (appearance, weight, gestures, etc.), mental characteristics (beliefs, tone of voice, etc.), and social position (relationships with other characters, affiliation, race, etc.). Also, a story refers to documents that appear in novels, manga, and scripts. Therefore, a story includes both lines spoken by characters and text that describes a scene. However, not all text in a story is related to a certain character.

[0014] In an embodiment of the present disclosure, a compressed representation of an existing character is learned from an appearance document in which the character appears and an explanatory document that explains the character. The appearance document and explanatory document may be obtained from documents published on the Internet, such as a website. The compressed representation of the existing character may be, for example, a numerical representation generated by embedding the appearance document and explanatory document using a network such as a deep neural network (DNN).

[0015] Furthermore, in an embodiment of the present disclosure, a network is used to generate a document that is a generated sentence based on the compressed representation generated based on the character's explanatory document and appearance document and the target document to be edited. At this time, attributes to be added to the generated sentence and keywords to be included in the generated sentence may be used as operators to control the generation of the generated sentence.

[0016] Furthermore, an embodiment of the present disclosure proposes a user interface that allows a user to easily apply compressed representations of existing characters to a target document.

[0017] The above-described configuration of the embodiment of the present disclosure enables a writer to generate a story related to a character by specifying it using flexible controls, thereby stimulating the writer's ideas, shortening the writing process, and supporting the creation of high-quality documents. Furthermore, by using a condensed representation of the character, when multiple writers write the same content, it becomes possible to generate a uniform story among multiple writers without having to deeply understand the character settings. Furthermore, when writing derivative content or a sequel to existing content, it becomes possible to flexibly provide a large amount of story material that can be difficult for writers.

[0018] Therefore, by applying the embodiments of the present disclosure, it becomes possible to easily generate a story that corresponds to the settings of existing characters.

[0019] (2. Configuration Applicable to Embodiments of the Present Disclosure) Next, a configuration applicable to embodiments of the present disclosure will be described. Fig. 1 is a block diagram showing the configuration of an example of an information processing system applicable to the embodiments.

[0020] In the following, neural networks such as DNN and SNN (Shallow Neural Network) will be simply referred to as "networks" to distinguish them from communication networks such as the Internet and LAN (Local Area Network).

[0021] 1, an information processing system 1 includes a server 3 and a user terminal 4 connected via a communication network 2. The communication network 2 may be the Internet or a local area network (LAN) established in a closed environment such as an in-house network.

[0022] The server 3 acquires training data to be used for training from, for example, a web server 5 connected to the communication network 2, and trains a network such as a DNN using the acquired training data. In addition, the server 3 generates a document using the trained training model in response to, for example, an instruction from a user terminal 4.

[0023] 1, the server 3 is shown as being configured in a cloud network 6 connected to the communication network 2. However, the server 3 is not limited to this and may be configured as a single computer or may be configured as being distributed across multiple computers.

[0024] The user terminal 4 is, for example, a general computer device that accepts input operations by the user and provides a user interface (UI) that presents information to the user by screen display or the like. The user terminal 4 may, for example, transmit commands in response to the user's input operations to the server 3 via the communication network 2. The user terminal 4 may also, for example, present information transmitted from the server 3 to the user by a display or the like.

[0025] FIG. 2 is a functional block diagram for explaining the functions of the server 3 and the user terminal 4 according to the embodiment.

[0026] (User Terminal) In FIG. 2, the user terminal 4 includes a document display control unit 40, a user input unit 41, and a display unit 42.

[0027] The user input unit 41 accepts input operations by the user. The user input unit 41 may be a pointing device such as a mouse, or a keyboard. The user input unit 41 may also be configured to accept voice input via a microphone or to detect gaze and enable gaze input.

[0028] The display unit 42 generates a display control signal for displaying a screen on a display device (not shown). The display unit 42 may include a display device. For example, the display unit 42 supplies the display control signal generated by the document display control unit 40 to the display device. The display device displays a screen corresponding to the supplied display control signal. Hereinafter, displaying a screen corresponding to the display control signal generated by the display unit 42 on the display device will be appropriately described as "the display unit 42 displays a screen," etc.

[0029] The document display control unit 40, for example, accepts input operations via the user input unit 41, generates a display control signal, and supplies the generated display control signal to the display unit 42. This allows the document display control unit 40 to form a user interface (UI). The document display control unit 40 may also communicate with the server 3 and send and receive data between the server 3. Furthermore, the document display control unit 40 can access the web server 5 via the communication network 20.

[0030] The document display control unit 40 may be, for example, a program that runs on the server 3 and performs input and output on a browser application installed on the user terminal 4. However, the document display control unit 40 may also be installed on the user terminal 4 as a dedicated program according to the embodiment.

[0031] (Server) In FIG. 2, the server 3 includes a learning data acquisition unit 31 , a learning unit 32 , a model evaluation unit 33 , a learning model storage unit 34 , a document search unit 36 ​​, and a document generation unit 37 .

[0032] The training data acquisition unit 31 acquires training data from the training document group 30. The training document group 30 may be documents held by the server 3, or may be documents collected or accessed via the communication network 2. For example, the training document group 30 may be documents made public on the web server 5.

[0033] The learning unit 32 learns a learning model based on the learning data acquired by the learning data acquisition unit 31. The learning model storage unit 34 stores the learned model learned by the learning unit 32 in a storage device or the like. Hereinafter, the storage of the learned model by the learning model storage unit 34 in a storage device or the like will be appropriately described as "the learning model storage unit 34 stores the learned model," etc.

[0034] The model evaluation unit 33 evaluates the model learned by the learning unit 32. The model evaluation unit 33 may use an SNN, which is smaller in scale than a DNN, to evaluate the model.

[0035] The document search unit 36 ​​searches for a target document to be edited from the group of documents to be edited 35 in response to a search operation from the document display control unit 40 of the user terminal 4. The group of documents to be edited 35 may be documents held by the server 3, or may be documents collected or accessed via the communication network 2. For example, the group of documents to be edited 35 may be documents made public on the web server 5.

[0036] The document search unit 36 ​​transmits the retrieved target document as a search result to the user terminal 4. The user terminal 4 passes the target document transmitted as a search result from the document search unit 36 ​​to the document display control unit 40. The document display control unit 40 transmits a generation instruction to the server 3 instructing the generation of a document, for example, in response to an input operation on the user input unit 41. The generation instruction may include various information related to the generation of the generated sentence, such as the edited portion of the target document, attributes to be assigned to the target document, keywords to be included in the generated sentence, and identification information of characters to be included in the generated sentence. The server 3 receives the generation instruction transmitted from the document display control unit 40 and passes it to the document generation unit 37.

[0037] The document generation unit 37 reads the trained model stored in the trained model storage unit 34 in response to a generation instruction transmitted from the document display control unit 40. The document generation unit 37 generates a document as a generated sentence using the read trained model based on the information included in the generation instruction. The document generation unit 37 may transmit the generated sentence to the user terminal 4 or may store it in the server 3.

[0038] In addition, in Figure 2, the document search unit 36 ​​is shown as being included in the server 3, but this is not limited to this example, and the document search unit 36 ​​may be included in the configuration of another computer that can communicate, for example, via the communication network 2.

[0039] 3 is a block diagram showing an example of a hardware configuration of the server 3 that can be applied to the embodiment. For the sake of explanation, it is assumed here that the server 3 is configured as a stand-alone piece of hardware.

[0040] In FIG. 3, the server 3 includes a CPU (Central Processing Unit) 3000, a ROM (Read Only Memory) 3001, a RAM (Random Access Memory) 3002, a storage device 3003, a data I / F (interface) 3004, and a communication I / F 3005, which are communicatively connected to each other via a bus 3010.

[0041] The storage device 3003 is a non-volatile storage medium such as a flash memory or a hard disk drive. The CPU 3000 operates in accordance with programs stored in the storage device 3003 and the ROM 3001, using the RAM 3002 as a work memory, and controls the overall operation of the server 3.

[0042] The data I / F 3004 controls the transmission and reception of data to and from external devices. The communication I / F 3005 controls communication via the communication network 2.

[0043] The server 3 may further be connected to an input device that accepts user input and a display device that displays information.

[0044] When the above-mentioned learning data acquisition unit 31, learning unit 32, model evaluation unit 33, document search unit 36, and document generation unit 37 are configured in server 3, the CPU 3000 in server 3 executes the information processing program according to the embodiment, thereby configuring these learning data acquisition unit 31, learning unit 32, model evaluation unit 33, document search unit 36, and document generation unit 37, for example, as modules, in the main memory area of ​​RAM 3002.

[0045] The information processing program can be acquired from outside via the communication network 2 by communication via the communication I / F 3005, for example, and installed in the server 3. However, the information processing program is not limited to this, and may be provided by being stored in a removable storage medium such as a CD (Compact Disk), a DVD (Digital Versatile Disk), or a USB (Universal Serial Bus) memory.

[0046] FIG. 4 is a block diagram showing an example of the hardware configuration of the user terminal 4 applicable to the embodiment.

[0047] In Figure 4, the user terminal 4 includes a CPU 4000, a ROM 4001, a RAM 4002, a storage device 4003, a data I / F 4004, a communication I / F 4005, an audio processing unit 4006, and a display control unit 4007, which are connected to each other so as to be able to communicate with each other via a bus 4010.

[0048] The storage device 4003 is a non-volatile storage medium such as a flash memory or a hard disk drive. The CPU 4000 operates in accordance with programs stored in the storage device 4003 and the ROM 4001, using the RAM 4002 as a work memory, and controls the overall operation of the server 3.

[0049] The data I / F 4004 controls transmission and reception of data with external devices. An input device 4020 that accepts user operations may be connected to the data I / F 4004. The input device 4020 may be a pointing device such as a mouse, or a keyboard. The input device 4020 and the display device 4030 may be integrally formed to form a so-called touch panel. The communication I / F 4005 controls communication via the communication network 2.

[0050] The audio processing unit 4006 performs audio signal processing including AD (Analog to Digital) conversion processing and DA (Digital to Analog) conversion processing. For example, the audio processing unit 4006 may be connected to a microphone 4040, convert an analog audio signal output from the microphone 4040 into a digital audio signal by AD conversion processing, and output the digital audio signal to the bus 4010. The audio processing unit 4006 may also convert a digital audio signal supplied via the bus 4010 into an analog audio signal by DA conversion processing, and output the analog audio signal to a speaker 4041 connected to the audio processing unit 4006.

[0051] The display control unit 4007 generates a display signal that can be handled by the display device 4030 connected to the display control unit 4007, based on a display control signal generated by, for example, the CPU 4000. The display device 4030 displays a screen according to the display signal generated by the display control unit 4007.

[0052] If the above-mentioned document display control unit 40 is a program dedicated to the embodiment, the CPU 4000 in the user terminal 4 configures the above-mentioned document display control unit 40, for example, as a module, in the main memory area of ​​RAM 3002 by executing the terminal program related to the embodiment.

[0053] The terminal program can be acquired from an external source, for example, via the communication network 2, by communication via the communication I / F 4005, and installed in the user terminal 4. However, the terminal program may be provided by being stored in a removable storage medium such as a CD (Compact Disk), a DVD (Digital Versatile Disk), or a USB (Universal Serial Bus) memory.

[0054] In addition, when the document display control unit 40 is a program that runs on a browser application installed on the user terminal 4, the browser application may be pre-installed on the user terminal 4 as a general-purpose application.

[0055] (3. Existing Technology) Here, for ease of understanding, existing technology will be described before specific description of the embodiment.

[0056] In recent years, the emergence of large language models (LLMs), such as GPT (Generative Pre-trained Transformer)-3 and ChatGPT, has made it easier to generate stories. However, when generating a story based on the settings of an existing character, the following limitations exist depending on whether the LLM has previously been trained on documents related to that character.

[0057] First, we will explain the limitations that arise when the LLM has previously been trained on documents related to the target existing character.

[0058] Note that here, "learning" refers to learning using MLM (Masked Language Modeling), which is common in LLM, and character settings are not taken into consideration. In the following, unless otherwise specified, "existing characters" will be simply referred to as "characters."

[0059] In this case, existing technologies have limited methods for conveying character settings to generated sentences, and it is not clear what kind of sentences should be given as character settings. Also, in the case of the Transformer model, which is common in LLM, the input length is limited because the cost of the document length that can be input to LLM is squared.

[0060] Furthermore, there was a limitation that the extent to which LLM used documents related to the character during pre-learning was not known to the public except for the developer. Furthermore, because it was not learned from the perspective of the character, even if explanatory documents to explain the character were included in the learning, there was also a limitation that the documents that were produced did not necessarily match the character settings.

[0061] Next, we will explain the limitations that exist when the LLM has not previously studied documents related to the character.

[0062] In this case, additional learning using documents related to character settings is required, but existing technology does not provide a clear learning method for character settings. Also, when dynamically assigning character settings during inference, there is a limitation that costs increase depending on the sequence length. Furthermore, there is a limitation in that the design of the setting documents in this case is not clear.

[0063] On the other hand, there are also limitations when multiple authors participate in writing the same content. For example, when the content is a game scenario, especially a game used on a mobile device (mobile game), it is common for multiple authors to participate in writing the same content. Even in the case of a single author, there are similar limitations when writing a large number of characters or when writing while remembering characters from old content.

[0064] In other words, in these cases, the writer needs to read the setting documents for the content and characters, which limits the time required for production. Also, when multiple writers are involved in the same content, it is difficult to ensure uniformity in the content written by each writer due to the variance in the degree to which the setting documents are read and the writers' skills.

[0065] In an embodiment of the present disclosure, a character representation CharEmb, which is a compressed representation of character settings, is trained, and a story is generated according to the character settings by using the trained model during inference. This reduces the cost according to the sequence length, allows the setting document to be formulated, and reduces the effort required for the writer to read the setting document, thereby eliminating the limitations mentioned above.

[0066] (4. Embodiments of the Present Disclosure) Next, embodiments of the present disclosure will be described in more detail.

[0067] In an embodiment of the present disclosure, two training methods, description-based character encoding and property-conditioned decoding, are used to train the character representation CharEmb, which is a compressed representation of the character settings. In other words, the character representation CharEmb corresponds to the setting document for the character settings.

[0068] In this embodiment, these two learning methods are used in a pipeline (e.g., learning (1) → learning (2) → learning (1) ...) and / or in simultaneous learning to train the character representation CharEmb. In this case, the training of the character representation CharEmb is considered to converge when a specific number of learning steps are performed or when the value of the loss function in the verification data used to verify the training results is minimized.

[0069] (4-1. Model Training Method According to the Embodiment) A model training method according to the embodiment will be described. In the embodiment, a method called description-based character encoding is used to train the model. Here, it is assumed that the character representation CharEmb of a specific existing character (target character) is trained.

[0070] Fig. 5 is a schematic diagram for explaining explanation-based character encoding according to an embodiment. Fig. 6 is a flowchart for explaining an example of processing by explanation-based character encoding according to an embodiment. The following will be described with reference to Fig. 5 along the processing flow of the flowchart in Fig. 6.

[0071] 6, the learning data acquisition unit 31 acquires the explanatory document 110a and the appearance document 120 from the learning document group 30. The explanatory document 110a is a document for explaining the target character. The appearance document 120 is a document in which the target character appears in the story.

[0072] As described above, the learning document group 30 may be a group of documents held by the server 3, or may be documents that are made publicly available to the communication network 2 on the web server 5 and acquired via the communication network 2. Examples of documents made publicly available to the communication network 2 include documents posted on a free encyclopedia site such as Wikipedia, a fan site related to the target character, and a site (such as Fandom (registered trademark)) that provides comprehensive information about characters similar to or in the same category as the target character.

[0073] The learning data acquisition unit 31 may, for example, search for information published on the communication network 2 based on keywords (such as character names) specified in response to user operation from the user terminal 4, and acquire the explanatory document 110a and appearance document 120 of the target character.

[0074] In addition, the learning data acquisition unit 31 may similarly acquire explanatory documents 110b of other characters different from the target character in step S100. The learning data acquisition unit 31 may acquire explanatory documents 110b of other characters in a limited manner in accordance with instructions from the user terminal 4 in response to user operation, or may arbitrarily search for and acquire explanatory documents 110b of other characters related to the target character.

[0075] In the next step S101, the learning data acquisition unit 31 samples sentences to be used for learning the character representation CharEmb from the explanatory document 110a and the appearance document 120 acquired in step S100.

[0076] Taking the explanatory document 110a as an example, the learning data acquisition unit 31 may select sentences at random positions from the acquired explanatory document 110a. In this case, the learning data acquisition unit 31 may select sentences on a sentence-by-sentence or paragraph-by-paragraph basis. Furthermore, for example, the learning data acquisition unit 31 may use an existing NLP (Natural Language Processing) tool to find is-a, has-a, or part-of relationships that represent relationships with the target character, and sample those sentences more than other sentences. This makes it possible to sample sentences that emphasize character setting information.

[0077] Furthermore, for example, the learning data acquisition unit 31 may select sentences to be used for learning the character representation CharEmb from various types of sentences using a bag-of-words or the results of encoding sentences using a universal sentence encoder that is generally used for determining sentence similarity. For example, when a universal sentence encoder is used, various types of sentences can be selected by adopting many sentences whose similarity is lower than a predetermined value, for example, sentences whose encoded vectors are different.

[0078] Furthermore, the training data acquisition unit 31 can also select sentences using existing entity linking technology. Entity linking technology links mentions (e.g., character names) in sentences to a knowledge base. The training data acquisition unit 31 may use entity linking technology to find documents containing the target person's name from the training document set 30 and select them with a certain probability. By using entity linking technology, it is possible to use not only text from sites such as Wikipedia and Fandom (registered trademark) mentioned above, but also data from short message posting sites and manga (speech bubbles), thereby enhancing the training data. The entity linking module used to utilize entity linking technology may be an existing trained one, or it may be retrained in advance if annotated explanatory documents or character appearance documents (e.g., Wikipedia) are available.

[0079] Furthermore, if the time series information for the target character is known, the learning data acquisition unit 31 may sample the same character by dividing it into different periods. For example, a character may be sampled by different periods, such as "up to 15 years old, 15 to 30 years old, 30 years old and over." This allows for cases where the character's characteristics change over time.

[0080] Returning to the explanation of FIG. 6, when the processing by the learning data acquisition unit 31 in step S101 ends, the processing proceeds to step S102.

[0081] The learning unit 32 repeatedly executes the processes of the following steps S102 and S103 until a predetermined condition is satisfied. For example, the learning unit 32 may end the processes of steps S102 and S103 after repeating the processes a specific number of times.

[0082] In step S102, the learning unit 32 embeds a sentence sampled from the explanatory document 110a by encoding it using a DNN 100 (first network, also shown as DNN #1 in FIG. 5 ) to obtain a character representation CharEmb 130a (first numerical representation) in a compressed form. Similarly, the learning unit 32 embeds a sentence sampled from the appearance document 120 by encoding it using a DNN 101 (second network, also shown as DNN #2 in FIG. 5 ) to obtain a character representation CharEmb 140 (second numerical representation) in a compressed form. Note that the character representations CharEmb 130a and 140 are each represented as a list of multiple (e.g., the same number) numerical values, as represented by multiple squares in FIG. 5 .

[0083] If a sentence sampled from the explanatory document 110a or a sentence sampled from the appearance document 120 includes the names of multiple people, the learning unit 32 may sandwich the name of the target character between special tokens and notify the DNN 100 and DNN 101 that the target character is a person of interest. For example, suppose the sentence contains two people, "Akira" and "Rika," and "Akira" is the target person. The special tokens are a pair of tags " <chara>" and "< / chara> "If so, "there <chara> Akira< / chara> The name of the target person "Akira" is enclosed between special tokens, such as "Then Rika appears..."

[0084] However, the learning unit 32 may mask a person's name (for example, "Akira") or a personal pronoun (for example, "he") with a certain probability when the sentence contains such a name. For example, the mask may be implemented by using a special token " <mask>", in the example above, <chara> <mask>< / mask> < / chara> And then Rika appears..."

[0085] By masking the target person's name and personal pronouns indicating the target person, the learning unit 32 can be encouraged to learn the character expression CharEmb from the expressions in the surrounding sentences, rather than from the person's name itself. Also, by masking the names of other people who are not the target person, it is possible to encourage the learning unit 32 to understand the sentence structure in accordance with the character setting.

[0086] A token is a basic unit used when processing text data, and a special token is a token that has a special meaning defined within the text data.

[0087] In the next step S103, the learning unit 32 trains the DNNs 100 and 101 so that the character representation CharEmb 130a encoded into a compressed representation by the DNN 100 and the character representation CharEmb 140 encoded into a compressed representation by the DNN 101 become similar. For example, the learning unit 32 calculates the inner product of the character representation CharEmb 130a and the character representation CharEmb 140, and trains the DNNs 100 and 101 so that this inner product becomes small. Alternatively, the learning unit 32 may train the DNNs 100 and 101 so that the similarity becomes high based on the cosine similarity between the character representation CharEmb 130a and the character representation CharEmb 140.

[0088] Furthermore, the learning unit 32 may embed the explanatory text 110b of a character different from the target character by encoding it using the DNN 100, thereby generating a character representation CharEmb 130b that is a compressed representation. The learning unit 32 may train the DNN 100 so that the character representation CharEmb 130a of the target character and the character representation CharEmb 130b of a character different from the target character become farther apart.

[0089] In step S103, the learning unit 32 encodes the sampled sentences from the explanatory document 110a using the DNN 100, and sets the encoded output as a character expression CharEmb that represents the target character.

[0090] Here, there may be cases where the document to be encoded by the DNN 100 is too long for the encoding capability of the DNN 100, making it difficult for the DNN 100 to encode it. In this case, the learning unit 32 may divide the document to be encoded into text chunks of a length that the DNN 100 can encode (e.g., 512 tokens), and encode each of the text chunks, and the average of each encoding result may be used as the character representation CharEmb.

[0091] In the next step S104, the model evaluation unit 33 uses the character verification model to verify the character representation CharEmb obtained in step S103. For example, the model evaluation unit 33 creates a Q&A (Question and Answer) statement using the is-a, has-a, and part-of relational statements detected during the sampling in step S101. This Q&A statement, along with the character representation CharEmb obtained in step S103, is input into a model such as GPT-3, and the answer is evaluated, thereby obtaining the reliability of the character representation CharEmb.

[0092] Alternatively, the model evaluation unit 33 may prepare an SNN that is smaller in scale than the DNN and use this SNN to evaluate the character representation CharEmb. The model evaluation unit 33 may train this SNN with the character representation CharEmb of a character that is not the target of evaluation (unknown character), and evaluate the character representation CharEmb obtained in step S103 using the character representation CharEmb of the unknown character.

[0093] 7 is a schematic diagram illustrating the evaluation of a character representation CharEmb using an SNN according to an embodiment. The model evaluation unit 33 inputs a verification statement 210 and the character representation CharEmb 130a output from the DNN 100 to an SNN 200 that has been trained using the character representation CharEmb of an unknown character. The verification statement 210 may be automatically acquired using the is-a and has-a relationships used in the sampling in step S101. The model evaluation unit 33 may then add a verification sentence to the original sentence "Akira is a student" by using a personal pronoun, such as "He is a student."

[0094] 7, a Q&A statement, "Is Akira a student?", is input to the SNN 200 as the verification statement 210. The SNN 200 verifies the character representation CharEmb 130a based on this verification statement 210 and outputs a verification result 211. In this example, the SNN 200 outputs a value of "1" if the character representation CharEmb 130a is determined to be correct (Yes), and outputs a value of "0" if it is determined to be incorrect (No).

[0095] If the verification result for the character representation CharEmb 130a is determined to be "correct," the model evaluation unit 33 may terminate the series of processes according to the flowchart of Fig. 6. On the other hand, if the verification result for the character representation CharEmb 130a is determined to be "incorrect," the model evaluation unit 33 may return the process to step S102.

[0096] (4-2. Story Generation Method According to the Embodiment) Next, a story generation method according to the embodiment will be described. In the embodiment, a story is created using a technique called property-conditioned decoding.

[0097] Here, as described above, it is assumed that DNNs 100 and 101 have been trained in learning unit 32, and a character representation CharEmb 130a of a target character (character name "Akira") and a character representation CharEmb 130c of another target character (character name "Rika") have been generated.

[0098] Fig. 8 is a schematic diagram for explaining the characteristic-conditional decoding according to the embodiment. Fig. 9 is a flowchart for explaining the processing by the characteristic-conditional decoding according to the embodiment. The following description will be given with reference to Fig. 8 along with the processing flow of the flowchart in Fig. 9.

[0099] 9 , in the user terminal 4, the document display control unit 40 performs a search operation for the document to be edited on the document search unit 36 ​​in response to a user operation on the user input unit 41, for example. In response to this search operation, the document search unit 36 ​​searches the group of documents to be edited 35 and acquires the document to be edited. The document to be edited by the document search unit 36 ​​is not limited to this, and may be stored locally on the user terminal 4. The document search unit 36 ​​transmits the acquired document to be edited to the document display control unit 40.

[0100] The user terminal 4 transmits a document generation instruction to the server 3 in response to, for example, a user operation on the user input unit 41. This generation instruction may include the document to be edited itself or access information in the edited document group 35 for the document to be edited.

[0101] 9, in step S110, the document generation unit 37 samples sentences from the document to be edited in response to a generation instruction from, for example, the user terminal 4. For the sampling here, for example, the method described in step S101 of FIG.

[0102] In the next step S111, the document generation unit 37 divides the document sampled from the document to be edited in step S110 into two documents, document A and document B. The position at which the document is divided may be arbitrary and is not particularly limited. Here, it is assumed that document A precedes document B.

[0103] In the next step S112, the document generation unit 37 detects people who appear in each of documents A and B. The document generation unit 37 may automatically detect people using, for example, Entity Linking technology, or may present documents A and B to the user terminal 4 and prompt the user to manually detect people who appear in the documents.

[0104] In the next step S113, the document generation unit 37 acquires a KW list 152, which is a list of keywords used in document B. The KW list 152 is a list of vocabulary used in document B. The KW list 152 may include synonyms for words used in document B. In the next step S114, the document generation unit 37 detects attributes of document B. The document generation unit 37 may use an existing classifier such as a Laplace classifier to detect the attributes of document B, or may present document B to the user terminal 4, for example, and prompt the user to manually detect or set attributes.

[0105] 10 is a schematic diagram illustrating an example of an attribute label 153 indicating the attribute detected in step S114 according to the embodiment. As illustrated in Fig. 10, the attribute label 153 is associated with a sentence related to the attribute label 153. The sentence related to the attribute label 153 is not limited to a sentence directly related to the attribute indicated by the attribute label 153, but may also include a context sentence surrounding the sentence.

[0106] It is possible to omit the acquisition of the KW list 152 in step S113 and the detection of the attributes in step S114.

[0107] In the next step S115, the document generation unit 37 inputs document A, a character expression CharEmb of a person appearing in document B (for example, a character with the character name "Akira"), a KW list 152, and an attribute label 153 to the DNN 102 (third network, also shown as DNN#3 in FIG. 8) as input data 150. The document generation unit 37 may further input to the DNN 102 the character string length Length of the generated sentence generated by the DNN 102.

[0108] In the next step S116, the document generation unit 37 applies the model to the input data 150 using the DNN 102 to generate output data 160. The output data 160 is document B' as a document following document A. It is assumed that the DNN 102 has been trained to a certain extent in the initial state.

[0109] In the next step S117, the document generation unit 37 calculates a loss function Loss based on document B' generated by the DNN 102 in step S116. In this embodiment, the document generation unit 37 calculates four loss functions Loss(1) to (4) based on document B' and the KW list 152, attribute label 153, and character representation CharEmb 130a included in the input data 150 and input to the DNN 102. The document generation unit 37 updates the DNN 102 and DNNs 100 and 101 based on the calculated loss functions Loss(1) to (4). The document generation unit 37 may use a known gradient descent method to update the DNNs 100 to 102.

[0110] After the process of step S117, the series of processes according to the flowchart of FIG. 9 ends. Note that document B' generated in step S116 is the final generated sentence. The server 3, for example, presents document B', which is the final generated sentence, to the user terminal 4. In the user terminal 4, the document display control unit 40 may insert document B' presented by the server 3 into a position in the document to be edited, for example, designated by a user operation.

[0111] The processes of steps S110 to S117 described above may be similarly executed for the compressed character representation CharEmb 130c of another target character with the character name "Rika."

[0112] (Loss Functions Loss(1) to (4)) Here, the above-mentioned loss functions Loss(1) to (4) will be explained in more detail.

[0113] Loss Function Loss(1) The loss function Loss(1) is the negative logarithmic likelihood of a generated string (generated sentence) generated by the DNN 102. The loss function Loss(1) is a loss function for training the DNN 102 so that the DNN 102 can generate the original sentence (trained sentence). The document generation unit 37 includes a likelihood calculator 170, which calculates this loss function Loss(1) based on, for example, the following equation (1):

[0114]

[0115] In formula (1), "L()" represents the log likelihood of a sentence composed of elements contained in the parentheses (), and each element "y1 to y M " represents each token included in sentence L(). Equation (1) indicates that calculation is performed for each token in the M-sequence from t=1 to M for the entire M-sequence. The document generation unit 37 uses this equation (1) to train the DNN 102 so as to maximize the probability of predicting the next word from the generated text.

[0116] Loss function Loss(2) The loss function Loss(2) is a loss function for training DNN102 so that the character representation CharEmb input to DNN102 and the character representation CharEmb generated based on the generated sentence generated by DNN102 become closer.

[0117] For example, consider a character representation CharEmb for a character named "Akira" (hereinafter referred to as "Akira's character representation CharEmb" where appropriate). In this case, document generation unit 37 generates character representation CharEmb 131a for Akira's character based on document B', which is a generated sentence generated by DNN 102, using DNNs 100 and 101 that were used when generating character representation CharEmb for Akira included in input data 150.

[0118] The document generation unit 37 trains the DNN 102 so that the character representation CharEmb 130a included in the input data 150 and the character representation CharEmb 131a generated using the DNNs 100 and 101 become closer to each other. For example, the training unit 32 calculates the inner product of the character representation CharEmb 130a and the character representation CharEmb 131a, and trains the DNN 102 so that this inner product becomes smaller. In this case, the loss function Loss(2) may be the inner product of the character representation CharEmb 130a and the character representation CharEmb 131a. Alternatively, the training unit 32 may train the DNN 102 so that the similarity becomes higher based on the cosine similarity between the character representation CharEmb 130a and the character representation CharEmb 131a.

[0119] Furthermore, the document generation unit 37 may train the DNN 102 and the DNNs 100 and 101 so that the character representation CharEmb 130a and the character representation CharEmb 131a become closer to each other.

[0120] Note that the input data 150 may also include character representations CharEmb of multiple people (for example, character representations CharEmb 130a and 130c). In this case, the document generation unit 37 may use the average of the inner products of the character representations CharEmb of the multiple people and the character representations CharEmb generated by encoding, by DNN 100 and 101, sentences generated by DNN 102 based on the character representations CharEmb of the multiple people as the loss function Loss(2).

[0121] Loss Function Loss(3) The loss function Loss(3) is a determination as to whether or not vocabulary (keywords) included in the KW list 152 appears in document B', which is a sentence generated by the DNN 102, when the KW list 152 is included in the input data 150. The loss function Loss(3) is a loss function for training the DNN 102 so that the vocabulary included in the KW list 152 is reflected in sentences generated by the DNN 102.

[0122] The document generation unit 37 includes a KW appearance determiner 172 that compares the vocabulary contained in document B' with the KW list 152, and uses this KW appearance determiner 172 to determine whether or not the vocabulary contained in the KW list 152 appears in document B'. At this time, the KW appearance determiner 172 may provide variation to the generated words by using a thesaurus or a known, trained Word2vec, etc. Furthermore, the KW appearance determiner 172 may make a determination using the inner product distance in the token embedding output by the DNN 102.

[0123] Loss Function Loss(4) The loss function Loss(4) is a cross-entropy loss related to sentence label prediction. The loss function Loss(4) is a loss function for training the DNN 102 so that the DNN 102 can generate sentences that match the attribute labels 153 included in the input data 150.

[0124] The document generation unit 37 includes an attribute classifier 171, such as a known Laplace classifier, and classifies document B', which is a sentence generated by the DNN 102, based on the attributes included in the input data 150, and determines whether or not the attribute indicated by the attribute label 153 is reflected in document B'. For example, if the attribute label 153 indicates the attribute "battle, tension," the attribute classifier 171 classifies the sentences included in document B' according to the sentence attributes, and detects sentences that reflect the attribute "battle, tension."

[0125] The document generation unit 37 trains the DNNs 100, 101, and 102 using, for example, gradient descent, based on the loss functions Loss(1) to (4) calculated as described above, so that the loss function Loss shown in the following equation (2) is minimized. In equation (2), the coefficients α, β, and γ are hyperparameters, and their values ​​are set by, for example, a learning developer. Loss=Loss(1)+α×Loss(2)+β×Loss(3)+γ×Loss(4) ... (2)

[0126] In existing technology, DNNs 100 to 102 are trained using only loss function Loss(1). However, this method of existing technology can make it difficult to reflect the controllability of character settings, KW list 152, and attribute label 153. In the embodiment, in addition to loss function Loss(1), loss functions Loss(2) to (4) are used to reinforce the learning of DNNs 100 to 102, making it easier to reflect the controllability.

[0127] It is possible that the document A included in the input data 150 does not contain any characters. Therefore, a character expression CharEmb indicating that the document A does not contain any characters may be defined in advance.

[0128] (4-3. User Interface According to Embodiment) Next, a user interface in the user terminal 4 according to the embodiment will be described.

[0129] 11 is an example flowchart illustrating the overall flow of document generation processing according to an embodiment. In FIG. 11, the left side (steps S200 to S211) shows processing in the user terminal 4, and the right side (steps S300 to S306) shows processing in the server 3. The processing in the user terminal 4 shown in FIG. 11 is mainly executed via a user interface (UI) provided by the user terminal 4 through display on the display unit 42 and input operations on the user input unit 41.

[0130] In step S200, the document display control unit 40 of the user terminal 4 acquires a document to be edited. For example, in the user terminal, the document display control unit 40 performs a search operation on the document search unit 36 ​​to search for a desired document in response to a user operation on the user input unit 41. In response to the search operation, the document search unit 36 ​​acquires, for example, information indicating the document to be edited from the group of documents to be edited 35 and returns the information as a search result to the document display control unit 40. The document display control unit 40 acquires, for example, the desired document to be edited from the group of documents to be edited 35 based on the search result passed from the document search unit 36. The user terminal 4 may display the acquired document to be edited on the display unit 42.

[0131] In the next step S201, the document display control unit 40 in the user terminal 4 determines a portion of the document to be edited acquired in step S200 where a sentence is to be generated. Fig. 12 is a schematic diagram showing an example of a UI screen for determining a portion of the document to be edited where a sentence is to be generated, according to an embodiment.

[0132] 12, the document display control unit 40 displays a group of documents to be edited, 3001, 3002, 3003, ..., on the display unit 42. The documents to be edited 3001, 3002, 3003, ... may be documents made public by the Web server 5 and obtained via the communication network 2, or may be documents held by the server 3 or the user terminal 4.

[0133] A user operating the user terminal 4 inputs information to the user input unit 41 to select a desired document from the documents to be edited 3001, 3002, 3003, .... In the example of Fig. 12, the input to select the desired document is made using natural language, as shown by "Where is the place where Tadashi is in trouble?" The input may be text input or voice input.

[0134] In response to the input in step S400, the document display control unit 40 selects the document to be edited 300 from the documents to be edited 3001, 3002, 3003, . . . in step S401. x Select .

[0135] At this time, the document display control unit 40 displays the document 300 to be edited in the example of FIG. x The reason for selecting the document to be edited 300 x The document to be edited 300 is shown by highlighting 301a and 301b. x The reason for the selection is that "Chapter 72 depicts the sentence 'Tadashi has been killed!' and his injuries," and a message is presented on the display unit 42 or as a sound.

[0136] The user who operates the user terminal 4 reads the document 300 to be edited that is presented in step S401. x In response to the message, in step S402, the presented document to be edited 300 x In order to determine the part of the sentence in the sentence list where the user wants to generate a sentence, the user may query the system using voice or text input. Alternatively, the user may determine the part of the sentence by operating the cursor 380, or, if the user terminal 4 has a touch panel, may determine the part of the sentence by touch input.

[0137] The document display control unit 40 also displays, for example, the document 300 to be edited. x 13 is a schematic diagram illustrating an example of visualizing and presenting the relationships between multiple characters according to an embodiment.

[0138] 13, multiple characters are displayed by icons 310a to 310d as identification information for identifying each character, and the relationships between the characters are indicated by arrows and words indicating the relationships (in the example shown, "protect," "trust," "cooperate," "obstruct," and "exclude"). Furthermore, as shown in the lower right of FIG. 13, the relationships and background explanations of the characters may be indicated using sentences.

[0139] For example, the document display control unit 40 can visualize the relationships between characters by classifying each character using character information preset for each character and a trained model. The character information for each character may be, for example, a character representation CharEmb acquired for each character.

[0140] For example, in response to an input such as "What is the relationship between character A and character B?", the learning model outputs "cooperation" as the relationship between character A and character B. Furthermore, in response to an input of character information for each of character A and character B, the learning model may output a value of "3" indicating cooperation at the class level.

[0141] 14 is a schematic diagram showing another UI screen for determining a portion for which a sentence is to be generated, according to an embodiment. The example in FIG. 14 shows an example of a UI (scenario viewer) that provides an overview of a story (scenario) in chronological order.

[0142] In FIG. 14 , section (a) shows, for example, an overview of the entire document selected for editing, with the horizontal axis representing the passage of time. Scenes are also classified by the corresponding explanatory text. In the example of section (a), three scenes are classified using different fill patterns, including periods when no fill pattern is present. For example, the document display control unit 40 may use a trained model to input chapter descriptions included in the document to be edited and output scene classifications such as "harmonious," "hostile," "neutral," and "positive / negative."

[0143] The document display control unit 40 may also perform text analysis on the document to be edited using a trained model or the like to list the characters. The document display control unit 40 may arrange the listed characters in chronological order at positions according to the time of their appearance using corresponding icons 310, and display them on the display unit 42. Note that the time of appearance of a character may be expressed in chronological order within the content of the document to be edited, or in chronological order related to the release of the content (such as when it was published in a magazine or aired on TV).

[0144] 14, section (b) shows a state in which the specified range in the time series of section (a) is zoomed in with respect to time. For example, if the user terminal 4 has a touch panel, the zooming process may be performed by pinching out at a corresponding position on the screen.

[0145] By zooming in on a specific time period in this way, the document display control unit 40 can display information that would be hidden in a bird's-eye view of the entire document. Furthermore, the document display control unit 40 can use the zoom display to display information indicating the attribute label 153 (the icon 312 indicating "battle") and character dialogue 311. The document display control unit 40 may also display a summary of the chapter description corresponding to the zoomed-in time period in the document being edited. The document display control unit 40 may also search for scenes, etc., in the chronological display using text input or voice input. The document display control unit 40 may accept information specifying a specific position in the chronological order, such as "only battle scenes" or "the part where character A is talking," input using text input or voice input.

[0146] 14, section (c) is an example of a UI screen 390 showing edited portions specified in chronological order in the document to be edited. The document display control unit 40 may, for example, determine the portion specified by the zoom process in section (b) in the document to be edited as the edited portion. In the example of section (c), keywords included in the edited portion are shown in bold in the text of the document to be edited ("Akira Kato", "Tadashi Oda", "Star Light Vessel", "curse ... R", "Dark Vessel Association").

[0147] In this way, by looking at the story in chronological order, it is possible to quickly find the part for which you want to generate a sentence.

[0148] 15 is a schematic diagram illustrating yet another UI screen for determining a section for which a sentence is to be generated, according to an embodiment. The example in FIG. 15 illustrates an example in which a chapter description is divided into units that are meaningful in terms of the development of the story, such as paragraphs, and then visualized together.

[0149] In FIG. 15, section (a) is the document to be edited 300. x The document to be edited 300 shown in section (a) of FIG. x For example, the document to be edited 300 shown in FIG. x and the document to be edited 300 x The reason for selecting the document 300 to be edited is x 3. The diagram is shown with highlighting 301a and 301b for .

[0150] Section (b) of FIG. 15 shows the document 300 to be edited in section (a). x 1 shows an example of a UI screen 320 that displays the chapter description written in the document 300 to be edited, divided into paragraphs. x The document display control unit 40 divides the sentences contained in the document 300 into units that are meaningful in terms of the development of the story, such as paragraphs, and creates summaries for each of the divided paragraphs. For example, the document display control unit 40 learns a model for dividing plain text into paragraphs and scenes based on the paragraph structure (blank lines, etc.) contained in the original document and manual annotations, and uses this model to divide the document 300 into paragraphs and scenes. x The chapter descriptions provided in may be divided into paragraphs.

[0151] The document display control unit 40 displays, on the UI screen 320, a summary that summarizes each divided paragraph and an icon 310 that indicates the characters that appear in that paragraph, in association with each other. By looking at the summary of each paragraph and the icons 310 of the characters that appear in each paragraph, the user can grasp the general flow of the story. x Since the flow of the story can be understood without reading the chapter explanations described in the text, the burden on the user is reduced and the desired sentence generation position can be quickly found.

[0152] As yet another UI for determining the portion of the text for which a sentence is to be generated according to the embodiment, the summary and character information shown in FIG. 15 may be directly overwritten and displayed on a web page or text view screen.

[0153] 16 is a schematic diagram showing an example of yet another UI screen 330 for determining a portion from which a sentence is to be generated, according to the embodiment. In the example of FIG. 16, x A UI screen 321 including a summary and icons 310 indicating the characters is overwritten and displayed on the lower half of the document 300. x By displaying these together, the user can more easily grasp the summary and the characters, which reduces the burden on the user and enables the user to quickly find the desired sentence generation position.

[0154] 11 , when the location where a sentence is to be generated is determined in step S201, the document display control unit 40 in the user terminal 4 transmits the document to be edited and information indicating the location where the sentence is to be generated in the document to be edited to the server 3 in step S202. In step S300, the server 3 acquires the document to be edited and the information indicating the location where the sentence is to be generated transmitted from the user terminal 4.

[0155] When the information is transmitted in step S202 to the user terminal 4, the process proceeds to step S203. In step S203, the document display control unit 40 in the user terminal 4 determines, using a UI screen (not shown), whether or not the user wishes to include the character in the generated sentence. On the other hand, if the document display control unit 40 determines, in response to the user's operation on the UI screen, that the user does not wish to include the character in the generated sentence ("No" in step S203), the process skips steps S204 and S205 and proceeds to step S206.

[0156] On the other hand, if the document display control unit 40 determines, in response to a user operation on the UI screen, that a character is to be included in the generated sentence (step S203, "Yes"), the process proceeds to step S204.

[0157] In step S204, the document display control unit 40 in the user terminal 4 selects a character to be included as a character in the game in response to a user operation on a UI screen (described later). The document display control unit 40 transmits information indicating the selected character to the server 3 (step S205).

[0158] In the next step S206, the document display control unit 40 in the user terminal 4 specifies keywords (shown as "KW" in the figure) of the sentence to be generated in response to user operations on a UI screen described later, and creates a KW list 152. In the next step S207, the document display control unit 40 in the user terminal 4 specifies attributes of the sentence to be generated in response to user operations on a UI screen described later, and creates an attribute label 153. In the next step S208, the document display control unit 40 in the user terminal 4 sets the sentence to be continued in response to user operations on a UI screen described later.

[0159] The processing order of steps S206 to S208 is not limited to the order shown in the flowchart of FIG. 11, and may be performed in another order.

[0160] In the next step S209, the document display control unit 40 in the user terminal 4 transmits the KW list 152, the attribute label 153, and the continuation sentence created by the processing in steps S206 to S208 to the server 3, and issues a document generation instruction to the server 3.

[0161] 17 is a schematic diagram showing an example of a UI screen for inputting information for document generation according to an embodiment. The character selection in step S204, the specification of keywords and attributes in steps S206 and S207, and the setting of a sentence for which a continuation is to be generated in step S208 are performed in response to user operations on a UI screen 350 shown in section (b) of FIG.

[0162] Section (a) of Fig. 17 shows an example of input data 150 input to the DNN 102. In the example of section (a), the input data 150 includes a peripheral sentence 151 as a sentence for which a continuation is to be generated, a character expression CharEmb 130a for the character name "Akira", a KW list 152, an attribute label 153, and a character string length Length 154 of the generated sentence generated by the DNN 102. The DNN 102 (also referred to as DNN#3 in the figure) applies a trained model to the input data 150 and outputs output data 161.

[0163] Section (b) of FIG. 17 shows an example of the UI screen 350 described above. In the UI screen 350, the peripheral sentence 151 is set in area 3501 (step S208 of FIG. 11 ). The document display control unit 40 may extract the peripheral sentence 151 from the document to be edited, for example, as described in step S101 of FIG. 6 , and input the extracted peripheral sentence 151 into area 3501. Alternatively, area 3501 may be a blank space into which text can be entered, or the user may manually enter the peripheral sentence 151. Furthermore, the document display control unit 40 may attach a text file to area 3506 by, for example, a so-called drag-and-drop operation, and input the contents of the attached text file into area 3501.

[0164] In the UI screen 350, an area 3502 includes icons 310 representing characters for which character representations CharEmb have been generated. An area 3503 is an area for arranging the icons 310 of characters to be included as characters. The document display control unit 40 can select the icon 310 of a character to be included as a character from the icons 310 included in area 3502, for example, by controlling a cursor 3520 in response to a user operation, and move (drag-and-drop) the selected character to area 3503, thereby selecting the character as a character (step S204 in FIG. 11 ).

[0165] It is also possible to select a character from content other than the content related to the currently edited document as a character. For example, by selecting an icon 310' representing a character from the different content and moving it to area 3503, the character represented by icon 310' can be selected as a character. It is assumed that a character representation CharEmb of the character represented by icon 310' is generated in advance and associated with icon 310'.

[0166] In the UI screen 350, an area 3504 is an area for specifying the attribute label 153 (step S207 in FIG. 11 ). The document display control unit 40 specifies, for example, an attribute checked in response to a user operation as the attribute label 153. The information on the selectable attributes in the area 3504 may be editable. In the UI screen 350, an area 3507 is an area for inputting keywords to be included in the KW list 152 (step S206 in FIG. 11 ). The user may manually input keywords into the area 3507, or words extracted from the document to be edited may be input into the area 3507 as keywords. In the UI screen 350, an area 3508 is an area for inputting the character string length Length 154 of the generated sentence.

[0167] On the UI screen 350, a button 3509 is a button for instructing the server 3 to generate a generated sentence (step S209 in FIG. 11 ). By operating this button 3509, the keyword list 152, the attribute label 153, the continuation sentence (peripheral sentence 151), and a document generation instruction are transmitted from the user terminal 4 to the server 3.

[0168] In the UI screen 350, an area 3510 displays a generated sentence generated in response to the operation of button 3509 as the output of DNN 102. Furthermore, a cursor 3530 can be used to specify an insertion position for the generated sentence in response to a user operation for the sentence displayed in area 3501. The document display control unit 40 may edit the document by, for example, inserting the generated sentence displayed in area 3510 at the position of cursor 3530 in the sentence displayed in area 3501.

[0169] Returning to the description of FIG. 11, the server 3 acquires, via the document generation unit 37, information indicating the selected character transmitted from the user terminal 4 in step S205 (step S301).

[0170] In the next step S302, the server 3 determines, by the document generation unit 37, whether or not the character indicated in the information acquired in step S301 is a learned character using the DNNs 100 and 101. In other words, in step S302, the server 3 determines, by the document generation unit 37, whether or not the character acquired in step S301 is a character for which a character representation CharEmb has been generated. If the document generation unit 37 determines in step S302 that the character indicated in the information acquired in step S301 is a learned character ("Yes" in step S302), the process proceeds to step S304.

[0171] On the other hand, if the document generating unit 37 determines in step S302 that the character is not a learned character (step S302, "No"), the process proceeds to step S303.

[0172] In step S303, the document generation unit 37 in the server 3 causes the learning unit 32 to generate a character representation CharEmb for the character acquired in step S301. In the next step S304, the learning unit 32 in the server 3 associates the generated character representation CharEmb with the character and sets the character representation CharEmb.

[0173] In the next step S305, the document generation unit 37 in the server 3 generates a generated sentence in accordance with the document generation instruction transmitted from the user terminal 4 in step S209. That is, the document generation unit 37 applies the KW list 152, attribute label 153, and subsequent sentence (peripheral sentence 151) transmitted from the user terminal 4 together with the document generation instruction in step S209, and the character expression CharEmb set in step S304 to the trained model by the DNN 102, thereby generating a generated sentence.

[0174] In the server 3, the document generation unit 37 transmits the generated sentence generated in step S305 to the user terminal 4 (step S306).

[0175] In step S210, the document display control unit 40 of the user terminal 4 causes the display unit 42 to display the generated sentence transmitted from the server 3 in step S306. For example, the document display control unit 40 causes the generated sentence to be displayed in area 3510 on the UI screen 350 shown in section (b) of FIG. 17. However, the document display control unit 40 may also display the generated sentence on a screen other than the UI screen 350.

[0176] In the next step S211, the document display control unit 40 in the user terminal 4 determines whether the generated sentence has been rewritten. The document display control unit 40 may determine whether the generated sentence has been rewritten, for example, in response to a user operation. If the document display control unit 40 determines that the generated sentence has not been rewritten (step S211, "Yes"), it ends the series of processes according to the flowchart in FIG. 11.

[0177] On the other hand, if the document display control unit 40 determines in step S211 that the generated sentence has been rewritten (step S211, "No"), it may return the processing to one of steps S200 to S208 and perform setting processing etc. again.

[0178] (UI screen allowing review according to variations in generated sentences) The generated sentences generated by the DNN 102 in step S305 may contain variations in expression. In this case, the DNN 102 may generate multiple generated sentences according to variations in expression. Figure 18 is a schematic diagram showing an example of a UI screen allowing review of generated sentences according to variations in expression, according to an embodiment.

[0179] In Fig. 18 , a UI screen 360 uses a tree diagram 361 to represent generated sentences that include variations in expression. The tree diagram 361 has a root on the left side and leaves on the right side. In the example of Fig. 18 , generated sentences branch out sequentially from the root of the tree diagram 361 according to variations, and each node displays a sentence resulting from the variations. In addition, for each leaf of the tree diagram 361, a summary of the sentence generated by tracing the nodes from the root to the leaf is displayed, along with an icon 310 indicating the characters that appear in the sentence. By looking at the summary and characters on this UI screen 360, the user can easily grasp the variations in generated sentences due to variations.

[0180] Furthermore, in response to a user operation, the document display control unit 40 can insert a character representation CharEmb into the middle of the tree diagram 361. In the example shown, the document display control unit 40 controls a cursor 363 in response to a user operation to move an icon 310 linked to the character representation CharEmb to be inserted to the position in the tree diagram 361 where the character representation CharEmb is to be inserted. This makes it possible to add the character linked to the character representation CharEmb as a character in the sentence along that path.

[0181] Furthermore, the document display control unit 40 may generate one sentence by connecting the sentences of the nodes on the path 362 by tracing each node and edge in order from the root to the leaf, as shown as a path 362 in Fig. 18, in response to a user operation. The generation of sentences in the tree diagram 361 is not limited to a tracing operation as in the path 362, but may also be performed by sequentially specifying nodes and edges by tapping or the like.

[0182] The generated sentence may be displayed, for example, in area 364a of UI screen 360. In the illustrated example, a portion of the sentence displayed in area 364a is corrected, and the corrected sentence is displayed in area 364b. In this example, the corrected sentence displayed in area 364b is reflected in the script (for example, the story).

[0183] According to the UI screen 360 shown in FIG. 18, the user can easily grasp the variations due to variations in the generated sentence, and can intuitively perform the process of generating a single sentence from the variations due to variations in the generated sentence.

[0184] (5. Effects of the Embodiments of the Present Disclosure) Next, effects of the embodiments of the present disclosure will be described in comparison with existing technologies.

[0185] The first method for generating a story according to the character settings using existing technology is as follows.

[0186] Existing technology, such as GPT-3, is trained using explanatory documents and documents in which the target character appears. More specifically, the LLM is trained using text as the input, with the character name and the sentence to be continued, and the output being the continuing document.

[0187] This first method had the following two problems: 1. Because learning is not based on character setting documents, the sentences used for learning do not necessarily conform to the character settings. 2. It is not possible to generate unknown characters that are not included in learning.

[0188] In response to this, by applying the embodiments of the present disclosure to story generation, it is possible to address the first problem by using a setting document to pre-learn a character representation CharEmb, which is a compressed representation of the target character, thereby ensuring that the document conforms to the setting. Furthermore, regarding the second problem, because DNNs 100 and 101 learn to extract settings from setting documents, it is possible to generate character representations CharEmb for unknown characters without learning.

[0189] Furthermore, the following is a second method for generating a story according to the character settings using existing technology.

[0190] A story is generated by inputting a character description document as a prompt into an LLM based on existing technology. More specifically, a setting document for the target character is added. In other words, the input is a "setting document" and a "sentence to explain the continuation," and the output is a "continuation document."

[0191] This second method had the following two problems: 1. How to create a setting document is not obvious (it is not formulated). 2. The setting document needs to include a lot of character information about the target character (traits, tone of voice, affiliation, etc.), but a typical LLM model requires memory costs that are the square of the input string length, making it difficult to input sentences that are too long.

[0192] In contrast, by applying an embodiment of the present disclosure to story generation, the first problem can be solved by using a compressed representation called CharEmb, which is a character representation learned in advance from various documents, so the user only needs to select the character. Furthermore, the second problem can be solved by eliminating the need to design a setting document each time. Therefore, there is no need to input long sentences with a lot of character information for the target character, which reduces memory costs.

[0193] (6. Modifications of the Embodiments of the Present Disclosure) Next, a modification of the embodiments of the present disclosure will be described. In the above-described embodiments, the generated sentence is generated based on the document to be edited. In contrast, in a modification of the embodiment, the generated sentence is generated using a modal different from the document. As an example of a modal different from the explanatory document and the appearance document, it is possible to apply an image of a cartoon.

[0194] In the modified embodiment, the character representation CharEmb of the character is also generated based on the description document and appearance document related to the character using the description-based character encoding described in the embodiment.

[0195] 19 is a schematic diagram illustrating an example of generating a generated sentence using an image of a comic, according to a modified example of the embodiment. In FIG. 19, an image of one frame of a comic is displayed in an area 3701 of a UI screen 370. The image to be displayed in an area 3710 may be manually selected and acquired by the user, for example. Also, an image of a character A and an image of a character X are included in the area 3710. Character X is speaking in a speech bubble 3703.

[0196] In the UI screen 370, an area 3702 includes icons 310 each representing a character for which a character representation CharEmb has been generated. For example, in the user terminal 4, the document display control unit 40 selects the icon 310 of the character to be linked to character X from the icons 310 included in area 3702, for example, by controlling a cursor 3730 in response to a user operation, and moves the icon 310 to the position of the image of character X included in area 3701 (drag-and-drop). This links the character representation CharEmb corresponding to the icon 310 to character X. This allows the behavior of character X to be based on the linked character representation CharEmb.

[0197] Linking of character representations CharEmb to characters A and X included in area 3701 is not limited to the method using icon 310. For example, a character representation CharEmb may be generated by the description-based character encoding described in the embodiment using document 3720a including an explanatory document and an appearance document for explaining character A, and linked to character A. Similarly, a character representation CharEmb may be generated for character X using document 3720b including an explanatory document and an appearance document for explaining character X, and linked to character A.

[0198] The method of linking the character representation CharEmb to the characters A and X included in the area 3701 is not limited to these. For example, character information may be conditioned for each of the characters A and X, and each of the conditioned character information may be used as an explanatory document for each of the characters A and X.

[0199] For example, character information may be assigned to each of characters A and X by image recognition of the image in area 3701, or character information may be assigned manually in response to a user operation. Also, scene recognition processing may be performed on the image in area 3701, and character information may be assigned to each of characters A and X in accordance with the recognized scene (outdoors, behind the school, etc.).

[0200] The document display control unit 40 may add a speech bubble 3704 in response to a user operation. For example, the document display control unit 40 may control the position of the cursor 3705 in response to a user operation to move the position of the speech bubble 3704. The document display control unit 40 may cause the added speech bubble 3704 to display a line from a character (character A in this example) located at a position corresponding to the speech bubble 3704. This processing is equivalent to processing that uses the character representation CharEmb of the character.

[0201] Furthermore, for the image displayed in area 3701, the pages immediately before and / or after the page containing the frame containing that image may be further used to condition the character information of each of characters A and X. Information on these pages may be obtained as image data, and dialogue may also be obtained using, for example, an OCR (Optical Character Recognition / Reader) (in the case of paper media) or character recognition. In this way, by using information on the immediately before and / or after pages, it is possible to make the flow of the story more natural.

[0202] As described above, based on the image in area 3701, the character information of each of characters A and X included in the image is conditioned, and character representations CharEmb of each of characters A and X are generated based on the character information of each of characters A and X. In server 3, document generation unit 37 may generate a story (generated text) based on the character representations CharEmb generated in this way.

[0203] In the server 3, the document generation unit 37 may embed lines included in the generated story, for example, by overwriting the lines in the speech bubble 3703 in the area 3701. At this time, the document generation unit 37 may determine the number of characters of the lines to be embedded, taking into consideration the size of the speech bubble 3703 and the font size.

[0204] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0205] Note that the present technology can also be configured as follows. (1) A learning device including: a first network for acquiring a first numerical expression by embedding based on an explanatory document for explaining a character, which is included in an existing document; and a second network for acquiring a second numerical expression by embedding based on a character appearance document, which is included in the existing document, in which the character appears in a story; and a learning unit that trains the first numerical expression and the second numerical expression so that the first numerical expression and the second numerical expression become similar to each other, and outputs the first numerical expression learned by the learning unit as a character numerical expression representing the character. (2) The learning device according to (1), wherein the learning unit acquires the existing document from a document published on the Internet. (3) The learning device according to (1), wherein the learning unit acquires the existing document from a document disclosed via a communication network. (4) The learning device according to any of (1) to (3), wherein the learning unit detects a word contained in the existing document that indicates a relationship with the character, selects more sentences containing the word from the existing document than sentences that do not contain the word, and extracts the explanatory document and the appearance document from the existing document based on the selected sentences. (5) The learning device according to any of (1) to (3), wherein the learning unit extracts the explanatory document and the appearance document from the existing document by selecting from the existing document sentences of different types to sentences containing the character contained in the existing document. (6) The learning device according to (5), wherein the learning unit extracts the explanatory document and the appearance document from the existing document based on sentences contained in the existing document that have a similarity lower than a predetermined value. (7) The learning device according to any one of (1) to (3), wherein the learning unit detects the name of the character from the existing document, selects a sentence containing the name from sentences included in the existing document with a certain probability, and extracts the explanatory document and the appearance document from the existing document.(8) The learning device according to any of (1) to (7), wherein, when the existing document includes time series information, the learning unit trains the first network and the second network for each period indicated in the time series information. (9) The learning device according to any of (1) to (8), wherein the learning unit further trains the first network so that a third numerical representation obtained using the first network by embedding based on an explanatory document for describing a character different from the character, which is obtained from the existing document, becomes distant from the first numerical representation. (10) The learning device according to any of (1) to (9), wherein, when the explanatory document and / or the appearance document include names of multiple people including the name of the character, the learning unit assigns a marker to the name of the character. (11) The learning device according to any of (1) to (10), wherein, when the explanatory document and / or the appearance document contains words that identify a person, the learning unit masks the words that identify the person with a certain probability. (12) The learning device according to any of (1) to (11), wherein the learning unit verifies the trained first numerical expression using a fourth network that is trained with a numerical expression different from the first numerical expression. (13) A generation device comprising: a first network that performs embedding based on an explanatory document to describe a character, and a second network that performs embedding based on an appearance document in which the character appears in a story, and a generation unit that generates a document that is a generated sentence using a third network based on the first numerical expression generated using these and a target document to be edited, wherein the generation unit trains the third network using a loss function based on the generated sentence. (14) The generating device according to (13), wherein the first numerical representation is acquired by training the first network and the second network so that the first numerical representation and a second numerical representation acquired by the second network become close to each other.(15) The generation device according to (13) or (14), wherein the generation unit divides a target document to be edited into a first document and a second document, and generates the generated sentence by inputting the first numerical representation of a person appearing in at least one of the first document and the second document and the first document to the third network. (16) The generation device according to (15), wherein the generation unit trains the third network using the loss function including a first loss function based on a negative log-likelihood of a character string of the generated sentence. (17) The generation device according to (16), wherein the generation unit embeds the generated sentence using the first network and the second network to obtain a third numerical representation, and trains the third network further using a second loss function calculated based on the first numerical representation and the third numerical representation. (18) The generation device according to (16) or (17), wherein the generation unit generates the generated sentence using the third network further based on a given keyword, and trains the third network further using a third loss function based on whether the keyword appears in the generated generated sentence. (19) The generation device according to any of (16) to (18), wherein the generation unit detects an attribute related to the character from the second document, and trains the third network further using a fourth loss function based on whether the generated generated sentence includes an expression corresponding to the attribute.(20) A terminal device comprising: an input unit that accepts user operations; and a display control unit that controls a display screen displayed on a display device, wherein the display control unit generates a display control signal for causing the display device to display a user interface screen including: a document area that displays a generated sentence generated based on a numerical expression representing the character generated using: a first network that performs embedding based on an explanatory document that explains a character; and a second network that performs embedding based on a character appearance document in which the character appears in a story; and a target document that is a document to be edited; and a character selection area for selecting a character to appear in the generated sentence displayed in the document area, and places identification information associated with the numerical expression in the character selection area. (21) The terminal device described in (20), wherein the display control unit further includes, in the user interface screen: an attribute designation area that designates an attribute to be assigned to the generated sentence. (22) The terminal device described in (20) or (21), wherein the display control unit further includes, in the user interface screen: a keyword designation area that designates a keyword to be included in the generated sentence. (23) The terminal device according to any of (20) to (22), wherein the display control unit sets a position at which the generated sentence generated using the numerical expression is to be inserted into the document displayed in the document area in response to a user operation. (24) The terminal device according to any of (20) to (23), wherein the display control unit generates the display control signal for displaying a tree diagram corresponding to a variation of the generated sentence in generating the generated sentence using the numerical expression on the user interface screen. (25) The terminal device according to (24), wherein the display control unit generates the generated sentence based on a path through the tree diagram in response to a user operation of tracing the nodes and edges of the tree diagram in order. (26) The terminal device according to (24) or (25), wherein the display control unit generates the display control signal for displaying, for each leaf of the tree diagram, a summary that summarizes documents generated along a path from the root of the tree diagram to the leaf.(27) The terminal device according to any of (24) to (26), wherein the display control unit generates the display control signal for displaying, for each leaf of the tree diagram, identification information that identifies characters appearing in documents generated along a path from the root of the tree diagram to the leaf. (28) The terminal device according to any of (20) to (23), wherein the display control unit generates the display control signal for displaying the target document in chronological order based on chronological order information related to the target document. (29) The terminal device according to (28), wherein the display control unit generates the display control signal for enlarging and displaying a period specified in accordance with the user operation in the chronological order, and determines a portion of the target document included in the period as a document to be used for generating the generated sentence. (30) The terminal device according to any of (20) to (23), wherein the display control unit generates the display control signal for dividing the target document into units that are meaningful in terms of the development of a story and displaying them, and displaying summaries of sentences included in the units and identification information that identifies characters in the units in association with the units.

[0206] REFERENCE SIGNS LIST 1 Information processing system 2 Communication network 3 Server 5 Web server 6 Cloud network 30 Training document group 31 Training data acquisition unit 32 Training unit 33 Model evaluation unit 34 Training model storage unit 35 Editing target document group 36 Document search unit 37 Document generation unit 40 Document display control unit 41 User input unit 42 Display unit 100, 101, 102 DNN 110a, 110b Explanatory document 120 Appearance document 130a, 130b, 130c, 140 Character expression CharEmb 150 Input data 151 Peripheral sentence 152 KW list 153 Attribute label 154 Character string length Length 160, 161 Output data 170 Likelihood calculator 171 Attribute classifier 172 KW appearance determiner 200 SNN 3001, 3002, 3003, 300 x Documents to be edited 310, 310', 310a, 310b, 310c, 310d, 312 Icons 320, 321, 330, 350, 360, 370, 390 UI screen 361 Tree diagram< / mask>

Claims

1. A learning device comprising: a first network for obtaining a first numerical expression by embedding based on an explanatory document for explaining a character, which is included in an existing document; a second network for obtaining a second numerical expression by embedding based on a character appearance document, which is included in the existing document, in which the character appears in the story; and a learning unit that trains the first numerical expression and the second numerical expression so that they become closer to each other; and outputs the first numerical expression learned by the learning unit as a character numerical expression representing the character.

2. The learning device according to claim 1, wherein the learning unit acquires the existing documents from documents published on the Internet.

3. The learning device according to claim 1, wherein the learning unit acquires the existing documents from documents disclosed via a communication network.

4. The learning device of claim 1, wherein the learning unit detects words contained in the existing document that indicate a relationship with the character, selects more sentences containing the words from the existing document than sentences not containing the words, and extracts the explanatory document and the appearance document from the existing document based on the selected sentences.

5. The learning device of claim 1, wherein the learning unit extracts the explanatory document and the appearance document from the existing document by selecting a different type of sentence from the existing document relative to a sentence indicating the character contained in the existing document.

6. The learning device according to claim 5, wherein the learning unit extracts the explanatory document and the featured document from the existing document based on sentences contained in the existing document that have a similarity lower than a predetermined value.

7. The learning device of claim 1, wherein the learning unit detects the name of the character from the existing document, selects sentences containing the name from sentences contained in the existing document with a certain probability, and extracts the explanatory document and the appearance document from the existing document.

8. The learning device according to claim 1, wherein the learning unit, when the existing document includes time-series information, trains the first network and the second network for each period indicated in the time-series information.

9. The learning device of claim 1, wherein the learning unit further trains the first network so that a third numerical representation obtained using the first network by embedding based on an explanatory document for explaining a character different from the character, obtained from the existing document, becomes distant from the first numerical representation.

10. The learning device according to claim 1, wherein the learning unit adds a marker to the name of the character when the explanatory document and / or the appearance document contains the names of multiple people including the name of the character.

11. The learning device according to claim 1, wherein the learning unit, when the explanatory document and / or the appearance document contains words that identify a person, masks the words that identify the person with a certain probability.

12. The learning device according to claim 1, wherein the learning unit verifies the learned first numerical representation using a fourth network trained with a numerical representation different from the first numerical representation.

13. A generation device comprising: a first network that performs embedding based on an explanatory document for explaining a character; a second network that performs embedding based on a document in which the character appears in a story; and a generation unit that generates a document that is a generated sentence using a third network based on a first numerical representation generated using the first network and a target document to be edited, wherein the generation unit trains the third network using a loss function based on the generated sentence.

14. The generating device of claim 13, wherein the first numerical representation is obtained by training the first network and the second network so that the first numerical representation and a second numerical representation obtained by the second network become close to each other.

15. The generating device according to claim 13, wherein the generating unit divides a target document to be edited into a first document and a second document, and generates the generated sentence by inputting the first numerical representation of a person appearing in at least one of the first document and the second document and the first document into the third network.

16. The generation device according to claim 15, wherein the generation unit trains the third network using the loss function including a first loss function based on a negative log-likelihood of a string of the generated sentence.

17. The generation device described in claim 16, wherein the generation unit performs embedding on the generated sentence using the first network and the second network to obtain a third numerical representation, and further trains the third network using a second loss function calculated based on the first numerical representation and the third numerical representation.

18. The generation device described in claim 16, wherein the generation unit generates the generated sentence using the third network further based on a given keyword, and trains the third network further using a third loss function based on whether or not the keyword appears in the generated generated sentence.

19. The generating device described in claim 16, wherein the generation unit detects attributes related to the character from the second document, and further trains the third network using a fourth loss function based on whether the generated sentence contains an expression corresponding to the attributes.

20. A terminal device comprising: an input unit for accepting user operations; and a display control unit for controlling a display screen displayed on a display device, wherein the display control unit generates a display control signal for causing a display device to display a user interface screen including: a first network for performing embedding based on an explanatory document for explaining a character; a second network for performing embedding based on an appearance document in which the character appears in a story; a document area for displaying a generated sentence generated based on a numerical representation representing the character generated using the first network and a target document which is a document to be edited; and a character selection area for selecting an appearance character to be displayed in the generated sentence displayed in the document area, and places identification information associated with the numerical representation in the character selection area.

21. The terminal device according to claim 20, wherein the display control unit further includes an attribute designation area for designating an attribute to be assigned to the generated sentence on the user interface screen.

22. The terminal device according to claim 20, wherein the display control unit further includes a keyword designation area for designating a keyword to be included in the generated sentence on the user interface screen.

23. The terminal device according to claim 20, wherein the display control unit sets a position at which a generated sentence generated using the numerical expression is to be inserted into a document displayed in the document area in response to a user operation.

24. The terminal device according to claim 20, wherein the display control unit generates the display control signal for displaying on the user interface screen a tree diagram corresponding to deviations in the generated sentence when the generated sentence uses the numerical expression.

25. The terminal device according to claim 24, wherein the display control unit generates the generated sentence based on a path through the tree diagram in response to a user operation of tracing the nodes and edges of the tree diagram in order.

26. The terminal device according to claim 24, wherein the display control unit generates the display control signal for displaying, for each leaf of the tree diagram, a summary that summarizes documents generated on a path from the root of the tree diagram to the leaf.

27. The terminal device according to claim 24, wherein the display control unit generates the display control signal for displaying, for each leaf of the tree diagram, identification information that identifies a character appearing in a document generated on a path from the root of the tree diagram to the leaf.

28. The terminal device according to claim 20, wherein the display control unit generates the display control signal for displaying the target documents in chronological order based on chronological information related to the target documents.

29. The terminal device according to claim 28, wherein the display control unit generates the display control signal for enlarging and displaying a period specified in accordance with the user operation in the timeline, and determines a portion of the target document included in the period as a document to be used in generating the generated sentence.

30. The terminal device according to claim 20, wherein the display control unit generates the display control signal for dividing the target document into units that are meaningful in terms of the development of a story and displaying them, and for displaying summaries of sentences contained in each unit and identification information identifying characters in the unit in association with the unit.

Citation Information

Patent Citations

  • Role-oriented story outcome generation method

    CN113268983A

  • Method and program for creating and / or editing work

    JP2023023978A

  • Information processing device and information processing method

    WO2022201943A1