Script generation system, script generation method, and program

JPWO2025258027A5Pending Publication Date: 2026-05-22
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2026-01-09
Publication Date
2026-05-22
Patent Text Reader

Abstract

An input information acquisition unit (101) of a script generation system (1) acquires input information that is to be input into a trained language model which is capable of generating a product described in a natural language and that pertains to a video in which merchandise or a service is introduced. A script generation unit (102) inputs the input information into the language model, so as to generate a script pertaining to the video.
Need to check novelty before this filing date? Find Prior Art

Description

Script generation system, script generation method, and program

[0001] The present disclosure relates to a script generation system, a script generation method, and a program.

[0002] Conventionally, there is known a technique for generating a script for a video for introducing a product or service. For example, Patent Document 1 describes a script production device that receives an engine including production information for content with embedded advertising data, generates a script using the content production engine with embedded identification information that identifies the script writer who created the script, and transmits the script to a viewing device used by a viewer. Patent Document 1 also describes a template related to the script.

[0003] Japanese Patent Application Laid-Open No. 2008-022256

[0004] However, with the script production device of Patent Document 1, even if a script writer uses a template to create a script, the content of the template is fixed, so the script writer can only create a script within the scope of the template. For this reason, the technology of Patent Document 1 does not allow for increased flexibility in the final script. This point is not limited to script production such as that of Patent Document 1, but is also true of conventional technologies in general that generate scripts for videos to introduce products or services.

[0005] One of the goals of this disclosure is to generate more flexible scripts.

[0006] The script generation system of the present disclosure includes an input information acquisition unit that acquires input information related to a video in which a product or service is introduced, the input information being input into a trained language model capable of generating products written in natural language, and a script generation unit that generates a script related to the video by inputting the input information into the language model.

[0007] According to the present disclosure, more flexible scripts can be generated.

[0008] FIG. 1 is a diagram showing an example of the hardware configuration of a script generation system; FIG. 2 is a diagram showing an example of how live streaming is performed; FIG. 3 is a diagram showing an example of functions realized by the script generation system; FIG. 4 is a diagram showing an example of a script database; FIG. 5 is a diagram showing an example of input and output to a language model; FIG. 6 is a diagram showing an example of input information based on extracted information extracted from content; FIG. 7 is a diagram showing an example of processing executed by the script generation system; and FIG. 8 is a diagram showing an example of functions realized in a modified example.

[0009] [1. Hardware Configuration of Script Generation System] An example of an embodiment of a script generation system, a script generation method, and a program according to the present disclosure will be described. In this embodiment, a case where the script generation system, the script generation method, and the program are applied to a live streaming service will be exemplified. A live streaming service is a service that distributes videos in real time to an unspecified number of people. Performers in the live stream proceed with the live stream according to a script prepared in advance. Viewers watch the videos that are distributed in real time or videos that are saved as archives.

[0010] Fig. 1 is a diagram showing an example of the hardware configuration of a script generation system. For example, the script generation system 1 includes a script generation device 10, a server 20, a performer device 30, and a viewer device 40. Each of the script generation device 10, the server 20, the performer device 30, and the viewer device 40 is connected to a network N such as the Internet or a LAN. Although Fig. 1 shows one each of the script generation device 10, the server 20, the performer device 30, and the viewer device 40, there may be multiple devices of at least one of these.

[0011] The script generation device 10 is a device that generates a script. For example, the script generation device 10 is a personal computer, a server computer, a tablet, or a smartphone. For example, the script generation device 10 includes a control unit 11, a memory unit 12, a communication unit 13, an operation unit 14, and a display unit 15. The control unit 11 includes at least one processor. The memory unit 12 includes at least one of a volatile memory such as RAM and a non-volatile memory such as flash memory. The communication unit 13 includes at least one of a communication interface for wired communication and a communication interface for wireless communication. The operation unit 14 is an input device such as a touch panel. The display unit 15 is a liquid crystal or organic EL display.

[0012] The server 20 is a server computer for a live streaming service. For example, the server 20 includes a control unit 21, a storage unit 22, and a communication unit 23. The hardware configurations of the control unit 21, the storage unit 22, and the communication unit 23 may be similar to those of the control unit 11, the storage unit 12, and the communication unit 13, respectively.

[0013] The performer device 30 is a device of a performer. For example, the performer device 30 is a personal computer, a tablet, or a smartphone. For example, the performer device 30 includes a control unit 31, a memory unit 32, a communication unit 33, an operation unit 34, and a display unit 35. The hardware configurations of the control unit 31, the memory unit 32, the communication unit 33, the operation unit 34, and the display unit 35 may be similar to those of the control unit 11, the memory unit 12, the communication unit 13, the operation unit 14, and the display unit 15, respectively. A camera unit 36 ​​is connected to the performer device 30. The camera unit 36 ​​includes at least one camera. The camera unit 36 ​​may be included inside the performer device 30.

[0014] The viewer device 40 is a viewer's device. For example, the viewer device 40 is a personal computer, a tablet, or a smartphone. For example, the viewer device 40 includes a control unit 41, a memory unit 42, a communication unit 43, an operation unit 44, and a display unit 45. The hardware configurations of the control unit 41, the memory unit 42, the communication unit 43, the operation unit 44, and the display unit 45 may be similar to those of the control unit 11, the memory unit 12, the communication unit 13, the operation unit 14, and the display unit 15, respectively.

[0015] The programs stored in the storage units 12, 22, 32, 42 may be supplied to the script generation device 10, the server 20, the performer device 30, or the viewer device 40 via the network N. Also, at least one of a reading unit (e.g., a memory card slot) that reads a computer-readable information storage medium and an input / output unit (e.g., a USB port) for inputting and outputting data to and from an external device may be included in the script generation device 10, the server 20, the performer device 30, or the viewer device 40. For example, a program stored in an information storage medium may be supplied to the script generation device 10, the server 20, the performer device 30, or the viewer device 40 via at least one of the reading unit and the input / output unit.

[0016] Furthermore, the script generation system 1 only needs to include at least one computer. The computers included in the script generation system 1 are not limited to the example of FIG. 1. For example, the script generation system 1 may include only the script generation device 10 and the server 20. In this case, the performer device 30 and the viewer device 40 exist outside the script generation system 1. The script generation system 1 may also include only the script generation device 10. In this case, the server 20, the performer device 30, and the viewer device 40 exist outside the script generation system 1. For example, the script generation system 1 may include the script generation device 10 and another computer not shown in FIG. 1.

[0017] [2. Overview of the Present Embodiment] In the present embodiment, a case where a product or service sold through an e-commerce service is introduced through a live streaming service is taken as an example. The live streaming service may be one of the services provided by the operator of the e-commerce service, or may be a separate service separate from the e-commerce service. For example, a store affiliated with the e-commerce service may perform a live stream to introduce the product or service it sells. The live stream may be performed by any party. For example, the live stream may be performed by a party other than the store, such as a product manufacturer, a service provider, or an influencer.

[0018] 2 is a diagram showing an example of how live streaming is performed. For example, a performer introduces a product or service in front of the filming unit 36 ​​according to a prepared script. The performer may be a store associate or another person requested by the store to appear. The performer device 30 transmits data (video data) showing the results of filming by the filming unit 36 ​​to the server 20. The server 20 distributes video to the viewer device 40 in real time based on the data received from the performer device 30. The viewer device 40 displays an introduction screen SC on the display unit 45, showing a video of the performer introducing the product or service.

[0019] In this embodiment, the operator of a live streaming service provides a script generation service that generates a script to a store that has applied for live streaming. The store staff may prepare the script themselves, or may use the script generation service to prepare a script. For example, when a store staff member uses the script generation service, the store staff member fills out the necessary information on an orientation sheet (described below) and requests the operator to use the script generation service. Based on the request from the store staff member, the operator generates a script corresponding to the product or service to be introduced in the live streaming.

[0020] For example, it is conceivable that the operator would generate a script by writing it manually, but this would require a lot of work on the part of the operator. It is also conceivable that the operator would create a template for the script so that they would not have to write it from scratch, but this would result in a lack of flexibility, as the operator would only be able to create a script within the scope of the template. Therefore, the script generation system 1 of this embodiment uses a trained language model to generate a more flexible script. Details of the script generation system 1 will be described below.

[0021] 3 is a diagram showing an example of functions realized by the script generation system 1. FIG. 3 shows an example of functions realized by the script generation device 10. For example, the script generation device 10 includes a data storage unit 100, an input information acquisition unit 101, and a script generation unit 102. The data storage unit 100 is realized by the storage unit 12. The input information acquisition unit 101 and the script generation unit 102 are each realized by the control unit 11.

[0022] [3-1. Data Storage Unit] The data storage unit 100 stores data necessary for generating a script. For example, the data storage unit 100 stores a trained language model M capable of generating a product written in a natural language, and a script database DB in which a script generated by the language model M is stored. Note that the data stored in the data storage unit 100 is not limited to these examples. The data storage unit 100 can store any data.

[0023] A natural language is a language that can be understood by humans. For example, the natural language may be any language, such as Japanese, English, or Chinese. A product is data generated by a language model M. The language model M can generate any product. For example, the product may be text, a table, a diagram, an image, a video, or a combination thereof. Text includes letters, numbers, symbols, or a combination thereof. Text may be in any format. For example, text may be a sentence, a list, a list of words, program code, or code written in a markup language.

[0024] The language model M is a model used in the field of natural language processing. For example, the language model M is a model that uses a machine learning technique. The language model M is also called a large-scale language model or generative AI (Artificial Intelligence). The language model M may be a model called by other names. For example, the language model M includes a program that executes information processing to process input information and generate a product, and parameters referenced by the program. The parameters are adjusted through learning. The parameters of the language model M may be known parameters. For example, the parameters of the language model M may be weights or biases. In this embodiment, it is assumed that the language model M has been previously trained using a large dataset.

[0025] The language model M may be any of various types of known models. For example, the language model M may be a Generative Pre-trained Transformer (GPT), a Transformer-based model other than the GPT (e.g., Bidirectional Encoder Representations from Transformers (BERT) or Text-To-Text Transfer Transformer (T5)), a neural network capable of natural language processing, or other models (e.g., Pegasus, UniLM, or Electra). The programs and parameters included in the language model M may be similar to those of these known models. For example, the language model M may be exactly the same as a known model, or may be a model fine-tuned using training data specialized for script generation.

[0026] In this embodiment, an example is given in which the data storage unit 100 stores the language model M, but the language model M may be stored in a device other than the script generation device 10. For example, the other device may be a device managed by a company that provides the functions of the language model M online. In this case, the script generation device 10 transmits input information to be input into the language model M to the other device. The other device inputs the input information received from the script generation device 10 into the language model M stored therein. The other device transmits a product output by the language model M to the script generation device 10. The script generation device 10 receives a product from the other device.

[0027] 4 is a diagram showing an example of a script database DB. For example, the script database DB stores a live streaming ID, an orientation sheet, and a script. The script database DB may store any information related to the script. For example, the script database DB may store questions and answers described in a modified example below, a video of a performer introducing a product or service according to a script (e.g., a video for archive distribution), or input information used when generating the script.

[0028] A live streaming ID is an ID that can identify an individual live streaming. For example, when a store staff member applies for a live streaming service, a new live streaming ID is issued. An orientation sheet is data that indicates basic information about the video to be streamed via the live streaming service. For example, the orientation sheet includes information about the product or service to be introduced. The information about the product or service may be any information, such as the product name, service name, characteristics of the product or service itself, price, inventory, color variations, size variations, or other information. The orientation sheet may also include input characteristic information, input setting information, and input session information, which will be described later.

[0029] The orientation sheet may be in any data format. For example, the orientation sheet may be in a spreadsheet data format, a CSV format, a text format, a document format, or another format. The orientation sheet may also include information other than the product or service. For example, the orientation sheet may include information about the store that wishes to live stream (e.g., the store's ID or name), the date and time of the live stream, the length (length) of the live stream, information about the performers (e.g., the performers' names, stage names, or profiles), whether a script has been requested, the level of detail of the script (described below), prohibited words, or other information.

[0030] For example, a store staff member who wishes to live stream operates their own terminal to enter the necessary information into an orientation sheet. They may be required to enter all items on the orientation sheet, or only some of the items. The store staff member's terminal sends the orientation sheet to the server 20. The server 20 receives the orientation sheet from the store staff member's terminal. The server 20 issues a live streaming ID and records the live streaming ID and the orientation sheet in the memory unit 22. The orientation sheet may be generated by any person. For example, the operator of the live streaming service may generate the orientation sheet based on the requests of the store.

[0031] For example, the operator of the live streaming service may review whether to permit a store to perform live streaming based on the contents of the orientation sheet. If the store passes the review and the orientation sheet indicates that the store will use the script generation service, the script generation device 10 obtains the live streaming ID and orientation sheet from the server 20 and stores them in the script database DB.

[0032] In this embodiment, a script is stored in a script database DB. In this embodiment, the term "script" refers to data indicating a script. A script may be in any data format. For example, a script may be in text format, document format, spreadsheet data format, CSV format, or other format. The data format of a script may be specified by a default prompt, which will be described later. The script database DB may store input information used when generating a script. A script may be divided into separate data for each session, which will be described later.

[0033] [3-2. Input Information Acquisition Unit] The input information acquisition unit 101 acquires input information to be input to the language model M, which is input information related to a video introducing a product or service. The input information indicates text (e.g., a sentence) written in a natural language. The input information is also called a prompt. The input information may be information in a format that can be processed by the language model M, and includes at least one character. The input information may include information other than the text written in a natural language (e.g., an image or a video). Other information other than the input information (e.g., information that serves as a sample of a product) may be input to the language model M together with the input information.

[0034] 5 is a diagram illustrating an example of input and output to the language model M. In this embodiment, a script for a video introducing a product or service is generated as a product of the language model M, and therefore the input information includes at least one piece of information related to the video for which the script is to be generated. In the example of FIG. 5, the input information includes four pieces of information: a default prompt, input feature information, input setting information, and input session information. The input information may include any number of pieces of information. For example, the input information may include one, two, three, five or more pieces of information.

[0035] A default prompt is a prepared prompt. A default prompt is a type of input information. A default prompt can include any information. For example, a default prompt includes text written in a natural language. A default prompt may include sentences, program code, code written in a markup language such as JSON, or other text. A default prompt may include information other than text (e.g., an image or video). The default prompt is stored in the data storage unit 100. A person in charge of generating a script may edit the content of the default prompt.

[0036] In this embodiment, an example is taken in which the language model M is not a model specialized for a specific purpose but a general-purpose model capable of generating various products. Therefore, in order for the general-purpose language model M to recognize the processing content to be executed, the input information includes a default prompt indicating the processing content to be executed by the language model M. The processing content to be executed by the language model M can also be referred to as the role (task) of the language model M or the type of product to be generated by the language model M.

[0037] For example, a general-purpose language model M recognizes the processing content it should perform based on the default prompt included in the input information. That is, the general-purpose language model M recognizes its own role or the type of product it should generate based on the default prompt included in the input information. In the example of Figure 5, the default prompt indicates that the language model M should generate a script based on the input information, such as "You are the scriptwriter for a live broadcast in which a product or service will be introduced. Please generate a script for the live broadcast based on the input information entered by you."

[0038] Note that the default prompt is not limited to the example of FIG. 5 . For example, the default prompt may indicate that the language model M should generate a script using wording other than that shown in FIG. 5 . The default prompt may include information other than information indicating that the language model M should generate a script. For example, the default prompt may include the language of the script, the length of the script (e.g., the number of characters or the number of pages), the layout of the script (e.g., a format in which lines are written after the names of the actors), the file format of the script, or other information. The default prompt may include settings of the language model M itself.

[0039] Furthermore, the language model M may be a model specialized for generating a script. In this case, the language model M can generate a script even if a default prompt does not explicitly indicate that the language model M should generate a script, and therefore the input information does not need to include a default prompt. When the language model M is a model specialized for generating a script, training data specialized for generating a script is assumed to have been learned in the language model M. For example, training data including pairs of training input information and a correct script is learned in the language model M. A large amount of training data may be learned in the language model M. The language model M specialized for generating a script can generate a script based on input information input to itself, based on parameters adjusted in advance, even without a default prompt.

[0040] In the example of Figure 5, the input information includes input feature information, input setting information, and input session information in addition to the default prompt. The input information may include only some of the input feature information, input setting information, and input session information. For example, the input information may include only input feature information, only input setting information, only input session information, only input feature information and input setting information, only input feature information and input session information, or input setting information and input session information. The input information may also include other information not shown in Figure 5.

[0041] For example, the input information acquisition unit 101 acquires, from the script database DB, an orientation sheet associated with the live streaming ID of the live streaming for which a script is to be generated. The live streaming for which a script is to be generated is specified via the operation unit 14 of the script generation device 10, but may be identified by other methods. For example, the input information acquisition unit 101 may refer to the script database DB and identify a live streaming for which a script has not yet been generated as the live streaming for which a script is to be generated. The input information acquisition unit 101 may identify a live streaming for which the length of the period until the streaming date and time is less than a threshold as the live streaming for which a script is to be generated.

[0042] For example, the input information acquiring unit 101 acquires input feature information, which is input information regarding the features of a product or service. The input feature information is a type of input information. The features of a product or service can also be referred to as a description of the product or service. The input feature information is written in text in a natural language. For example, the input feature information indicates the features of the product or service using letters, numbers, symbols, or a combination thereof. For example, the input feature information may indicate the product or service's identification information (e.g., name, model number, manufacturer, or JAN code), classification (e.g., genre, category, attribute, or attribute value), appearance (e.g., design, size, color, or pattern), quality, function, material, price, discount rate, appealing points, store information, or other information. The input feature information may also be the title, description, search information, or other information of the product or service listed on the e-commerce service.

[0043] For example, the input information acquiring unit 101 acquires input feature information included in an orientation sheet. The input information acquiring unit 101 may generate input feature information based on information included in the orientation sheet, rather than acquiring the input feature information included in the orientation sheet. The input information acquiring unit 101 may generate input feature information based on information not included in the orientation sheet. The input information acquiring unit 101 may generate input feature information when the amount of input feature information included in the orientation sheet is insufficient or when the orientation sheet does not include input feature information.

[0044] For example, the input information acquiring unit 101 may acquire input feature information generated by inputting extracted information extracted from content related to the features of a product or service into the language model M or another language model. The content is electronic information. For example, the content may be a website, a web advertisement, a digital catalog, a digital pamphlet, a digital flyer, an image showing a paper captured by a scanner, or other information. In this embodiment, a website will be described as an example of content.

[0045] 6 is a diagram showing an example of input information based on extracted information extracted from content. For example, the input information acquisition unit 101 uses a known web search service to perform a search for a product or service to be introduced. The search query may be entered by the person operating the script generation device 10, or may be acquired based on input feature information included in an orientation sheet. The input information acquisition unit 101 acquires, as content, an image I showing the product or service from a website found in the search.

[0046] For example, the input information acquiring unit 101 performs optical character recognition on the content and extracts text from the content as extracted information. The extracted information may be information other than text (e.g., a table or a diagram). The input information acquiring unit 101 may acquire the extracted information extracted from the content as input feature information as it is. If the content includes text, the input information acquiring unit 101 may extract the extracted information by extracting the text from the content without particularly performing optical character recognition. Extraction of the extracted information may be performed by a function other than the input information acquiring unit 101.

[0047] For example, the input information acquisition unit 101 may acquire first input information, which is input information prepared in advance, and second input information, which is input information generated by inputting the first input information into language model M or another language model. The first input information may be information included in an orientation sheet. For example, input feature information included in an orientation sheet corresponds to the first input information. The other language model is a language model different from language model M that generates the script. The other language model, like language model M, may be any model. To describe the other language model, the term "language model M" in the description of language model M can be replaced with "other language model."

[0048] In this embodiment, since the extracted information extracted from the content may include information unrelated to the features of the product or service, the input information acquisition unit 101 inputs the extracted information to the language model M or another language model and acquires input feature information output from the language model M or another language model. For example, the input information acquisition unit 101 may input input feature information included in an orientation sheet together with the extracted information to the language model M or another language model. In this case, the input feature information included in the orientation sheet corresponds to the first input information.

[0049] For example, the input information acquisition unit 101 may input a default prompt to the language model M or another language model, indicating that input feature information should be generated. For example, the default prompt may indicate the processing content to be performed by the language model M or another language model (the role of the language model M or another language model, or the type of product to be generated by the language model M or another language model), such as "You are an AI that generates input feature information indicating the features of a product or service based on product or service information extracted from content." Note that if the language model M or another language model is a model specialized in generating input feature information from extracted information, such a default prompt may not be input.

[0050] For example, the language model M or another language model divides information such as extracted information input thereto into tokens and calculates embeddings for each token based on parameters adjusted through training. The token division method may be a known method. The embeddings indicate the features of the tokens. For example, the embeddings may be multidimensional vectors or may be in other formats. The language model M or another language model predicts the next text as needed based on the order of the embeddings of the tokens, and then generates the input feature information as a product.

[0051] For example, based on a default prompt, language model M or another language model extracts information from the extracted information that does not overlap with the input feature information included in the orientation sheet, and outputs the information as input feature information. The input information acquisition unit 101 acquires the input feature information output from language model M or another language model. In this case, the input feature information output from language model M or another language model corresponds to second input information. In the example of Figure 6, characters extracted by optical character recognition from image I of a website, which is an example of content, are formatted by language model M. Language model M outputs input feature information indicating features not included in the orientation sheet.

[0052] For example, the input information acquisition unit 101 acquires input setting information, which is input information related to the settings of a video. The input setting information is a type of input information. The video settings can be said to be characteristics of the video itself, rather than characteristics of the product or service itself. The video settings can also be said to be stored information associated with the video. The input setting information is written in text in a natural language. For example, the input setting information indicates the video settings using letters, numbers, symbols, or a combination thereof. For example, the input setting information is the distribution date and time of the video, its length (length), target demographic (e.g., age group or gender), performer information (e.g., name, profile, or role), prohibited expressions, the title of the video, a summary of the video, or other information.

[0053] For example, the input information acquisition unit 101 acquires input setting information included in an orientation sheet. The input information acquisition unit 101 may generate the input setting information based on the information included in the orientation sheet, rather than acquiring the input setting information included in the orientation sheet. The input information acquisition unit 101 may generate the input setting information based on information not included in the orientation sheet. The input information acquisition unit 101 may generate the input setting information when the amount of input setting information included in the orientation sheet is insufficient or when the orientation sheet does not include the input setting information.

[0054] For example, the input information acquisition unit 101 may input information included in an orientation sheet (e.g., input feature information, input setting information, or input session information included in the orientation sheet) to the language model M or another language model, and acquire input setting information output from the language model M or another language model. In this case, the information included in the orientation sheet corresponds to the first input information. When information not included in the orientation sheet is input to the language model M or another language model, the information not included in the orientation sheet corresponds to the first input information.

[0055] For example, the input information acquisition unit 101 may input a default prompt to the language model M or another language model, indicating that input setting information should be generated. For example, the default prompt may indicate the processing content to be executed by the language model M or another language model (the role of the language model M or another language model, or the type of product to be generated by the language model M or another language model), such as "You are an AI that generates input setting information indicating video settings based on information input to you."

[0056] For example, the language model M or another language model divides information input thereto (information used to generate the input setting information) into tokens and calculates embeddings for each token based on parameters adjusted through learning. The language model M or another language model predicts subsequent text as necessary based on the order of the embedded representations of the tokens, and then generates the input setting information as a product. For example, the language model M or another language model generates a video title and summary, etc., corresponding to the information input thereto based on a default prompt, and outputs the generated information as the input setting information. The input information acquisition unit 101 acquires the input setting information output from the language model M or another language model. In this case, the input setting information output from the language model M or another language model corresponds to second input information.

[0057] For example, the input information acquiring unit 101 acquires input session information, which is input information related to each of multiple sessions in a video. A session is an individual part that constitutes a video. A session may also be called by other names, such as a section. A video may be divided into multiple sessions from any perspective. For example, sessions may be divided by topic or by time. The input session information is a type of input information. The input session information is written in text in a natural language. For example, the input session information indicates a session using letters, numbers, symbols, or a combination thereof. The input session information may indicate an outline introduced in each session. For example, the session information may be a number indicating the order of the session, a heading, text indicating an outline, a time length, or other information.

[0058] For example, the input information acquiring unit 101 acquires input session information included in an orientation sheet. The input information acquiring unit 101 may generate input session information based on information included in the orientation sheet, rather than acquiring the input session information included in the orientation sheet. The input information acquiring unit 101 may generate input session information based on information not included in the orientation sheet. The input information acquiring unit 101 may generate input session information when the amount of input session information included in the orientation sheet is insufficient or when the orientation sheet does not include input session information.

[0059] For example, the input information acquisition unit 101 inputs information included in an orientation sheet (e.g., input feature information, input setting information, or input session information included in the orientation sheet) into the language model M or another language model, and acquires input session information output from the language model M or another language model. In this case, the information included in the orientation sheet corresponds to the first input information. When information not included in the orientation sheet is input to the language model M or another language model, the information not included in the orientation sheet corresponds to the first input information.

[0060] For example, the input information acquiring unit 101 may input a default prompt to the language model M or another language model, indicating that input session information should be generated. For example, the default prompt may indicate the processing content to be performed by the language model M or another language model (the role of the language model M or another language model, or the type of product to be generated by the language model M or another language model), such as "You are an AI that generates input session information indicating a video session based on information input to you." The default prompt may specify the number of sessions.

[0061] For example, the language model M or another language model divides information input thereto (information used to generate input session information) into tokens and calculates embeddings for each token based on parameters adjusted through learning. The language model M or another language model predicts subsequent text as necessary based on the order of the embedded representations of the tokens, and then generates the input session information as a product. For example, the language model M or another language model generates a session heading or the like corresponding to the information input thereto based on a default prompt, and outputs the generated heading as input session information. The input information acquisition unit 101 acquires the input session information output from the language model M or another language model. In this case, the input session information output from the language model M or another language model corresponds to second input information.

[0062] In this embodiment, the input information acquisition unit 101 acquires input session information according to the level of detail related to the script. The level of detail is the degree of detail of the script. The level of detail can also be referred to as the volume of the script. For example, there may be two levels of detail, such as a simple version and a detailed version, or there may be three or more levels of detail. The level of detail can be specified by any person. For example, the store staff may specify the level of detail. The input information acquisition unit 101 acquires input session information such that the higher the level of detail, the more finely divided the video is. The level of detail is indicated by letters, numbers, symbols, or a combination thereof.

[0063] For example, if the orientation sheet includes a level of detail, the input information acquiring unit 101 acquires the level of detail included in the orientation sheet. If the orientation sheet already includes input session information corresponding to the level of detail, the input information acquiring unit 101 acquires the input session information corresponding to the level of detail included in the orientation sheet. The input information acquiring unit 101 may acquire the input session information corresponding to the level of detail based on the language model M or another language model.

[0064] For example, the higher the level of detail, the greater the number of sessions. The higher the level of detail, the more sub-sessions are generated by dividing the session into smaller parts. A sub-session can also be considered a lower level of a session. A session may have three or more levels of hierarchy, not just two. The higher the level of detail, the greater the number of levels of a session. For example, when the level of detail indicates a simplified version, the input information acquisition unit 101 generates input session information using the method described above, and ends the process of acquiring input session information without dividing the session into smaller parts.

[0065] For example, when the level of detail indicates the detailed version, the input information acquiring unit 101 causes the language model M or another language model to generate sub-sessions by further dividing the session based on the input session information generated by the above-described method. For example, the input information acquiring unit 101 may input the input session information to the language model M or another language model, and acquire the input session information output from the language model M or another language model.

[0066] For example, the input information acquiring unit 101 may input a default prompt to the language model M or another language model, indicating that input session information according to the level of detail should be generated. For example, the default prompt may indicate the processing content to be executed by the language model M or another language model (the role of the language model M or another language model, or the type of product to be generated by the language model M or another language model), such as "You are an AI that generates input session information indicating sub-sessions into which a session is further divided, based on the input session information input to you."

[0067] For example, the language model M or another language model divides input session information input thereto into tokens and calculates embedding representations of each token based on parameters adjusted through learning. The language model M or another language model predicts subsequent text as necessary based on the order of the embedded representations of the tokens, and generates, as a product, input session information that is more detailed than the input session information input thereto. For example, the language model M or another language model generates, based on a default prompt, a sub-session heading or the like corresponding to the input session information input thereto, and outputs the sub-session heading as input session information. The input information acquisition unit 101 acquires input session information output from the language model M or another language model.

[0068] The input information acquiring unit 101 may input the level of detail to the language model M or another language model. In this case, the language model M or another language model may acquire input session information indicating the heading of a sub-session, etc., based on the level of detail input thereto.

[0069] Script Generation Unit The script generation unit 102 generates a script for a video by inputting input information into the language model M. The language model M divides the input information into tokens and calculates embedded representations for each token based on parameters adjusted through learning. The language model M predicts the next text as needed based on the order of the embedded representations of the tokens, and then generates a script as a product. For example, the language model M generates and outputs a script corresponding to the input information based on a default prompt. The input information acquisition unit 101 acquires the script output from the language model M.

[0070] For example, the script generation unit 102 generates a script by inputting input feature information into a language model M. The language model M divides the input feature information into tokens and calculates embedded representations for each token based on parameters adjusted through learning. The language model M predicts subsequent text as necessary based on the order of the embedded representations of the tokens, and then generates a script as a product. For example, the language model M generates and outputs a script corresponding to the input feature information based on a default prompt. The input feature information acquisition unit acquires the script output from the language model M.

[0071] For example, the script generation unit 102 generates a script by inputting input setting information into a language model M. The language model M divides the input setting information into tokens and calculates embedded representations for each token based on parameters adjusted by learning. The language model M predicts subsequent text as necessary based on the order of the embedded representations of the tokens, and then generates a script as a product. For example, the language model M generates and outputs a script corresponding to the input setting information based on a default prompt. The input setting information acquisition unit acquires the script output from the language model M.

[0072] For example, the script generation unit 102 generates a script by inputting input session information into a language model M. The language model M divides the input session information into tokens and calculates embedded representations for each token based on parameters adjusted through learning. The language model M predicts subsequent text as necessary based on the order of the embedded representations of the tokens, and then generates a script as a product. For example, the language model M generates and outputs a script corresponding to the input session information based on a default prompt. The input session information acquisition unit acquires the script output from the language model M.

[0073] For example, the script generation unit 102 generates a script according to the level of detail by inputting input session information according to the level of detail into the language model M. The language model M divides the input session information according to the level of detail into tokens and calculates embedded representations of each token based on parameters adjusted by learning. The language model M predicts the next text as necessary based on the order of the embedded representations of the tokens, and then generates a script as a product. For example, the language model M generates and outputs a script according to the input session information according to the level of detail based on a default prompt. The input session information acquisition unit acquires the script output from the language model M.

[0074] For example, for each session, the script generation unit 102 generates a script portion that is a part of the script, by inputting input session information for that session into the language model M, and for the second or subsequent session among the multiple sessions, the script generation unit 102 generates a script portion for the second or subsequent session by also inputting script portions of sessions before the second or subsequent session into the language model M, thereby generating a script based on the script portions of each of the multiple sessions. In this way, the default prompt may indicate that a script portion should be generated for each individual session.

[0075] For example, the script generation unit 102 generates a script by inputting first input information and second input information into a language model M. The language model M divides the first input information and second input information into tokens and calculates embedded representations of each token based on parameters adjusted through learning. The language model M predicts subsequent text as necessary based on the order of the embedded representations of the tokens, and then generates a script as a product. For example, the language model M generates and outputs a script corresponding to the first input information and the second input information based on a default prompt. The input session information acquisition unit acquires the script output from the language model M.

[0076] Note that the script generation unit 102 may generate a script including all sessions at once, rather than generating a script portion for each session. In this case, a default prompt may indicate that the script should be generated at once. Also, the script may not be particularly divided into multiple sessions. In this case, input session information is not acquired. The script generation unit 102 may generate a script based only on input feature information. The script generation unit 102 may generate a script based only on input setting information. The script generation unit 102 may generate a script based only on input session information. The script generation unit 102 may generate a script based on input information. The script generation unit 102 may generate a script without any particular regard to the level of detail.

[0077] 7 is a diagram showing an example of processing executed in the script generation system 1. FIG. 7 shows processing of the script generation device 10 among the processing executed in the script generation system 1. The control unit 11 executes a program stored in the storage unit 12 to execute the processing of FIG. 7. Each step of FIG. 7 is an example of a step included in the script generation method according to the present disclosure. For example, when a person in charge of generating a script in a script generation service operates the script generation device 10 to specify a live broadcast for which a script is to be generated, the processing of FIG. 7 is executed.

[0078] 7 , the script generation device 10 obtains, from the script database DB, an orientation sheet associated with the live streaming ID of the live streaming for which a script is to be generated (S1). The script generation device 10 obtains input feature information based on the orientation sheet (S2). In S2, the script generation device 10 obtains, as input feature information, information entered in a field on the orientation sheet that indicates the features of the product or service.

[0079] The script generation device 10 searches for content of products or services to be introduced in the live broadcast based on the input feature information acquired in S2 (S3). In S3, the script generation device 10 searches publicly known web search services using information such as product or service names indicated in the input feature information as a search query. The script generation device 10 acquires extracted information from the content searched in S3 (S4). The script generation device 10 acquires input feature information by inputting a default prompt indicating that input feature information will be generated and the extracted information extracted in S4 into a language model M (S5). In S5, the language model M divides these into tokens and outputs input feature information based on the sequence of embedded expressions of the tokens. The script generation device 10 acquires the input feature information output from the language model M.

[0080] Note that if a sufficient amount of input feature information is included in the orientation sheet, steps S3 to S5 do not need to be executed. For example, the person in charge of generating the script may specify whether or not to execute steps S3 to S5 by operating the operation unit 14. If information has been entered in all or a predetermined number of items in the orientation sheet that indicate the features of the product or service, the script generation device 10 does not need to execute steps S3 to S5.

[0081] The script generation device 10 acquires input setting information based on the orientation sheet (S6). In S6, if the orientation sheet contains input setting information, the script generation device 10 acquires the input setting information included in the orientation sheet. If the orientation sheet does not contain input setting information or if the orientation sheet does not contain a sufficient amount of input setting information, the script generation device 10 inputs a default prompt indicating that input setting information will be generated and the information included in the orientation sheet to the language model M. The language model M divides these into tokens and outputs input setting information based on the sequence of embedded expressions of the tokens. The script generation device 10 acquires the input setting information output from the language model M. In S6, the script generation device 10 may also acquire information such as a summary of a live broadcast as input setting information.

[0082] The script generation device 10 acquires input session information based on the orientation sheet (S7). In S7, if the orientation sheet contains input session information, the script generation device 10 acquires the input session information contained in the orientation sheet. If the orientation sheet does not contain input session information or if the orientation sheet does not contain a sufficient amount of input session information, the script generation device 10 inputs a default prompt indicating that input session information will be generated and the information contained in the orientation sheet into the language model M. The language model M divides these into tokens and outputs input session information based on the sequence of embedded expressions of the tokens. The script generation device 10 acquires the input session information output from the language model M.

[0083] The script generation device 10 determines whether to generate a simplified or detailed script based on the orientation sheet (S8). In S8, the script generation device 10 determines whether the level of detail of the script included in the orientation sheet indicates a simplified or detailed script. If the orientation sheet does not include the level of detail of the script, the person in charge of generating the script may specify the level of detail of the script by operating the operation unit 14. The script generation device 10 may make the determination in S8 based on the level of detail of the script specified by the person in charge.

[0084] If it is determined in S8 that a simplified script is to be generated (S8: simplified version), the script generation device 10 generates a script portion for the first session by inputting input information including input feature information, input setting information, and input session information into the language model M (S9). In S9, the script generation device 10 also inputs a default prompt indicating that a script portion corresponding to the session will be generated into the language model M. The language model M divides the input information, etc. into tokens and outputs the script portion for the first session based on the sequence of the embedded representations of the tokens. Since there are no other sessions before the first session, the language model M outputs the script portion for the first session without based on the script portions of other sessions. The script generation device 10 acquires the script portion for the first session output from the language model M.

[0085] The script generation device 10 generates a script portion for the next session by inputting input information including input feature information, input setting information, and input session information, and a script portion for a session that has already been generated, into the language model M (S10). In S10, the script generation device 10 also inputs a default prompt indicating that a script portion for the next session will be generated into the language model M. The language model M divides the input information and the script portion into tokens and outputs the script portion for the next session based on the sequence of the embedded representations of the tokens. For example, the script generation device 10 inputs the script portions of all sessions that have been generated up to that point into the language model M. The script generation device 10 obtains the script portion for the next session output from the language model M.

[0086] The script generation device 10 determines whether or not the script portion for the last session has been generated based on the input session information (S11). If it is determined in S11 that the script portion for the last session has not been generated (S11: N), the process returns to S10, and the script portion for the next session is generated. If it is determined that the script portion for the last session has been generated (S11: Y), the script generation device 10 generates a final script based on the script portions of each session (S12), and this process ends. In S12, the script generation device 10 generates the final script by connecting the script portions of each session. The script generation device 10 stores the final script in the script database DB.

[0087] If it is determined in S8 that a detailed version of the script is needed (S8: detailed version), the script generation device 10 acquires input session information indicating each sub-session of the multiple sessions by generating sub-sessions for each session indicated by the input session information (S13). In S13, the script generation device 10 inputs a default prompt indicating that a sub-session will be generated for each session and the input session information to the language model M. The language model M divides these into tokens and outputs input session information indicating the sub-sessions based on the sequence of the embedded expressions of the tokens. The script generation device 10 acquires the input session information output from the language model M.

[0088] The script generation device 10 generates a script portion of the first sub-session of the session to be processed by inputting input information, including input feature information, input setting information, and input session information, for the session to be processed into the language model M (S14). The session to be processed is the session that is the target of the loop from S14 to S17. The sessions to be processed are selected in order, starting with the first session. In S14, the script generation device 10 also inputs a default prompt to the language model M, indicating that a script portion will be generated according to the sub-session. The language model M divides the input information, etc. into tokens and outputs the script portion of the sub-session of the first session based on the sequence of the embedded representations of the tokens. Since there are no other sub-sessions before the sub-session of the first session, the language model M outputs the script portion of the first sub-session without using the script portions of the other sub-sessions. The script generation device 10 obtains the script portion of the first sub-session output from the language model M.

[0089] The script generation device 10 generates a script portion of the next sub-session of the session to be processed based on input information including input feature information, input setting information, and input session information for the session to be processed, and on the script portion of the sub-session that has already been created (S15). In S15, the script generation device 10 also inputs a default prompt indicating that a script portion corresponding to the next sub-session will be generated to the language model M. The language model M divides the input information and the script portion into tokens and outputs the script portion of the next sub-session based on the sequence of the embedded representations of the tokens. For example, the script generation device 10 inputs the script portions of all sub-sessions that have been generated up to that point into the language model M. The script generation device 10 obtains the script portion of the next sub-session output from the language model M.

[0090] The script generation device 10 determines, based on the input session information, whether or not the script portion of the last sub-session of the session being processed has been generated (S16). If it is determined that the script portion of the last sub-session of the session being processed has not been generated (N in S16), the process returns to S15, and the script portion of the next sub-session of the session being processed is generated. If it is determined that the script portion of the last sub-session of the session being processed has been generated (Y in S16), the script generation device 10 determines, based on the input session information, whether or not the script portion of the last session has been generated (S17).

[0091] If it is determined in S17 that the script portion for the last session has not yet been generated (S17: N), the process returns to S14, and the script portion for the first sub-session of the next session is generated. That is, the next session becomes the session to be processed. If it is determined that the script portion for the last session has been generated (S17: Y), the script generation device 10 generates a final script based on the script portions of each sub-session of each session (S18), and this process ends. In S18, the script generation device 10 generates the final script by connecting the script portions of each sub-session of each session. The script generation device 10 stores the final script in the script database DB.

[0092] 5. Summary of the Script Generation System The script generation system 1 of this embodiment acquires input information related to a video introducing a product or service. The script generation system 1 generates a script related to the video by inputting the input information into a language model M. The script generation system 1 can generate a flexible script according to the input information by using the language model M. For example, if a person providing a script generation service generates a script based on a template, the script cannot be generated flexibly because it can only be generated within the scope of the template. However, the language model M can generate a more flexible script by performing flexible natural language processing according to the input information. The script generation system 1 can also reduce the workload of the person providing the script, thereby reducing the costs associated with script generation. For example, if a person providing a script generation service outsources the generation of a script to an external company, costs will be incurred for the external company. However, the script generation system 1 can avoid such costs.

[0093] Furthermore, the script generation system 1 generates a script by inputting input feature information into the language model M. This enables the language model M to flexibly generate a script according to the features of a product or service, and the script generation system 1 can improve the accuracy of the script, which is the product of the language model M. For example, if the input feature information indicates an appearance feature, the script generation system 1 can generate a script according to the appearance feature. If the input feature information indicates a functional feature, the script generation system 1 can generate a script according to the functional feature.

[0094] Furthermore, the script generation system 1 acquires input feature information generated by inputting extracted information extracted from content related to the features of a product or service into the language model M or another language model. This allows the script generation system 1 to generate a script from more features, thereby improving the accuracy of the script that is the product of the language model M. For example, even if there is no input feature information in the orientation sheet or the orientation sheet does not contain a sufficient amount of input feature information, the script generation system 1 can acquire input feature information from the content.

[0095] Furthermore, the script generation system 1 generates a script by inputting input setting information into the language model M. Because the language model M can generate a flexible script according to the settings of the video, the script generation system 1 can improve the accuracy of the script that is the product of the language model M. For example, if the input setting information indicates the distribution date and time of the video, the script generation system 1 can generate a script of a length that corresponds to the distribution date and time of the video. If the input setting information indicates the length of the video, the script generation system 1 can generate a script of a length that corresponds to the length of the video.

[0096] Furthermore, the script generation system 1 generates a script by inputting input session information into the language model M. Because the language model M can flexibly generate a script according to the video session, the script generation system 1 can improve the accuracy of the script, which is the product of the language model M. For example, if the language model M attempts to generate a large number of products at once, the accuracy of the products may decrease. In this regard, if the language model M generates a script for each session, the amount of products generated by the language model M at one time can be reduced, and the script generation system 1 can improve the accuracy of the script.

[0097] Furthermore, the script generation system 1 generates a script according to the level of detail by inputting input session information according to the level of detail into the language model M. Because the language model M can flexibly generate a script according to the level of detail, the script generation system 1 can improve the accuracy of the script, which is the product of the language model M. For example, the script generation system 1 can generate a simplified script requested by a store for which a simplified script showing the overall progress of a live broadcast is sufficient. In this case, the amount of text generated by the language model M can be reduced, thereby reducing the processing load on the script generation device 10. If the language model M is stored in a computer other than the script generation device 10, the processing load on the other computer can be reduced. For example, the script generation system 1 can generate a detailed script requested by a store that requests a detailed script showing the progress of a live broadcast in detail.

[0098] Furthermore, for each session, the script generation system 1 generates a script portion for that session by inputting input session information for that session into the language model M. For a second or subsequent session among multiple sessions, the script generation system 1 generates the script portion for the second or subsequent session by also inputting script portions of sessions prior to the second or subsequent session into the language model M. The script generation system 1 generates a script based on each script portion of multiple sessions. This allows the script generation system 1 to generate a script that has connections between sessions. For example, if the language model M attempts to generate a large amount of text at once, the processing load on the script generation device 10 may increase. However, by having the language model M generate small script portions for each session, the processing load on the script generation device 10 can be reduced. Furthermore, if the language model M attempts to generate a large amount of text at once, the accuracy of the generated text may decrease. However, by having the language model M generate small script portions for each session, the script generation system 1 can improve the accuracy of the final script.

[0099] The script generation system 1 also acquires first input information prepared in advance and second input information generated by inputting the first input information into the language model M or another language model. The script generation system 1 generates a script by inputting the first input information and the second input information into the language model M. This allows the script generation system 1 to input a wide variety of input information into the language model M, thereby improving the accuracy of the script.

[0100] [6. Modifications] The present disclosure is not limited to the above-described embodiments. The present disclosure can be modified as appropriate without departing from the spirit of the present disclosure.

[0101] 8 is a diagram showing an example of functions realized in the modified example. For example, the script generation device 10 of the modified example is realized by a question and answer generation unit 103, a first sample acquisition unit 104, and a second sample acquisition unit 105. Each of the question and answer generation unit 103, the first sample acquisition unit 104, and the second sample acquisition unit 105 is realized by the control unit 11.

[0102] [6-1. Variation 1] For example, in the example of the introduction screen SC in FIG. 2, the live streaming service accepts comments from viewers. Viewers may enter questions about the products or services introduced by the performers as comments. The performers may answer questions from viewers during the live streaming. In Variation 1, questions that are expected from viewers and answers to those questions are generated from the contents of the script generated by the script generation unit 102 so that the performers can prepare for questions that are expected from viewers.

[0103] The script generation system 1 of Modification 1 includes a question and answer generation unit 103. The question and answer generation unit 103 generates questions and answers related to a video by inputting a script into the language model M or another language model. The question and answer generation unit 103 may input information other than the script into the language model M or another language model. For example, the question and answer generation unit 103 may input at least one of input feature information, input setting information, and input session information into the language model M or another language model.

[0104] The questions and answers are data including at least one question and at least one answer to the question. The questions and answers may be in any data format. For example, the questions and answers may be in a spreadsheet data format, a CSV format, a text format, a document format, or another format. The question and answer generation unit 103 can generate any number of questions and answers. The number of questions and answers to be generated by the question and answer generation unit 103 may be specified in advance, or may not be specified in particular.

[0105] For example, if language model M or another language model is not a model specialized in generating questions and answers, the question and answer generation unit 103 may input a default prompt to language model M or another language model indicating that it should generate questions and answers. The default prompt may indicate the processing content to be performed by language model M or another language model (the role of language model M or another language model, or the type of product to be generated by language model M or another language model), such as "You are an AI that generates questions and answers from a script entered for you."

[0106] For example, the language model M or another language model divides a script input thereto into tokens and calculates an embedding representation (feature vector) of each token based on parameters adjusted through learning. The language model M or another language model predicts subsequent text as necessary based on the order of the embedding representations of the tokens, and then generates questions and answers as products. For example, the language model M or another language model generates and outputs questions and answers according to the script input thereto based on a default prompt. The question and answer generation unit 103 acquires the questions and answers output from the language model M or another language model.

[0107] For example, the question and answer generation unit 103 stores the questions and answers in the script database DB. The server 20 transmits the questions and answers stored in the script database DB to the performer device 30. The question and answer generation unit 103 may store the questions and answers in a database other than the script database DB. The question and answer generation unit 103 may record the questions and answers in a computer other than the script generation device 10 or in an external information storage medium. The questions and answers generated by the question and answer generation unit 103 may be transmitted to a computer other than the performer device 30.

[0108] The script generation system 1 of the first modification generates questions and answers related to a video by inputting a script into the language model M or another language model. The script generation system 1 can support the progress of the performers by generating questions and answers according to the script.

[0109] [6-2. Modification 2] For example, in the embodiment, an example has been given of a case where input information, one example of which is input feature information, is input to the language model M. Any information that can be used to generate a script may be input to the language model M. Modification 2 takes as an example a case where a script sample is input to the language model M. The sample is script data used as a reference for the language model M. Like the script generated by the script generation unit 102, the sample may be in any data format. For example, the sample may be a script format or template. The sample may be generated manually or by the language model M. The sample may be a transcript of the audio of a video.

[0110] The script generation system 1 of variant example 2 includes a first sample acquisition unit 104. The first sample acquisition unit 104 acquires a sample of a video introducing a product different from the product introduced in the video for which a script is to be generated, or a service different from the service introduced in the video for which a script is to be generated, in which the features of the other product or service are omitted. The other video is a video other than the video for which a script is to be generated. The other video may be a video that has been distributed in the past, or a video for which a script has been generated but has not yet been distributed. The other video may be a video distributed by a service other than the live streaming service. The other product or service is a product or service introduced in the other video.

[0111] In the second modification, an example is taken of a case where the data storage unit 100 stores samples. The first sample acquisition unit 104 acquires samples from the data storage unit 100. The samples may be stored in a computer other than the script generation device 10 or in an external information storage medium. In this case, the first sample acquisition unit 104 may acquire the samples from the other computer or the external information storage medium.

[0112] For example, the first sample acquisition unit 104 acquires a sample from a script previously generated by the script generation unit 102, in which portions of words included in the input feature information used when generating the script are masked. Masking involves hiding specific portions (for example, filling the portions with spaces or specific symbols). The first sample acquisition unit 104 may acquire a sample by masking the script based on the input feature information. The first sample acquisition unit 104 may also acquire a sample in which masking has been completed by a function other than the first sample acquisition unit 104.

[0113] Note that the method of omitting features of other products or services is not limited to masking. For example, the first sample acquisition unit 104 may acquire a sample by deleting, from a script previously generated by the script generation unit 102, a portion of a word included in the input feature information used when generating the script. The first sample acquisition unit 104 may acquire a sample from which the portion has been deleted by a function other than the first sample acquisition unit 104.

[0114] The script generation unit 102 of Modification 2 generates a script by further inputting samples into the language model M. That is, the script generation unit 102 inputs the input information and samples described in the embodiment into the language model M. The language model M divides the input information and samples into tokens and calculates embedded representations of each token based on parameters adjusted by learning. The language model M predicts subsequent text as necessary based on the order of the embedded representations of the tokens, and then generates a script corresponding to the sample as a product. For example, the language model M may generate and output a script corresponding to the sample based on a default prompt. The default prompt may include text indicating that the language model M should refer to the sample, such as "Please refer to this sample to generate a script." The script generation unit 102 acquires the script output from the language model M.

[0115] The script generation system 1 of Modification 2 generates a script by further inputting samples from which the features of other products or services have been omitted into the language model M. This makes the language model M less susceptible to the effects of the features of other products or services, allowing the script generation system 1 to generate a script that reflects the features of the products or services introduced in the video for which the script is being generated. In other words, the script generation system 1 can improve the accuracy of the script.

[0116] [6-3. Variation 3] For example, a performer may have unique characteristics in the flow of speech or speaking style. If a script is generated that matches the performer's own characteristics, the performer is likely to be able to proceed with the live broadcast more easily. For this reason, scripts of other videos in which the same performer as the performer of the video for which a script is to be generated may be acquired as samples. In Variation 3, it is assumed that the scripts of other videos and the performers of those other videos are stored in the script database DB.

[0117] The script generation system 1 of variant example 3 includes a second sample acquisition unit 105. The second sample acquisition unit 105 acquires samples of other videos that feature the same performers as the performers in the video. The meaning of the sample is the same as in variant example 2. For example, the second sample acquisition unit 105 identifies the performers of the video for which a script is to be generated, based on an orientation sheet stored in the script database DB. The second sample acquisition unit 105 identifies other videos that feature the same performers as the identified performers, based on the orientation sheet stored in the script database DB. The second sample acquisition unit 105 acquires the scripts of the other videos that feature the identified performers as samples.

[0118] In addition, when there are multiple other videos featuring the same performer, the second sample acquisition unit 105 may acquire, as samples, the scripts of all of the other videos featuring the same performer, or may acquire, as samples, the scripts of some of the other videos featuring the same performer. Furthermore, when combining variants 2 and 3, the second sample acquisition unit 105 may acquire, as samples, scripts of other videos featuring the same performer as the video for which a script is to be generated, from which information about other products or services has been omitted. The second sample acquisition unit 105 may acquire, as samples, scripts of videos featuring the same performer that have been distributed via a service other than a live distribution service.

[0119] The script generation unit 102 of Modification 3 generates a script by further inputting samples into the language model M. That is, the script generation unit 102 inputs the input information and samples described in the embodiment into the language model M. The language model M divides the input information and samples into tokens and calculates embedded representations of each token based on parameters adjusted by learning. The language model M predicts the subsequent text as necessary based on the order of the embedded representations of the tokens, and then generates a script corresponding to the sample as a product. For example, the language model M may generate and output a script corresponding to the sample based on a default prompt similar to that of Modification 2. The script generation unit 102 acquires the script output from the language model M.

[0120] The script generation system 1 of the third modification generates a script by further inputting samples of other videos in which the same performer appears into the language model M. This allows the script generation system 1 to reflect the characteristics of the performer in the script. In other words, the script generation system 1 can improve the accuracy of the script.

[0121] [6-4. Other Modifications] For example, the above modifications may be combined.

[0122] For example, although the example has been given in which the script generation system 1 is applied to a script generation service provided within a live streaming service, the script generation system 1 can also be applied to other situations. The script generation system 1 may generate scripts for videos other than live streaming services. For example, the script generation system 1 may generate scripts for videos that are not distributed in real time, for videos that advertise products or services, for videos in which AI (artificial voice) speaks instead of a human, for videos that introduce products or services sold through services other than e-commerce services, or for other videos.

[0123] For example, the script generation system 1 may generate a script for a video introducing products or services of services other than e-commerce services. The script generation system 1 may generate a script for a video introducing products or services sold in physical stores, services provided at accommodation facilities such as hotels, communication services, payment services, financial services, e-book services, services provided at beauty salons or restaurants, or products or services introduced on social media. The products introduced in the video are not limited to tangible objects and may be content such as music or movies.

[0124] For example, the functions described as being implemented by the script generation device 10 may be implemented by the server 20, the performer device 30, or another computer. The processing described as being implemented by the script generation device 10 may be shared among multiple computers.

[0125] [7. Supplementary Note] For example, the script generation system according to the present disclosure can also be configured as follows.

[0126] (1) A script generation system including: an input information acquisition unit that acquires input information related to a video introducing a product or service, the input information being input to a trained language model capable of generating a product described in natural language; and a script generation unit that generates a script related to the video by inputting the input information to the language model. (2) The script generation system described in (1), wherein the input information acquisition unit acquires input feature information that is the input information related to features of the product or service, and the script generation unit generates the script by inputting the input feature information to the language model. (3) The script generation system described in (2), wherein the input information acquisition unit acquires the input feature information generated by inputting extracted information extracted from content related to features of the product or service into the language model or another language model. (4) The script generation system according to any of (1) to (3), wherein the input information acquisition unit acquires input setting information, which is the input information related to settings of the video, and the script generation unit generates the script by inputting the input setting information to the language model. (5) The script generation system according to any of (1) to (4), wherein the input information acquisition unit acquires input session information, which is the input information related to each of a plurality of sessions in the video, and the script generation unit generates the script by inputting the input session information to the language model. (6) The script generation system according to (5), wherein the input information acquisition unit acquires the input session information according to a level of detail related to the script, and the script generation unit generates the script according to the level of detail by inputting the input session information according to the level of detail to the language model.(7) The script generation system according to (5) or (6), wherein the script generation unit: for each session, generates the script portion of the session, which is a script portion that is part of the script, by inputting the input session information of the session into the language model, and for a second or subsequent session among the plurality of sessions, generates the script portion of the second or subsequent session by also inputting the script portion of the session before the second or subsequent session into the language model, and generates the script based on the script portion of each of the plurality of sessions. (8) The script generation system according to any of (1) to (7), wherein the input information acquisition unit acquires first input information, which is the input information prepared in advance, and second input information, which is the input information generated by inputting the first input information to the language model or another language model, and the script generation unit generates the script by inputting the first input information and the second input information to the language model. (9) The script generation system according to any of (1) to (8), further including a question and answer generation unit that generates questions and answers related to the video by inputting the script into the language model or another language model. (10) The script generation system according to any of (1) to (9), further including a first sample acquisition unit that acquires a sample related to another video introducing a product different from the product or a service different from the service, the sample omitting features of the other product or service, and the script generation unit generates the script by further inputting the sample into the language model. (11) The script generation system according to any of (1) to (10), further including a second sample acquisition unit that acquires a sample related to another video featuring the same performers as those in the video, and the script generation unit generates the script by further inputting the sample into the language model.

Claims

1. Input information to be input to a trained language model capable of generating products described in natural language, the input information acquisition unit acquires the input information relating to a video introducing a product or service, A script generation unit generates a script for the video by inputting the input information into the language model, A script generation system that includes [this].

2. The input information acquisition unit acquires input characteristic information, which is the input information relating to the characteristics of the product or the service. The script generation unit generates the script by inputting the input feature information into the language model. The script generation system according to claim 1.

3. The input information acquisition unit acquires the input feature information generated when extracted information extracted from content relating to the characteristics of the product or the service is input into the language model or another language model. The script generation system according to claim 2.

4. The input information acquisition unit acquires the input setting information, which is the input information relating to the settings of the video. The script generation unit generates the script by inputting the input setting information into the language model. A script generation system according to any one of claims 1 to 3.

5. The input information acquisition unit acquires input session information, which is the input information for each of the multiple sessions in the video. The script generation unit generates the script by inputting the input session information into the language model. A script generation system according to any one of claims 1 to 3.

6. The input information acquisition unit acquires the input session information corresponding to the level of detail of the script, which is the level of detail specified by the user from among multiple levels of detail. The script generation unit generates the script according to the level of detail by inputting the input session information according to the level of detail specified by the user into the language model. The script generation system according to claim 5.

7. The script generation unit, For each session, the input session information for that session is input to the language model to generate a script portion which is part of the script, and the script portion for that session is generated. For the second and subsequent sessions among the aforementioned multiple sessions, the script portion of the sessions preceding the second and subsequent sessions is also input into the language model to generate the script portion of the second and subsequent sessions. Based on the script portion of each of the aforementioned multiple sessions, the script is generated. The script generation system according to claim 5.

8. The input information acquisition unit acquires first input information, which is pre-prepared input information, and second input information, which is input information generated when the first input information is input to the language model or another language model. The script generation unit generates the script by inputting the first input information and the second input information into the language model. A script generation system according to any one of claims 1 to 3.

9. The script generation system further includes a question and answer generation unit that generates questions and answers relating to the video by inputting the script into the language model or another language model. A script generation system according to any one of claims 1 to 3.

10. The script generation system further includes a first sample acquisition unit that acquires a sample relating to another video in which a product different from the product or a service different from the service is introduced, wherein the characteristics of the other product or the other service are omitted. The script generation unit generates the script by further inputting the sample into the language model. A script generation system according to any one of claims 1 to 3.

11. The script generation system further includes a second sample acquisition unit that acquires samples related to other videos featuring the same performers as the performers in the aforementioned video. The script generation unit generates the script by further inputting the sample into the language model. A script generation system according to any one of claims 1 to 3.

12. Input information acquisition step, which involves acquiring input information relating to a video introducing a product or service, which is input to a trained language model capable of generating products described in natural language, A script generation step in which a script for the video is generated by inputting the input information into the language model, A script generation method that includes [this].

13. Input information to be input to a trained language model capable of generating products described in natural language, the input information acquisition unit acquires the input information relating to a video introducing a product or service. A script generation unit generates a script for the video by inputting the input information into the language model. A program that makes a computer function.