Script generation system, script generation method, and program
By using a language model generation system that has been trained, the limitations of scriptwriters in creating scripts within a fixed template range have been overcome, enabling more flexible script generation and improving the diversity and adaptability of scripts.
Patent Information
- Application Number
- CN202480044429.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-13
- Publication Date
- 2026-02-13
AI Technical Summary
In existing technologies, scriptwriters can only create scripts within a fixed template range, lacking flexibility and unable to generate more flexible scripts.
By employing a learned language model generation system, more flexible scripts can be generated by inputting information related to dynamic images introducing products or services.
It enables more flexible script generation, improving the diversity and adaptability of scripts.
Smart Images

Figure CN121532789A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to script generation systems, script generation methods, and procedures. Background Technology
[0002] Previously, techniques for generating scripts related to moving images used to introduce goods or services were known. For example, Patent Document 1 describes a script production apparatus that receives an engine-generating engine containing performance information including content embedded with advertising data, generates a script embedded with identification information recognizing the scriptwriter who created the script using the content creation engine, and sends it to a viewing device used by the audience. Patent Document 1 also describes templates related to the script.
[0003] Existing technical documents
[0004] Patent documents
[0005] Patent Document 1: Japanese Patent Application Publication No. 2008-022256 Summary of the Invention
[0006] The problem to be solved by the present invention
[0007] However, in the scriptwriting apparatus of Patent Document 1, even if the scriptwriter creates a script using a template, the content of the template is fixed, so the scriptwriter can only create the script within the scope of the template. Therefore, the technology of Patent Document 1 cannot improve the flexibility of the final script. This is not limited to the scriptwriting of Patent Document 1, but also applies to the prior art as a whole to generating scripts related to moving images used to introduce goods or services.
[0008] One purpose of this disclosure is to generate more flexible scripts.
[0009] Methods for solving problems
[0010] The script generation system according to this disclosure includes: an input information acquisition unit that acquires input information to be input into a learned language model, the learned language model being capable of generating products described in natural language, and the input information being related to a dynamic image introducing a product or service; and a script generation unit that generates a script related to the dynamic image by inputting the input information into the language model.
[0011] The effects of the invention
[0012] According to this disclosure, more flexible scripts can be generated. Attached Figure Description
[0013] Figure 1 This is a diagram illustrating an example of the hardware structure of a script generation system.
[0014] Figure 2 This is an example diagram showing a live broadcast in progress.
[0015] Figure 3 This is a diagram illustrating one example of the functionality implemented by the script generation system.
[0016] Figure 4 This is a diagram illustrating an example of a script database.
[0017] Figure 5 This is a diagram illustrating an example of the input and output of a language model.
[0018] Figure 6 This is a diagram illustrating an example of input information based on information extracted from content.
[0019] Figure 7 This is a diagram illustrating an example of the processing performed by the script generation system.
[0020] Figure 8 This is a diagram illustrating one example of the functionality implemented in the variant. Detailed Implementation
[0021] [1. Hardware Structure of the Script Generation System]
[0022] An example of an implementation of the script generation system, script generation method, and program according to this disclosure will be described. In this embodiment, as an example, the script generation system, script generation method, and program are applied to a live streaming service. The live streaming service is a service that publishes moving images in real time to an unspecified number of people. Actors perform live according to a pre-prepared script. Viewers watch the live-published moving images or the moving images saved as archives.
[0023] Figure 1 This diagram illustrates an example of the hardware structure of a script generation system. For example, script generation system 1 includes a script generation device 10, a server 20, an actor device 30, and a viewing device 40. Each of the script generation device 10, server 20, actor device 30, and viewing device 40 is connected to a network N such as the Internet or a LAN. Figure 1 In this document, although one of each of the script generation device 10, server 20, actor device 30 and viewing device 40 is shown, at least one of these devices may exist in multiples.
[0024] The script generation apparatus 10 is an apparatus for generating scripts. For example, the script generation apparatus 10 may be a personal computer, server computer, tablet computer, or smartphone. For example, the script generation apparatus 10 includes a control unit 11, a storage unit 12, a communication unit 13, an operation unit 14, and a display unit 15. The control unit 11 includes at least one processor. The storage unit 12 includes at least one of volatile memory such as RAM and non-volatile memory such as flash memory. The communication unit 13 includes at least one of a communication interface for wired communication and a communication interface for wireless communication. The operation unit 14 is an input device such as a touch panel. The display unit 15 is a liquid crystal display or an organic EL display.
[0025] Server 20 is a server computer for the live streaming service. For example, server 20 includes a control unit 21, a storage unit 22, and a communication unit 23. The hardware structure of the control unit 21, storage unit 22, and communication unit 23 can be the same as that of the control unit 11, storage unit 12, and communication unit 13, respectively.
[0026] The actor device 30 is an apparatus for the actor. For example, the actor device 30 may be a personal computer, tablet computer, or smartphone. For example, the actor device 30 includes a control unit 31, a storage unit 32, a communication unit 33, an operation unit 34, and a display unit 35. The hardware structure of the control unit 31, storage unit 32, communication unit 33, operation unit 34, and display unit 35 may be the same as that of the control unit 11, storage unit 12, communication unit 13, operation unit 14, and display unit 15, respectively. A shooting unit 36 is connected to the actor device 30. The shooting unit 36 includes at least one camera. The shooting unit 36 may be included inside the actor device 30.
[0027] Viewing device 40 is a device for the viewer. For example, viewing device 40 is a personal computer, tablet computer, or smartphone. For example, viewing device 40 includes a control unit 41, a storage unit 42, a communication unit 43, an operation unit 44, and a display unit 45. The hardware structure of the control unit 41, storage unit 42, communication unit 43, operation unit 44, and display unit 45 may be the same as that of the control unit 11, storage unit 12, communication unit 13, operation unit 14, and display unit 15, respectively.
[0028] Additionally, the program stored in storage units 12, 22, 32, and 42 can be provided to the script generation apparatus 10, server 20, actor device 30, or viewing device 40 via network N. Furthermore, at least one of a reading unit (e.g., a memory card slot) for reading computer-readable information storage media and an input / output unit (e.g., a USB port) for inputting and outputting data to external devices can be included in the script generation apparatus 10, server 20, actor device 30, or viewing device 40. For example, the program stored in the information storage medium can be provided to the script generation apparatus 10, server 20, actor device 30, or viewing device 40 via at least one of the reading unit and the input / output unit.
[0029] Furthermore, the script generation system 1 may include at least one computer. The computer included in the script generation system 1 is not limited to… Figure 1 Examples. For instance, script generation system 1 may consist only of script generation device 10 and server 20. In this case, actor device 30 and viewing device 40 exist outside of script generation system 1. Script generation system 1 may consist only of script generation device 10. In this case, server 20, actor device 30, and viewing device 40 exist outside of script generation system 1. For example, script generation system 1 may include script generation device 10 and... Figure 1 Another computer not shown.
[0030] [2. Overview of this implementation method]
[0031] In this embodiment, the example will be a scenario where a product or service sold within an e-commerce service is presented through a live-streaming service. The live-streaming service can be one of the services provided by the operator of the e-commerce service, or it can be a separate service independent of the e-commerce service. For example, a store participating in the e-commerce service might conduct a live stream to introduce the goods or services it sells. The live stream can be performed by anyone. For instance, the live stream could not be conducted by a store, but by a manufacturer of goods, a provider of services, or an influencer, among other people.
[0032] Figure 2 This diagram illustrates an example of a live broadcast. For instance, an actor introduces a product or service in front of a shooting unit 36 according to a pre-prepared script. The actor can be a staff member of the store or another person commissioned by the store to perform. The actor device 30 sends data (moving image data) representing the shooting results of the shooting unit 36 to the server 20. The server 20 publishes the moving image to the viewing device 40 in real time based on the data received from the actor device 30. The viewing device 40 causes the display unit 45 to display an introduction screen SC, which shows the moving image of the actor introducing the product or service.
[0033] In this embodiment, the operator of the live streaming service provides a script generation service to stores that have applied for live streaming. The store manager can prepare the script themselves or use the script generation service. For example, if the store manager uses the script generation service, they may note necessary information on a presentation sheet (described later) and request the operator to use the script generation service. Based on the request from the store manager, the operator generates a script corresponding to the products or services to be introduced in the live stream.
[0034] For example, it's also possible for operators to generate scripts manually, but this is time-consuming. Alternatively, operators could use script templates, eliminating the need to write scripts from scratch, but this limits their script creation to the templates, lacking flexibility. Therefore, the script generation system 1 according to this embodiment uses a learned language model to generate more flexible scripts. The script generation system 1 will be described in detail below.
[0035] [3. Functions implemented by the script generation system]
[0036] Figure 3 This is a diagram illustrating an example of the functionality implemented by the script generation system 1. Figure 3 An example of the functions implemented by the script generation apparatus 10 is shown. For example, the script generation apparatus 10 includes a data storage unit 100, an input information acquisition unit 101, and a script generation unit 102. The data storage unit 100 is implemented by a storage unit 12. Each of the input information acquisition unit 101 and the script generation unit 102 is implemented by a control unit 11.
[0037] [3-1. Data Storage Department]
[0038] The data storage unit 100 stores the data required to generate the script. For example, the data storage unit 100 stores a learned language model M capable of generating products described in natural language, and a script database DB storing scripts generated from the language model M. However, the data stored in the data storage unit 100 is not limited to these examples. The data storage unit 100 can store any type of data.
[0039] Natural language is a language that humans can recognize. For example, natural language can be any language such as Japanese, English, or Chinese. The output is the data generated by the language model M. The language model M can generate any output. For example, the output can be text, tables, graphs, images, animated images, or combinations thereof. Text includes characters, numbers, symbols, or combinations thereof. Text can be in any form. For example, text can be an article, a record, a word sequence, program code, or code described in a markup language.
[0040] A language model M is a model used in the field of natural language processing. For example, a language model M is a model that utilizes machine learning methods. Language model M is sometimes also referred to as a large-scale language model or generative AI (Artificial Intelligence). Language model M can also be called a model by other names. For example, language model M includes a program that performs information processing to generate outputs from input information and parameters referenced by that program. The parameters are adjusted through learning. The parameters of language model M can be well-known parameters. For example, the parameters of language model M can be weights or biases. In this embodiment, language model M is a model that has been pre-learned based on a large-scale dataset.
[0041] Furthermore, the language model M can be any known type of model. For example, the language model M can be a GPT (Generative Pre-trained Transformer), a variable model other than GPT (e.g., BERT: Bidirectional Encoder Representations from Transformers, or T5: Text-To-Text Transfer Transformer), a neural network or other model capable of natural language processing (e.g., Pegasus, UniLM, or Electra). The procedures and parameters included in the language model M can be the same as these known models. For example, the language model M can be exactly the same as a known model, or it can be a model fine-tuned using training data specifically designed for script generation.
[0042] In this embodiment, the data storage unit 100 stores the language model M, but the language model M can be stored in a device other than the script generation device 10. For example, the other device could be a device managed by a company that provides online functionality for the language model M. In this case, the script generation device 10 sends the input information to the language model M to the other device. The other device inputs the input information received from the script generation device 10 into its own stored language model M. The other device sends the output of the language model M to the script generation device 10. The script generation device 10 receives the output from the other device.
[0043] Figure 4This diagram illustrates an example of a script database (DB). For example, the script database DB stores the live stream ID, the target audience, and the script. Any information about the script can be stored in the script database DB. For example, the script database DB can store questions and answers as described in the variations described later, moving images of actors introducing products or services according to the script (e.g., moving images for document release), or input information used when generating the script.
[0044] A live stream ID is an identifier that recognizes each live stream. For example, when a store manager applies for live streaming services, a new live stream ID is issued. A profile picture is data representing basic information about the dynamic images displayed in the live streaming service. For example, a profile picture contains information related to the product or service being featured. This information can be arbitrary, such as the product name, service name, characteristics of the product or service itself, price, inventory, color variations, size variations, or other information. Profile pictures may include input feature information, input setting information, and input session information, as described later.
[0045] Furthermore, the orientation video can be in any data format. For example, it can be in the form of spreadsheet data, CSV, text, document, or other formats. Additionally, the orientation video can contain information beyond the product or service. For example, it may include information about the store to be livestreamed (e.g., store ID or name), the date and time of the livestream, the length of the livestream (during runtime), actor information (e.g., actor's name, stage name, or bio), whether a scriptwriting commission exists, the level of detail in the script (described later), outtakes, or other information.
[0046] For example, the person in charge of a store wishing to conduct a live stream uses their own terminal to input the necessary information into the orientation video. They may be required to input all items in the orientation video, or only a portion of them. The store manager's terminal sends the orientation video to server 20. Server 20 receives the orientation video from the store manager's terminal. Server 20 issues a live stream ID and records the live stream ID and the orientation video in storage unit 22. The orientation video can be generated by anyone. For example, the operator of the live stream service can also generate the orientation video based on the wishes of the store.
[0047] For example, the operator of the live streaming service reviews whether to approve a store's live stream based on the content of the orientation video. When the store passes the review and the orientation video shows that the store is using the script generation service, the script generation device 10 obtains the live stream ID and orientation video from the server 20 and stores them in the script database DB.
[0048] In this embodiment, the script is stored in a script database DB. In this embodiment, the portion recorded as a script refers to the data representing the script. The script can be in any data format. For example, a script can be in text format, document format, spreadsheet software data format, CSV format, or other formats. The data format of the script can also be specified in the default prompts described later. The input information used when generating the script can be stored in the script database DB. The script can be divided into separate data for each scene, as described later.
[0049] [3-2. Input Information Acquisition Unit]
[0050] The input information acquisition unit 101 acquires the input information input to the language model M, namely, input information related to dynamic images introducing goods or services. The input information represents text described in natural language (e.g., an article). The input information may also be referred to as a prompt. The input information can be any form of information that the language model M can process, and must contain at least one character. The input information may also contain other information besides text described in natural language (e.g., images or dynamic images). In the language model M, other information besides the input information (e.g., information about a sample of the product) may also be input along with the input information.
[0051] Figure 5 This diagram illustrates an example of the input and output of a language model M. In this embodiment, since a script introducing a dynamic image of a product or service is generated as the product of the language model M, the input information includes at least one piece of information related to the dynamic image that becomes the object of the script generation. Figure 5 In this example, the input information includes four pieces of information: default prompt, input feature information, input setting information, and input session information. The input information can include any number of pieces of information. For example, the input information can include one, two, three, or more than five pieces of information.
[0052] A default prompt is a pre-prepared prompt. A default prompt is a type of input information. A default prompt can include any information. For example, a default prompt may include text described in natural language. A default prompt may include articles, program code, code described in markup languages such as JSON, or other text. A default prompt may include other information besides text (e.g., images or animated GIFs). Default prompts are stored in data storage unit 100. The person responsible for generating the script can edit the content of the default prompts.
[0053] In this embodiment, an example is given where the language model M is not a model specific to a particular purpose, but rather a general model capable of generating various products. Therefore, in order for the general language model M to recognize the processing it should perform, the input information includes a default prompt indicating the processing that the language model M should perform. The processing that the language model M should perform can also be the function (task) of the language model M, or the type of product that the language model M should generate.
[0054] For example, a general language model M identifies the processing it should perform based on default prompts included in the input information. In other words, a general language model M can identify its role or the type of output it should produce based on default prompts included in the input information. Figure 5 In the example, the default prompt could be, "You are a playwright introducing a product or service in a live stream. Please generate a live stream script based on your input information." This indicates that the language model M should generate a script based on the input information.
[0055] In addition, the default prompt is not limited to Figure 5 For example, the default suggestion can be used in addition to... Figure 5 Other words besides those in the text indicate that language model M should generate a script. Default prompts can include information other than what indicates that language model M should generate a script. For example, default prompts can include the language of the script, the number of scripts (e.g., number of characters or pages), the layout of the script (e.g., the format of lines following actors' names), the file format of the script, or other information. Default prompts can also include settings of the language model M itself.
[0056] Furthermore, the language model M can be a model specifically designed for script generation. In this case, even if the default prompt does not explicitly state that the language model M should generate a script, the input information may not include the default prompt because the language model M is capable of generating a script. When the language model M is a model specifically designed for script generation, it is assumed that training data specifically for script generation is learned into the language model M. For example, training data including the input information used for training and pairs of scripts that become the correct answers is learned into the language model M. Multiple training data sets can also be learned in the language model M. Even without a default prompt, the language model M specifically designed for script generation can generate a script based on pre-tuned parameters and the input information given to itself.
[0057] exist Figure 5In the example, the input information includes not only the default prompt but also input feature information, input setting information, and input session information. The input information may include only a portion of the input feature information, input setting information, and input session information. For example, the input information may include only input feature information, only input setting information, only input session information, only input feature information and input setting information, only input feature information and input session information, or only input setting information and input session information. The input information may include... Figure 5 Other information not shown in the document.
[0058] For example, the input information acquisition unit 101 retrieves the orientation image associated with the live stream ID of the live stream that has become a target for script generation from the script database DB. The live streams that have become targets for script generation are specified by the operation unit 14 of the script generation apparatus 10, but can also be determined by other methods. For example, the input information acquisition unit 101 can refer to the script database DB and determine live streams that have not yet been scripted as targets for script generation. The input information acquisition unit 101 can also determine live streams whose duration up to the publication date is less than a threshold as targets for script generation.
[0059] For example, the input information acquisition unit 101 acquires input information related to the characteristics of a product or service, namely, input feature information. Input feature information is a type of input information. The characteristics of a product or service can also be referred to as the description of the product or service. Input feature information is described in natural language text. For example, input feature information can represent the characteristics of a product or service through characters, numbers, symbols, or combinations thereof. For example, input feature information may show the product or service's identification information (e.g., name, model, manufacturer, or JAN code), classification (e.g., category, class, attribute, or attribute value), appearance (e.g., design, size, color, or pattern), quality, function, material, price, discount rate, selling points, store information, or other information. Input feature information may be the title, description, search information, or other information of a product or service published in an e-commerce service.
[0060] For example, the input information acquisition unit 101 acquires the input feature information contained in the orientation sheet. The input information acquisition unit 101 may also generate input feature information based on information contained in the orientation sheet, without acquiring the input feature information contained in the orientation sheet. The input information acquisition unit 101 may also generate input feature information based on information not contained in the orientation sheet. The input information acquisition unit 101 may also generate input feature information when the amount of information in the input feature information contained in the orientation sheet is insufficient, or when the orientation sheet does not contain input feature information.
[0061] For example, the input information acquisition unit 101 can acquire input feature information generated by inputting extracted information from content related to the characteristics of goods or services into a language model M or other language model. The content is electronic information. For example, the content may be a website, online advertisement, digital catalog, digital brochure, digital flyer, an image representing a sheet of paper taken by a scanner, or other information. In this embodiment, a website will be described as an example of content.
[0062] Figure 6 This diagram illustrates an example of input information based on information extracted from content. For example, the input information acquisition unit 101 uses a known online search service to perform a search for goods or services that are the subject of the search. The search query during the search can be entered by a person using the script generation device 10, or it can be obtained based on input feature information included in the search results. The input information acquisition unit 101 obtains an image I representing the goods or services as content from the website that was found through the search.
[0063] For example, the input information acquisition unit 101 performs optical character recognition on the content and extracts text from the content as extracted information. The extracted information can also be other information besides text (e.g., a table or graph). The input information acquisition unit 101 can directly acquire the extracted information from the content as input feature information. If the content contains text, the input information acquisition unit 101 may also extract the extracted information by extracting text from the content without performing optical character recognition. The extraction of extracted information can also be performed through functions other than the input information acquisition unit 101.
[0064] For example, the input information acquisition unit 101 can also acquire pre-prepared input information, i.e., first input information, and input information generated by inputting the first input information into language model M or another language model, i.e., second input information. The first input information can also be information contained in the orientation slice. For example, the input feature information contained in the orientation slice is equivalent to the first input information. Other language models are language models different from the language model M that generates the script. Like language model M, other language models can be any model. The description of other language models simply requires replacing the description of "language model M" in the description of language model M with "other language model".
[0065] In this embodiment, since the extracted information from the content may include information unrelated to the characteristics of the goods or services, the input information acquisition unit 101 inputs the extracted information into the language model M or other language model, and acquires the input feature information output from the language model M or other language model. For example, the input information acquisition unit 101 may also input the input feature information contained in the orientation patch along with the extracted information for the language model M or other language model. In this case, the input feature information contained in the orientation patch is equivalent to the first input information.
[0066] For example, the input information acquisition unit 101 can input a default prompt to the language model M or other language model indicating that input feature information should be generated. For example, the default prompt could be, "You are an AI that generates input feature information representing the characteristics of goods or services based on information about goods or services extracted from content." This indicates the processing content that the language model M or other language model should perform (the function of the language model M or other language model, or the type of product that the language model M or other language model should generate). Furthermore, if the language model M or other language model is specifically designed to generate input feature information based on extracted information, such a default prompt may not be required.
[0067] For example, a language model M or other language model segments the extracted information from its input into tokens and computes the embedding representation of each token based on parameters adjusted through learning. The token segmentation method can be a well-known one. The embedding representation represents the features of the token. For example, the embedding representation can be a multi-dimensional vector or other form. Based on the order of the token embedding representations, the language model M or other language model generates the input feature information as output, predicting subsequent text as needed.
[0068] For example, language model M or another language model can extract information that does not overlap with the input feature information included in the orientation slice based on default prompts, and output it as input feature information. The input information acquisition unit 101 acquires the input feature information output from language model M or another language model. In this case, the input feature information output from language model M or another language model is equivalent to the second input information. Figure 6 In the example, characters extracted from image I of a website (which serves as content) using optical character recognition are styled using language model M. Language model M outputs input feature information representing features not included in the style sheet.
[0069] For example, the input information acquisition unit 101 acquires input information related to the settings of a moving image, namely, input setting information. Input setting information is a type of input information. The settings of a moving image may not be a feature of the product or service itself, but rather a feature of the moving image itself. The settings of a moving image may also be referred to as information stored in association with the moving image. Input setting information is described in natural language text. For example, input setting information represents the settings of a moving image through characters, numbers, symbols, or combinations thereof. For example, input setting information may be the release date and time of the moving image, its length (during runtime), target audience (e.g., age group or gender), actor information (e.g., name, introduction, or role), NG expressions, the title of the moving image, the summary of the moving image, or other information.
[0070] For example, the input information acquisition unit 101 acquires the input setting information contained in the orientation sheet. The input information acquisition unit 101 may also generate input setting information based on information contained in the orientation sheet without acquiring the input setting information contained in the orientation sheet. The input information acquisition unit 101 may also generate input setting information based on information not contained in the orientation sheet. The input information acquisition unit 101 may also generate input setting information when the amount of input setting information contained in the orientation sheet is insufficient, or when the orientation sheet does not contain input setting information.
[0071] For example, the input information acquisition unit 101 may also input information contained in the orientation patch (e.g., input feature information, input setting information, or input scene information contained in the orientation patch) into the language model M or other language model to acquire input setting information output from the language model M or other language model. In this case, the information contained in the orientation patch is equivalent to the first input information. If information not contained in the orientation patch is input into the language model M or other language model, the information not contained in the orientation patch is equivalent to the first input information.
[0072] For example, the input information acquisition unit 101 may also input a default prompt to the language model M or other language model indicating that input setting information should be generated. For example, the default prompt may be "You are an AI that generates input setting information representing the settings of a dynamic image based on the information input to you." In this way, it indicates the processing content that the language model M or other language model should perform (the function of the language model M or other language model, or the type of product that the language model M or other language model should generate).
[0073] For example, language model M or another language model segments the information input to itself (the information used to generate input setting information) into tokens and calculates the embedding representation of each token based on parameters adjusted through learning. Based on the order of the token embedding representations, language model M or another language model generates input setting information as output, predicting subsequent text as needed. For example, language model M or another language model generates a title and summary of a dynamic image corresponding to the information input to itself based on default prompts, and outputs it as input setting information. Input information acquisition unit 101 acquires the input setting information output from language model M or another language model. In this case, the input setting information output from language model M or another language model is equivalent to the second input information.
[0074] For example, the input information acquisition unit 101 acquires input scene information, which is input information related to each of the multiple scenes in a moving image. A scene is a component of a moving image. A scene can also be called a section or other names. A moving image can be divided into multiple scenes from any viewpoint. For example, scenes can be divided according to each topic or according to time. Input scene information is a type of input information. Input scene information is written in natural language text. For example, input scene information can represent scenes using characters, numbers, symbols, or combinations thereof. Input scene information can represent an overview presented in each scene. For example, scene information can be a number indicating the sequence of scenes, a title, text indicating an overview, time length, or other information.
[0075] For example, the input information acquisition unit 101 acquires input field information contained in the orientation sheet. The input information acquisition unit 101 may also generate input field information based on information contained in the orientation sheet without acquiring the input field information contained in the orientation sheet. The input information acquisition unit 101 may also generate input field information based on information not contained in the orientation sheet. The input information acquisition unit 101 may also generate input field information when the amount of input field information contained in the orientation sheet is insufficient, or when the orientation sheet does not contain input field information.
[0076] For example, the input information acquisition unit 101 inputs information contained in the orientation patch (e.g., input feature information, input setting information, or input scene information contained in the orientation patch) to the language model M or other language model, and acquires the input scene information output from the language model M or other language model. In this case, the information contained in the orientation patch is equivalent to the first input information. If information not contained in the orientation patch is input to the language model M or other language model, the information not contained in the orientation patch is equivalent to the first input information.
[0077] For example, the input information acquisition unit 101 may also provide a default prompt to the language model M or other language model indicating that input scene information should be generated. For example, the default prompt could be something like, "You are an AI that generates input scene information representing scenes of a moving image based on the information you input." This indicates the processing content that the language model M or other language model should perform (the function of the language model M or other language model, or the type of output that the language model M or other language model should generate). The number of scenes can be specified in the default prompt.
[0078] For example, language model M or another language model segments the information input to itself (the information used to generate input scene information) into tokens and calculates the embedding representation of each token based on parameters adjusted through learning. Based on the order of the token embedding representations, language model M or another language model generates input scene information as output, predicting subsequent text as needed. For example, language model M or another language model generates a title for the scene corresponding to the information input to itself based on default prompts and outputs it as input scene information. Input information acquisition unit 101 acquires the input scene information output from language model M or another language model. In this case, the input scene information output from language model M or another language model is equivalent to the second input information.
[0079] In this embodiment, the input information acquisition unit 101 acquires input scene information corresponding to the level of detail in the script. Detail level refers to the degree of detail in the script. Detail level can also be referred to as the script's volume. For example, it can be a two-stage level of detail, such as a simplified version or a detailed version, or a level of detail with three or more stages. The level of detail can be specified by any person. For example, the person in charge of the shop can specify the level of detail. The input information acquisition unit 101 acquires the input scene information such that the higher the level of detail, the more finely the moving images are segmented. The level of detail is represented by characters, numbers, symbols, or a combination thereof.
[0080] For example, if the orientation slice includes detail, the input information acquisition unit 101 acquires the detail included in the orientation slice. If the input scene information corresponding to the detail is already included in the orientation slice, the input information acquisition unit 101 acquires the input scene information corresponding to the detail included in the orientation slice. The input information acquisition unit 101 may acquire the input scene information corresponding to the detail based on language model M or other language models.
[0081] For example, higher detail levels result in more sessions. Higher detail levels also lead to more sub-sessions that are further subdivided into smaller sessions. These sub-sessions can also be referred to as lower levels of a session. A session may not have two stages, but rather three or more levels. Higher detail levels also allow for more levels of sessions. For example, if the detail level represents a simplified version, the input information acquisition unit 101 generates input session information using the method described above and ends the process of acquiring the input session information without further subdividing the sessions.
[0082] For example, in the case of a detailed representation, the input information acquisition unit 101, based on the input scene information generated by the above method, causes the language model M or other language model to generate sub-scenes obtained by further refining the segmentation of the scenes. For example, the input information acquisition unit 101 can input the input scene information into the language model M or other language model to obtain the input scene information output from the language model M or other language model.
[0083] For example, the input information acquisition unit 101 can input a default prompt to the language model M or other language model, which indicates that input scene information corresponding to the level of detail should be generated. For example, the default prompt could be, "You are generating AI that represents sub-scenes obtained by further segmenting the scenes based on the input scene information input to you." In this way, it indicates the processing content that the language model M or other language model should perform (the role of the language model M or other language model, or the type of output that the language model M or other language model should generate).
[0084] For example, language model M or another language model divides the input field information into tokens and calculates the embedding representation of each token based on parameters adjusted through learning. Based on the order of the token embedding representations, language model M or another language model generates more detailed input field information than the input field information itself as output, predicting subsequent text as needed. For example, language model M or another language model generates titles, etc., of sub-fields corresponding to the input field information itself based on default prompts, and outputs them as input field information. The input information acquisition unit 101 acquires the input field information output from language model M or another language model.
[0085] Additionally, the input information acquisition unit 101 can input the level of detail into the language model M or another language model. In this case, the language model M or other language model can also acquire input scene information, such as the title of the sub-scene, based on the level of detail input to itself.
[0086] [3-3. Script Generation Department]
[0087] The script generation unit 102 generates a script about a moving image by inputting input information into a language model M. The language model M divides the input information into tokens and calculates the embedding representation of each token based on parameters adjusted through learning. Based on the order of the token embedding representations and predicting subsequent text as needed, the language model M generates the script as an output. For example, the language model M generates and outputs a script corresponding to the input information based on default prompts. The input information acquisition unit 101 acquires the script output from the language model M.
[0088] For example, the script generation unit 102 generates a script by inputting input feature information into a language model M. The language model M segments the input feature information into tokens and calculates the embedding representation of each token based on parameters adjusted through learning. Based on the order of the token embedding representations and predicting subsequent text as needed, the language model M generates the script as an output. For example, the language model M generates and outputs a script corresponding to the input feature information based on default prompts. The input feature information acquisition unit acquires the script output from the language model M.
[0089] For example, the script generation unit 102 generates a script by inputting input setting information into a language model M. The language model M segments the input setting information into tokens and calculates the embedding representation of each token based on parameters adjusted through learning. Based on the order of the token embedding representations and predicting subsequent text as needed, the language model M generates the script as an output. For example, the language model M generates and outputs a script corresponding to the input setting information based on default prompts. The input setting information acquisition unit acquires the script output from the language model M.
[0090] For example, the script generation unit 102 generates a script by inputting the input scene information into a language model M. The language model M segments the input scene information into tokens and calculates the embedding representation of each token based on parameters adjusted through learning. Based on the order of the token embedding representations and predicting subsequent text as needed, the language model M generates the script as an output. For example, the language model M generates and outputs a script corresponding to the input scene information based on default prompts. The input scene information acquisition unit acquires the script output from the language model M.
[0091] For example, by inputting scene information corresponding to the level of detail into the language model M, the script generation unit 102 generates a script corresponding to the level of detail. The language model M segments the scene information corresponding to the level of detail into tokens and calculates the embedding representation of each token based on parameters adjusted through learning. Based on the order of the token embedding representations and predicting subsequent text as needed, the language model M generates the script as an output. For example, the language model M generates and outputs the script corresponding to the scene information of the level of detail based on default prompts. The scene information acquisition unit acquires the script output from the language model M.
[0092] For example, the script generation unit 102 generates a script portion for each scene by inputting the scene information into the language model M. This script portion is part of the script. For scenes after the second one in a plurality of scenes, the script portions of the scenes preceding the second one are input into the language model M. Script portions for scenes after the second one are generated, and the script is generated based on the script portions of each of the plurality of scenes. In this way, the default prompt can indicate that a script portion should be generated for each individual scene.
[0093] For example, the script generation unit 102 generates a script by inputting the first input information and the second input information into a language model M. The language model M segments the first and second input information into tokens and calculates the embedding representation of each token based on parameters adjusted through learning. Based on the order of the token embedding representations and predicting subsequent text as needed, the language model M generates the script as an output. For example, the language model M can generate and output a script corresponding to the first and second input information based on default prompts. The scene information acquisition unit acquires the script output from the language model M.
[0094] Furthermore, the script generation unit 102 can generate a script including all scenes at once, instead of generating script portions for each scene. In this case, a default prompt can be displayed indicating that the script should be generated all at once. Additionally, the script does not need to be specifically divided into multiple scenes. In this case, input scene information is not obtained. The script generation unit 102 can generate a script based solely on input feature information. The script generation unit 102 can generate a script based solely on input setting information. The script generation unit 102 can generate a script based solely on input scene information. The script generation unit 102 can generate a script based on input information. The script generation unit 102 can generate a script, regardless of the level of detail.
[0095] [4. Processing performed by the script generation system]
[0096] Figure 7This is a diagram illustrating an example of the processing performed by the script generation system 1. Figure 7 The processing of the script generation device 10 in the process executed by the script generation system 1 is shown. The control unit 11 executes the program stored in the storage unit 12. Figure 7 The processing. Figure 7 Each step in this process is an example of a step included in the script generation method of this disclosure. For example, when the person in charge of generating a script in the script generation service operates the script generation device 10 to designate a live stream as the object of script generation, the following steps are executed: Figure 7 The processing in the middle.
[0097] like Figure 7 As shown, the script generation device 10 obtains an orientation piece (S1) from the script database DB that is associated with the live stream ID of the live stream that becomes the script generation object. The script generation device 10 obtains input feature information based on the orientation piece (S2). In S2, the script generation device 10 obtains information about items representing features of goods or services that are input into the orientation piece as input feature information.
[0098] The script generation device 10 retrieves content about goods or services introduced in the live broadcast based on the input feature information obtained in S2 (S3). In S3, the script generation device 10 retrieves well-known online search services by using information such as the name of the goods or services shown in the input feature information as a search query. The script generation device 10 obtains extracted information from the content retrieved in S3 (S4). The script generation device 10 obtains input feature information by inputting a default prompt indicating the generation of input feature information and the extracted information obtained in S4 into the language model M (S5). In S5, the language model M segments these into tokens and outputs the input feature information according to the arrangement of the token embeddings. The script generation device 10 obtains the input feature information output from the language model M.
[0099] Furthermore, if the orientation sheet contains a sufficient amount of input feature information, the processing steps S3 to S5 may not need to be performed. For example, the person in charge of generating the script can specify whether the processing steps S3 to S5 need to be performed by operating the operation unit 14. When information is input into all items or more than a predetermined number of items representing the features of a product or service in the orientation sheet, the script generation device 10 does not need to perform the processing steps S3 to S5.
[0100] The script generation device 10 obtains input setting information based on the orientation piece (S6). In S6, if the orientation piece contains input setting information, the script generation device 10 obtains the input setting information contained in the orientation piece. If the orientation piece does not contain input setting information, or if the orientation piece does not contain a sufficient amount of input setting information, the script generation device 10 inputs the default prompt representing the generation of input setting information and the information contained in the orientation piece into the language model M. The language model M segments these into tokens and outputs the input setting information according to the arrangement of the token embeddings. The script generation device 10 obtains the input setting information output from the language model M. In S6, the script generation device 10 can also obtain information such as a live broadcast summary as input setting information.
[0101] The script generation device 10 obtains input scene information based on the orientation piece (S7). In S7, if the orientation piece contains input scene information, the script generation device 10 obtains the input scene information contained in the orientation piece. If the orientation piece does not contain input scene information, or if the orientation piece does not contain a sufficient amount of input scene information, the script generation device 10 inputs the default prompt representing the generation of input scene information and the information contained in the orientation piece into the language model M. The language model M segments these into tokens and outputs the input scene information according to the arrangement of the token embeddings. The script generation device 10 obtains the input scene information output from the language model M.
[0102] The script generation device 10 determines, based on the orientation sheet, which version of the script to be generated—a simplified version or a detailed version (S8). In S8, the script generation device 10 determines whether the level of detail of the script included in the orientation sheet represents a simplified version or a detailed version. If the orientation sheet does not contain any level of detail for the script, the person in charge of generating the script can specify the level of detail by operating the operation unit 14. The script generation device 10 can perform the determination in S8 based on the level of detail of the script specified by the person in charge.
[0103] In S8, if it is determined that a simplified version of the script is to be generated (S8: simplified version), the script generation device 10 generates the script portion of the initial scene by inputting input information including input feature information, input setting information, and input scene information into the language model M (S9). In S9, the script generation device 10 also inputs a default prompt to the language model M indicating the generation of a script portion corresponding to the scene. The language model M segments the input information into tokens and outputs the script portion of the initial scene based on the permutation of the embedding expressions of the tokens. Since there are no other scenes before the initial scene, the language model M outputs the script portion of the initial scene without basing it on the script portions of other scenes. The script generation device 10 obtains the script portion of the initial scene output from the language model M.
[0104] The script generation device 10 generates the script portion for the next scene by inputting input information, including input feature information, input setting information, and input scene information, along with the script portion of the completed scene, into the language model M (S10). In S10, the script generation device 10 also inputs a default prompt to the language model M indicating the generation of the script portion corresponding to the next scene. The language model M segments the input information and script portion into tokens and outputs the script portion for the next scene based on the permutation of the token embeddings. For example, the script generation device 10 inputs the script portions of all scenes generated so far into the language model M. The script generation device 10 then obtains the script portion for the next scene output from the language model M.
[0105] Based on the input scene information, the script generation device 10 determines whether the script portion for the final scene has been generated (S11). If it is determined in S11 that the script portion for the final scene has not been generated (S11: No), the process returns to S10, and the script portion for the next scene is generated. If it is determined that the script portion for the final scene has been generated (S11: Yes), the script generation device 10 generates the final script based on the script portions for each scene (S12), and this process ends. In S12, the script generation device 10 generates the final script by combining the script portions for each scene. The script generation device 10 saves the final script in the script database DB.
[0106] In S8, if a detailed version of the script is determined to be required (S8: detailed version), the script generation device 10 generates sub-scenes according to each scene represented by the input scene information, thereby obtaining input scene information representing the sub-scenes of each of the multiple scenes (S13). In S13, the script generation device 10 inputs the default prompts representing the generation of sub-scenes for each scene and the input scene information into the language model M. The language model M segments these into tokens and outputs the input scene information representing the sub-scenes based on the permutation of the token embeddings. The script generation device 10 obtains the input scene information output from the language model M.
[0107] The script generation device 10 generates the script portion of the initial sub-scene of the scene to be processed by inputting input information, including input feature information, input setting information, and input scene information, into the language model M (S14). The scenes to be processed are the scenes of the object that form a cycle from S14 to S17. The scenes to be processed are selected sequentially starting from the initial scene. In S14, the script generation device 10 also inputs a default prompt to the language model M indicating the generation of the script portion corresponding to the sub-scene. The language model M divides the input information into tokens, etc., and outputs the script portion of the initial sub-scene based on the permutation of the embedding expressions of the tokens. Since there are no other sub-scenes before the initial sub-scene, the language model M outputs the script portion of the initial sub-scene without based on the script portions of other sub-scenes. The script generation device 10 obtains the script portion of the initial sub-scene output from the language model M.
[0108] Regarding the scene being processed, the script generation device 10 generates the script portion for the next sub-scene of the scene being processed, based on input information including input feature information, input setting information, and input scene information, as well as the script portion of the completed sub-scene (S15). In S15, the script generation device 10 also inputs a default prompt indicating the generation of the script portion corresponding to the next sub-scene into the language model M. The language model M segments the input information and script portions into tokens, and outputs the script portion of the next sub-scene based on the permutation of the embedding expressions of the tokens. For example, the script generation device 10 inputs the script portions of all sub-scenes generated so far into the language model M. The script generation device 10 then obtains the script portion of the next sub-scene output from the language model M.
[0109] The script generation device 10 determines, based on the input scene information, whether to generate the script portion of the last sub-scene of the scene to be processed (S16). If it is not determined that the script portion of the last sub-scene of the scene to be processed should be generated (S16: No), the process returns to S15, and the script portion of the next sub-scene of the scene to be processed is generated. If it is determined that the script portion of the last sub-scene of the scene to be processed should be generated (S16: Yes), the script generation device 10 determines, based on the input scene information, whether to generate the script portion of the last scene (S17).
[0110] If it is determined in S17 that the script portion for the final scene has not been generated (S17: No), the process returns to S14, and the script portion for the initial sub-scene of the next scene is generated. That is, the next scene becomes the scene to be processed. If it is determined that the script portion for the final scene has been generated (S17: Yes), the script generation device 10 generates the final script based on the script portions for each sub-scene of each scene (S18), and this process ends. In S18, the script generation device 10 generates the final script by combining the script portions for each sub-scene of each scene. The script generation device 10 stores the final script in the script database DB.
[0111] [5. Summary of the script generation system]
[0112] The script generation system 1 according to this embodiment acquires input information related to dynamic images introducing goods or services. The script generation system 1 generates a script for the dynamic images by inputting the input information into a language model M. By using the language model M, the script generation system 1 can generate flexible scripts based on the input information. For example, when a person responsible for script generation services generates a script based on a template, a flexible script cannot be generated because the script can only be generated within the scope of the template. However, the language model M can generate more flexible scripts by performing flexible natural language processing based on the input information. Because the script generation system 1 can reduce the workload of the person in charge, the cost of generating scripts can be reduced. For example, if the person responsible for script generation services outsources the script generation to an external vendor, the external vendor incurs costs, but the script generation system 1 can avoid such costs.
[0113] Furthermore, the script generation system 1 generates scripts by inputting input feature information into a language model M. Thus, the language model M can generate flexible scripts corresponding to the features of goods or services, thereby improving the accuracy of the scripts produced by the language model M. For example, when the input feature information indicates appearance features, the script generation system 1 can generate a script corresponding to the appearance features. When the input feature information represents functional features, the script generation system 1 can generate a script corresponding to the functional features.
[0114] Furthermore, the script generation system 1 acquires input feature information, which is generated by inputting extracted information from content related to the features of goods or services into a language model M or another language model. Therefore, since the script generation system 1 can generate scripts based on more features, the accuracy of the scripts produced by the language model M can be improved. For example, even if there is no input feature information in the orientation piece, or if there is not a sufficient amount of input feature information in the orientation piece, the script generation system 1 can still obtain input feature information from the content.
[0115] Furthermore, the script generation system 1 generates scripts by inputting input setting information into a language model M. Because the language model M can generate flexible scripts corresponding to the settings of a moving image, the script generation system 1 can improve the accuracy of the scripts generated by the language model M. For example, when the input setting information represents the release date and time of the moving image, the script generation system 1 can generate a number of scripts corresponding to the release date and time of the moving image. When the input setting information represents the length of the moving image, the script generation system 1 can generate a number of scripts corresponding to the length of the moving image.
[0116] Furthermore, the script generation system 1 generates a script by inputting scene information into a language model M. Because the language model M can generate flexible scripts corresponding to the scenes in the motion graphics, the script generation system 1 can improve the accuracy of the scripts produced by the language model M. For example, if the language model M tries... Figure 1 Generating a large number of products at once may reduce the accuracy of the products. In this regard, when the language model M generates the script for each scene, the amount of products generated by the language model M at one time can be suppressed, so the script generation system 1 can improve the accuracy of the script.
[0117] Furthermore, by inputting scene information corresponding to the level of detail into the language model M, the script generation system 1 generates a script corresponding to the level of detail. Because the language model M can generate flexible scripts corresponding to different levels of detail, the script generation system 1 can improve the accuracy of the scripts generated by the language model M. For example, the script generation system 1 can generate a simplified version of the script requested by a store, sufficient for a store to represent the overall progress of the live stream. In this case, the processing load on the script generation device 10 can be reduced because the amount of text generated by the language model M can be reduced. If the language model M is stored in a computer other than the script generation device 10, the processing load on that computer can be reduced. For example, the script generation system 1 can generate a detailed version of the script requested by a store that wants a detailed version of the script representing the detailed progress of the live stream.
[0118] Furthermore, the script generation system 1 inputs the input scene information for each scene into the language model M, thereby generating the script portion for that scene. For scenes after the second one in a series of scenes, the script generation system 1 also inputs the script portions of the scenes preceding the second one into the language model M, thus generating the script portions for the second and subsequent scenes. The script generation system 1 generates a script based on the script portions of each of the multiple scenes. Therefore, the script generation system 1 can generate scripts with connections between scenes. For example, when the language model M attempts... Figure 1Generating a large amount of text at once may increase the processing load on the script generation device 10. However, when the language model M generates subdivided script segments for each scene, the processing load on the script generation device 10 can be reduced. Furthermore, when the language model M attempts to... Figure 1 When generating large amounts of text repeatedly, the accuracy of the output may decrease. However, by generating subdivided script segments for each scene using a language model M, script generation system 1 can improve the accuracy of the final script output.
[0119] Furthermore, the script generation system 1 obtains pre-prepared first input information and second input information generated by inputting the first input information into a language model M or another language model. The script generation system 1 generates a script by inputting the first and second input information into the language model M. Therefore, since the script generation system 1 can input a wide variety of rich input information into the language model M, the accuracy of the script can be improved.
[0120] [6. Variations]
[0121] This disclosure is not limited to the embodiments described above. This disclosure can be appropriately modified without departing from its spirit.
[0122] Figure 8 This diagram illustrates one example of the functions implemented in the modified example. For example, in the script generation apparatus 10 of the modified example, a question answer generation unit 103, a first sample acquisition unit 104, and a second sample acquisition unit 105 are implemented. The question answer generation unit 103, the first sample acquisition unit 104, and the second sample acquisition unit 105 are each implemented by the control unit 11.
[0123] [6-1. Variation Example 1]
[0124] For example, in Figure 2 In the example of the introduction screen SC, the live streaming service accepts suggestions from viewers. Viewers sometimes input questions as suggestions related to the products or services being introduced by the actors. Actors can answer these questions during the live stream. In Variation 1, the content of the script generated by the script generation unit 102 is used to generate hypothetical questions from the viewers and answers to those questions, allowing actors to prepare for the hypothetical questions.
[0125] The script generation system 1 in Variation 1 includes a question-and-answer generation unit 103. The question-and-answer generation unit 103 generates questions and answers related to motion images by inputting a script into a language model M or another language model. The question-and-answer generation unit 103 can also input other information besides the script into the language model M or other language model. For example, the question-and-answer generation unit 103 can input at least one of input feature information, input setting information, and input scene information into the language model M or other language model.
[0126] Questions and answers are data containing at least one question and at least one answer to that question. Questions and answers can be in any data format. For example, questions and answers can be in the data format of a table calculation software, CSV format, text format, document format, or other formats. The question-answer generation unit 103 is capable of generating any number of questions and answers. The number of questions and answers that the question-answer generation unit 103 should generate can be pre-specified or not specifically specified.
[0127] For example, if the language model M or other language model is not specifically designed for generating questions and answers, the question-answer generation unit 103 may input a default prompt indicating that questions and answers should be generated into the language model M or other language model. The default prompt could be something like, "You are an AI that generates questions and answers based on the script input into you." This indicates the processing that the language model M or other language model should perform (the function of the language model M or other language model, or the type of output that the language model M or other language model should generate).
[0128] For example, language model M or another language model segments the script input to itself into tokens and computes the embedding representation (feature vector) of each token based on parameters adjusted through learning. Based on the order of the token embedding representations, language model M or another language model generates questions and answers as outputs, predicting subsequent text as needed. For example, language model M or another language model generates and outputs questions and answers corresponding to the script input to itself based on default prompts. Question and answer generation unit 103 obtains the questions and answers output from language model M or another language model.
[0129] For example, the question-and-answer generation unit 103 stores the questions and answers in the script database DB. The server 20 sends the questions and answers stored in the script database DB to the actor device 30. The question-and-answer generation unit 103 may also store the questions and answers in a database other than the script database DB. The question-and-answer generation unit 103 may record the questions and answers on a computer other than the script generation device 10, or on an external information storage medium. The questions and answers generated by the question-and-answer generation unit 103 may be sent to a computer other than the actor device 30.
[0130] The script generation system 1 in Variation 1 generates questions and answers related to motion graphics by inputting a script into a language model M or another language model. The script generation system 1 can generate questions and answers corresponding to the script to assist the actors' progress.
[0131] [6-2. Variation Example 2]
[0132] For example, in this embodiment, an example of inputting feature information into the language model M is given. Any information that can be used to generate a script can be input into the language model M. In Variation 2, an example of inputting a script sample into the language model M is given. The sample is script data used as a reference for the language model M. Similar to the script generated by the script generation unit 102, the sample can have any data format. For example, the sample can be a script format or template. The sample can be generated manually or by the language model M. The sample can also be a sample obtained by transcribing the sound of a moving image.
[0133] The script generation system 1 in Modification 2 includes a first sample acquisition unit 104. The first sample acquisition unit 104 acquires samples related to other moving images, in which different goods are described compared to those described in the moving images that are the objects of script generation, or different services are described compared to those described in the moving images that are the objects of script generation, and the sample omits features of the other goods or services. The other moving images are moving images other than those that are the objects of script generation. The other moving images may be moving images that have been previously published, or moving images for which a script has been generated but not yet published. The other moving images may also be moving images published through services other than live streaming services. The other goods or services are goods or services described in the other moving images.
[0134] In Variation 2, taking the case where the data storage unit 100 stores a sample as an example, the first sample acquisition unit 104 acquires the sample from the data storage unit 100. The sample can be stored in a computer other than the script generation device 10, or in an external information storage medium. In this case, the first sample acquisition unit 104 can acquire the sample from another computer or an external information storage medium.
[0135] For example, the first sample acquisition unit 104 acquires samples from a script previously generated by the script generation unit 102 where parts of the words included in the input feature information used during script generation are masked. Masking is hiding specific parts (e.g., filling the part with spaces or specific symbols). The first sample acquisition unit 104 can perform masking on the script based on this input feature information to acquire samples. The first sample acquisition unit 104 can also acquire samples after masking through other functions besides the first sample acquisition unit 104.
[0136] Furthermore, the method of omitting features of other goods or services is not limited to masking. For example, the first sample acquisition unit 104 can acquire samples by deleting a portion of the input feature information used when generating the script in a script previously generated by the script generation unit 102. The first sample acquisition unit 104 can also acquire samples with this portion deleted through other functions besides the first sample acquisition unit 104.
[0137] In Variation 2, the script generation unit 102 generates a script by further inputting samples into the language model M. Specifically, the script generation unit 102 inputs the input information and samples described in the embodiment into the language model M. The language model M segments the input information and samples into tokens and calculates the embedding representation of each token based on parameters adjusted through learning. Based on the order of the token embedding representations, the language model M generates a script corresponding to the sample as output, predicting subsequent text as needed. For example, the language model M can generate and output a script corresponding to the sample based on a default prompt. The default prompt is such as "Please refer to this sample to generate a script." This indicates that the language model M should refer to the text of the sample. The script generation unit 102 obtains the script output from the language model M.
[0138] The script generation system 1 in Variation 2 generates a script by further inputting samples into the language model M, omitting features of other products or services. Therefore, the language model M is less likely to be influenced by features of other products or services, and thus, script generation system 1 can generate a script that reflects the features of the goods or services presented in the dynamic image that becomes the subject of script generation. That is, script generation system 1 can improve the accuracy of the script.
[0139] [6-3. Variation Example 3]
[0140] For example, an actor's speaking flow and style can possess characteristics unique to that actor. If a script is generated that matches these characteristics, it's assumed the actor can easily perform a live broadcast. Therefore, scripts for other animated images, performed by the same actor as the actor in the animated image that becomes the subject of script generation, can be obtained as samples. In Variation 3, it's assumed that the scripts for other animated images and the actors in those other animated images are stored in a script database DB.
[0141] The script generation system 1 in Modification 3 includes a second sample acquisition unit 105. The second sample acquisition unit 105 acquires samples related to other moving images featuring the same actor as the actor in the moving image. The meaning of the samples is the same as in Modification 2. For example, the second sample acquisition unit 105 determines the actor in the moving image that becomes the subject of script generation based on orientation images stored in the script database DB. The second sample acquisition unit 105, based on orientation images stored in the script database DB, determines other moving images featuring the same actor as the determined actor. The second sample acquisition unit 105 acquires the scripts containing the other moving images of the determined same actor as samples.
[0142] Furthermore, when multiple other animated images of the same actor exist, the second sample acquisition unit 105 can acquire the script of all other animated images of the same actor as a sample, or it can acquire the script of a portion of other animated images of the same actor as a sample. Additionally, in the case of combined variations 2 and 3, the second sample acquisition unit 105 can acquire the script of other animated images of the same actor who is the object of script generation, omitting information about other goods or services. The second sample acquisition unit 105 can acquire the script of animated images of the same actor published through services other than live streaming services as a sample.
[0143] In Variation 3, the script generation unit 102 generates a script by further inputting samples into the language model M. Specifically, the script generation unit 102 inputs the input information and samples described in the embodiment into the language model M. The language model M segments the input information and samples into tokens and calculates the embedding representation of each token based on parameters adjusted through learning. Based on the order of the token embedding representations, the language model M generates a script corresponding to the sample as output, predicting subsequent text as needed. For example, the language model M can also generate and output a script corresponding to the sample based on the same default prompts as in Variation 2. The script generation unit 102 obtains the script output from the language model M.
[0144] In variation 3, the script generation system 1 generates a script by further inputting samples related to other moving images of the same actor into the language model M. Therefore, the script generation system 1 can reflect the characteristics of the actors in the script. That is, the script generation system 1 can improve the accuracy of the script.
[0145] [6-4. Other variations]
[0146] For example, the above variations can also be combined.
[0147] For example, the example given is the application of script generation system 1 to script generation services provided within a live streaming service; however, script generation system 1 can also be applied to other scenarios. Script generation system 1 can generate scripts for animated images other than those used in live streaming services. For instance, script generation system 1 can generate scripts for animated images not used in live streaming, scripts for animated images representing advertisements for goods or services, scripts for animated images featuring AI (artificial speech) rather than human speech, scripts for animated images introducing goods or services sold through services other than e-commerce services, or scripts for other animated images.
[0148] For example, script generation system 1 can generate scripts with animated images introducing goods or services other than e-commerce services. Script generation system 1 can generate scripts with animated images that introduce goods or services sold in physical stores, services offered in accommodation facilities such as hotels, communication services, payment services, financial services, e-book services, services offered in beauty salons or restaurants, or goods or services introduced on social networking sites (SNS). The goods introduced in the animated images are not limited to tangible objects; they can also be content such as music or movies.
[0149] For example, the functions described as being implemented by the script generation apparatus 10 can be implemented by the server 20, the actor apparatus 30, or other computers. The processing described as being implemented by the script generation apparatus 10 can be shared by multiple computers.
[0150] [7. Postscript]
[0151] For example, the script generation system according to this disclosure may have the following structure. (1)
[0153] A script generation system, comprising:
[0154] An input information acquisition unit acquires input information, which is input into a learned language model and is related to a dynamic image describing a product or service. The learned language model is capable of generating a product described in natural language.
[0155] The script generation unit generates a script related to the motion picture by inputting the input information into the language model. (2)
[0157] According to the script generation system described in (1), wherein,
[0158] The input information acquisition unit acquires input feature information, which is input information related to the characteristics of the product or the service.
[0159] The script generation unit generates the script by inputting the input feature information into the language model. (3)
[0161] According to the script generation system described in (2), wherein,
[0162] The input information acquisition unit acquires the input feature information generated by inputting extracted information extracted from content related to the features of the product or service into the language model or other language model. (4)
[0164] According to any one of (1) to (3) of the script generation system, wherein,
[0165] The input information acquisition unit acquires input setting information, which is input information related to the settings of the dynamic image.
[0166] The script generation unit generates the script by inputting the input setting information into the language model. (5)
[0168] According to any one of (1) to (4) of the script generation system, wherein,
[0169] The input information acquisition unit acquires input scene information, which is the input information related to each of the multiple scenes in the dynamic image.
[0170] The script generation unit generates the script by inputting the input scene information into the language model. (6)
[0172] According to the script generation system described in (5), wherein,
[0173] The input information acquisition unit acquires the input scene information corresponding to the level of detail of the script.
[0174] The script generation unit generates the script corresponding to the level of detail by inputting the input scene information corresponding to the level of detail into the language model. (7)
[0176] According to the script generation system described in (5) or (6), wherein,
[0177] The script generation department
[0178] For each scene, the input scene information for that scene is input into the language model, thereby generating a script portion for that scene, which is a part of the script.
[0179] For each of the multiple scenes from the second to the subsequent scenes, the script portions of the scenes preceding the second to the subsequent scenes are also input into the language model, thereby generating the script portions of the second to subsequent scenes.
[0180] The script is generated based on the script portion of each of the plurality of scenes. (8)
[0182] According to any one of (1) to (7) of the script generation system, wherein,
[0183] The input information acquisition unit acquires first input information and second input information. The first input information is pre-prepared input information, and the second input information is input information generated by inputting the first input information into the language model or other language model.
[0184] The script generation unit generates the script by inputting the first input information and the second input information into the language model. (9)
[0186] According to any one of (1) to (8) of the script generation system, wherein,
[0187] The script generation system also includes a question and answer generation unit, which generates questions and answers related to the motion picture by inputting the script into the language model or other language model. (10)
[0189] According to any one of (1) to (9) of the script generation system, wherein,
[0190] The script generation system further includes a first sample acquisition unit, which acquires samples related to other moving images. These other moving images introduce other products or services that are different from the stated product, and the samples omit features of these other products or services.
[0191] The script generation unit generates the script by further inputting the sample into the language model. (11)
[0193] The script generation system according to any one of (1) to (10), wherein,
[0194] The script generation system further includes a second sample acquisition unit, which acquires samples related to other moving images, wherein the other moving images are moving images performed by the same actors as the actors in the moving images.
[0195] The script generation unit generates the script by further inputting the sample into the language model.
Claims
1. A script generation system, comprising: The input information acquisition unit acquires input information, which is input into the learned language model and is related to a dynamic image introducing a product or service. The learned language model is able to generate a product described in natural language. as well as The script generation unit generates a script related to the motion picture by inputting the input information into the language model.
2. The script generation system according to claim 1, wherein, The input information acquisition unit acquires input feature information, which is input information related to the characteristics of the product or the service. The script generation unit generates the script by inputting the input feature information into the language model.
3. The script generation system according to claim 2, wherein, The input information acquisition unit acquires the input feature information generated by inputting extracted information extracted from content related to the features of the product or service into the language model or other language model.
4. The script generation system according to any one of claims 1 to 3, wherein, The input information acquisition unit acquires input setting information, which is input information related to the settings of the dynamic image. The script generation unit generates the script by inputting the input setting information into the language model.
5. The script generation system according to any one of claims 1 to 3, wherein, The input information acquisition unit acquires input scene information, which is the input information related to each of the multiple scenes in the dynamic image. The script generation unit generates the script by inputting the input scene information into the language model.
6. The script generation system according to claim 5, wherein, The input information acquisition unit acquires the input scene information corresponding to the level of detail of the script. The script generation unit generates the script corresponding to the level of detail by inputting the input scene information corresponding to the level of detail into the language model.
7. The script generation system according to claim 5, wherein, The script generation unit inputs the input scene information for each scene into the language model, thereby generating a script portion for that scene, which is a part of the script. For the second and subsequent scenes among the plurality of scenes, the script generation unit also inputs the script portions of the scenes preceding the second and subsequent scenes into the language model, thereby generating the script portions of those second and subsequent scenes. The script generation unit generates the script based on the script portions of each of the multiple scenes.
8. The script generation system according to any one of claims 1 to 3, wherein, The input information acquisition unit acquires first input information and second input information. The first input information is pre-prepared input information, and the second input information is input information generated by inputting the first input information into the language model or other language model. The script generation unit generates the script by inputting the first input information and the second input information into the language model.
9. The script generation system according to any one of claims 1 to 3, wherein, The script generation system also includes a question and answer generation unit, which generates questions and answers related to the motion picture by inputting the script into the language model or other language model.
10. The script generation system according to any one of claims 1 to 3, wherein, The script generation system further includes a first sample acquisition unit, which acquires samples related to other moving images. These other moving images introduce other products or services that are different from the stated product, and the samples omit features of these other products or services. The script generation unit generates the script by further inputting the sample into the language model.
11. The script generation system according to any one of claims 1 to 3, wherein, The script generation system further includes a second sample acquisition unit, which acquires samples related to other moving images, wherein the other moving images are moving images performed by the same actors as the actors in the moving images. The script generation unit generates the script by further inputting the sample into the language model.
12. A script generation method, comprising: The input information acquisition step involves acquiring input information, which is then input into the learned language model and is related to dynamic images introducing products or services. The learned language model is capable of generating products described in natural language. as well as The script generation step involves inputting the input information into the language model to generate a script related to the motion picture.
13. A program that causes a computer to function as: An input information acquisition unit acquires input information, which is input into a learned language model and is related to a dynamic image describing a product or service. The learned language model is capable of generating a product described in natural language. The script generation unit generates a script related to the motion picture by inputting the input information into the language model.
Citation Information
Patent Citations
Advertisement distribution system, advertisement creating device, content creating engine creating device, script creating device, viewing device, and advertisement distributing program
JP2008022256A