Information processing system, information processing method, and program

The information processing system clarifies user intent through dialogues and generates content based on user input and context, addressing the challenge of skill requirements in content creation and intent matching.

WO2026028751A1PCT designated stage Publication Date: 2026-02-05SONY GROUP CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/024590
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-31
Filing Date
2025-07-09
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Users face challenges in creating content, such as video editing, due to the need for high skills and experience, limiting content creation to a small number of users, and the difficulty in generating content that matches their intent using generative AI, especially for beginners.

Method used

An information processing system that evaluates the clarity of user intention through user input information, generates dialogues to clarify intent when unclear, and generates content based on clear intent using a combination of user input and context information.

Benefits of technology

Enables users to easily create content that aligns with their intentions by clarifying their input through dialogues, reducing the need for advanced editing skills and expanding content creation opportunities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025024590_05022026_PF_FP_ABST
    Figure JP2025024590_05022026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing system comprises: an evaluation unit that evaluates a degree of clarity of an intention of a user regarding generation of content indicated by user input information; a dialogue generation unit that generates a dialog for obtaining additional user input information when an evaluation that the intention of the user is not clear is obtained; and a content generation unit that generates content on the basis of the user input information when an evaluation that the intention of the user is clear is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system, information processing method, and program

[0001] The present technology relates to an information processing system, an information processing method, and a program, and in particular to an information processing system, an information processing method, and a program that enable a user to more easily obtain content that is closer to his or her intention.

[0002] BACKGROUND ART In recent years, users have been creating a lot of content, such as video editing (moving image editing) using gameplay as material, for the purpose of posting to SNS (Social Networking Service) or the like.

[0003] For example, one technology proposed for creating such content is to acquire information about the relationship between the poster of the content and the viewers of the content, and generate comments for a virtual commentator to utter based on the information about the relationship (see, for example, Patent Document 1). This technology makes it possible to obtain video content in which the virtual commentator's comments are added to game footage.

[0004] International Publication No. 2022 / 249522

[0005] However, when a user edits a video, they need to have high skills in using video editing tools and come up with content ideas, so content creation depends heavily on the user's experience and editing sense. For this reason, the use of video editing tools is limited to a small number of users.

[0006] As such, creating content is a high hurdle for beginners and it is not something they can easily get started on, which is why the number of people posting content is not expanding.

[0007] Additionally, while image generation from natural language or text using generative AI (Artificial Intelligence) is becoming increasingly popular, it is often difficult to generate content that matches the user's intent when the user's input is abstract.

[0008] Specifically, for example, when editing video, i.e., creating content, the user will need to think about and input the information necessary to create the content, such as what scenes to extract and the purpose of the video they intend to create.

[0009] However, in such cases, the amount of information to be entered becomes enormous, which not only increases the burden on the user, but also makes it difficult for the user to obtain the content they intended if the information entered is insufficient.

[0010] The present technology has been made in view of such circumstances, and makes it possible to more easily obtain content that is closer to the user's intention.

[0011] An information processing system according to a first aspect of the present technology includes an evaluation unit that evaluates the degree of clarity of a user's intention regarding content generation, as indicated by user input information; a dialogue generation unit that generates a dialogue to obtain additional user input information when the evaluation indicates that the user's intention is unclear; and a content generation unit that generates the content based on the user input information when the evaluation indicates that the user's intention is clear.

[0012] An information processing method or program according to a first aspect of the present technology includes evaluating the degree of clarity of a user's intention regarding content generation as indicated by user input information, generating a dialogue to obtain additional user input information if the evaluation indicates that the user's intention is unclear, and generating the content based on the user input information if the evaluation indicates that the user's intention is clear.

[0013] A first aspect of the present technology includes evaluating the degree of clarity of a user's intention regarding content generation as indicated by user input information, generating a dialogue to obtain additional user input information if the evaluation indicates that the user's intention is unclear, and generating the content based on the user input information if the evaluation indicates that the user's intention is clear.

[0014] An information processing system according to a second aspect of the present technology includes a content generation unit that generates content based on material data that serves as the raw material for the content; an evaluation unit that, when a user inputs user input information regarding modification of the content, evaluates the degree of clarity of the user's intention regarding modification of the content as indicated by the user input information; and a dialogue generation unit that, when the evaluation indicates that the user's intention is not clear, generates a dialogue to obtain additional user input information. When the evaluation indicates that the user's intention is clear, the content generation unit generates the content that reflects the user's intention to modify the content based on the material data and the user input information.

[0015] An information processing method or program according to a second aspect of the present technology includes generating content based on material data that is the raw material of the content; when a user inputs user input information regarding modification of the content, evaluating the degree of clarity of the user's intention regarding modification of the content as indicated by the user input information; when the evaluation indicates that the user's intention is unclear, generating a dialogue to obtain additional user input information; and when the evaluation indicates that the user's intention is clear, generating the content that reflects the user's intention to modify based on the material data and the user input information.

[0016] A second aspect of the present technology includes generating content based on material data that is the raw material of the content; when a user inputs user input information regarding modification of the content, evaluating the clarity of the user's intention regarding modification of the content as indicated by the user input information; when the evaluation indicates that the user's intention is unclear, generating a dialogue to obtain additional user input information; and when the evaluation indicates that the user's intention is clear, generating the content that reflects the user's intention to modify based on the material data and the user input information.

[0017] FIG. 1 is a diagram illustrating a flow of content production. FIG. 2 is a diagram illustrating an example of the configuration of an information processing system. FIG. 3 is a flowchart illustrating content generation processing. FIG. 4 is a diagram illustrating calculation of a clarity score. FIG. 5 is a diagram illustrating generation of a dialogue. FIG. 6 is a diagram illustrating use of context information. FIG. 7 is a diagram illustrating generation of content. FIG. 8 is a diagram illustrating modification of content. FIG. 9 is a diagram illustrating selection of a template and a dramatic effect. FIG. 10 is a diagram illustrating an example of input of user input information. FIG. 11 is a diagram illustrating an example of presentation of content. FIG. 12 is a diagram illustrating an example of setting a clarity degree. FIG. 13 is a diagram illustrating adjustment of a required level based on a modification history. FIG. 14 is a diagram illustrating another example of the configuration of an information processing system. FIG. 15 is a flowchart illustrating content generation processing. FIG. 16 is a diagram illustrating an example of the configuration of computer hardware.

[0018] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.

[0019] <First embodiment> <About the present technology> When generating content based on content material, the present technology evaluates whether user input information input by a user sufficiently contains information necessary for generating the content, in other words, whether the user's intention regarding content generation is clear, and by conducting a dialogue based on the evaluation result, it becomes possible to more easily obtain content that is closer to the user's intention.

[0020] That is, in this technology, when the evaluation result of the user input information indicates that the user's intention regarding content generation is not sufficiently clear (that is, that necessary information is not sufficiently included), a dialogue prompting the user to input additional information required for content generation is generated, and the dialogue is presented to the user. In other words, a dialogue with the user is held.

[0021] When the user inputs additional user input information in response to the presentation of the dialogue, the additional user input information and the previously input user input information are evaluated, and the dialogue is repeatedly generated and presented until an evaluation result is obtained that indicates that the user's intention is sufficiently clear.

[0022] Then, if the evaluation result indicates that the user's intention is sufficiently clear, content is generated in accordance with the user's intention based on the user input information and content material that have been input so far.

[0023] In this way, with this technology, the user can obtain content that is close to what the user intended, simply by providing input in response to the presented dialogue. In other words, the user can obtain content that is close to what the user intended more easily.

[0024] The content to be generated by this technology may be any type of content, such as moving images (video content), still images, audio content, text content, etc. Below, as a specific example, a case will be described in which video content with audio is generated using video (moving images) and audio data of game play as content materials.

[0025] For example, video editing using game footage and other materials requires editing skills, composition ideas, and the ability to use editing tools, making it difficult for beginners to create content through video editing.

[0026] Therefore, in the present technology, the user's intention for video editing is input by utilizing natural language input from the user. That is, user input information is acquired through the above-described dialogue.

[0027] In addition, in order to extract and clarify the user's editing intentions, the clarity of the user's spoken text is scored, and a dialogue (dialogue flow) is generated that prompts the user for further input based on the resulting clarity score.

[0028] This technology combines a mechanism that appropriately extracts the user's intention from one or more user inputs, by conducting appropriate dialogue based on the clarity score, with text generation technology such as LLM (Large Language Models), making it possible to accurately obtain the user's editing intention.

[0029] Furthermore, this technology utilizes context information from gameplay, i.e., user input other than speech, video analysis, etc., to specify video scenes, which is difficult to do using natural language. For example, context information here refers to game operations, user emotional expressions, user vital information, etc.

[0030] By utilizing context information, this technology can provide user-friendly editing operations. In other words, even without explicit user input, scene extraction, effect addition, background music selection, and other operations are performed in line with the user's intentions, enabling video editing that closely matches the user's intent.

[0031] With this technology, even users (game players) without video editing skills can easily create fun and cool game videos, which was previously limited to users with advanced video editing skills. This will invigorate the gaming community and the UGC (User Generated Content) area, accelerating the growth of the game title business.

[0032] Furthermore, by utilizing voice and natural language, this technology allows users to create content without interrupting their gameplay, and allows them to work on another personal computer (PC) without having to leave the game. This can improve game engagement. Furthermore, video production (content production) can increase the amount of time users spend immersed in the game (game title), which can lead to increased gameplay time and increased purchases of the next title.

[0033] Now, the present technology will be described in more detail below.

[0034] First, the flow of content creation according to the present technology will be described with reference to FIG.

[0035] For example, as indicated by arrow Q11, when a user plays a desired game, the video footage of the game being played is stored as material used to generate content, that is, content material (material data).

[0036] After playing the game, as indicated by arrow Q12, when the user speaks to the game device "AAA" along with a call phrase to instruct it to generate a highlight video of the game play, the content of the speech is evaluated.

[0037] Specifically, in the game device "AAA," the user's utterance "Hey AAA, Create a nice highlight of my game play!" is acquired as user input information that is an input indicating the user's intention to generate content.

[0038] Then, based on the user input information, a clarity score indicating the clarity of the user's intention to instruct content generation indicated by the user input information, i.e., the user's intention regarding content generation, is calculated, and threshold processing is performed on the clarity score. In other words, an evaluation is performed to determine whether the user input information entered by the user sufficiently contains information indicating the user's intention necessary for content generation.

[0039] Here, for example, the clearer the user's intention regarding content generation as indicated by the user input information, the larger the clarity score value. In other words, the more abstract the user's intention as indicated by the user input information and the more insufficient the information required for content generation, the smaller the clarity score value.

[0040] Furthermore, if the clarity score is equal to or greater than a predetermined threshold, it is determined that the user's intention indicated by the user input information is sufficiently clear, that is, that sufficient information necessary for generating content has been obtained from the user input information.

[0041] If the clarity score is below the threshold, as shown by arrow Q13, a dialogue is generated and presented to the user to obtain information about the user's intent that is missing for content generation, i.e., additional information needed to make the user's intent clearer.

[0042] The dialogue generated at this time is an intention extraction dialogue for extracting (clarifying) the user's intention, and is, for example, a question-type utterance (message) from the game device prompting the user to speak (input) necessary information. Note that the intention extraction dialogue may be presented by voice utterance, by text message, or by both voice and text message.

[0043] Here, the user is presented with the message "Sure thing! Tell me how you enjoyed the fight of BBB?" In other words, a dialogue is presented to the user asking about what was good (fun) about the fight (play) in the game "BBB."

[0044] When such a dialogue is presented, the user inputs additional user input information (hereinafter, also referred to as additional user input information) by making an utterance in response to the presented dialogue. That is, the additional user input information is input in natural language, that is, a user utterance (voice).

[0045] Then, the game device "AAA" calculates the clarity score again based on the initially entered user input information and the additionally entered additional user input information, and determines whether the clarity score is greater than or equal to the threshold value.

[0046] If the clarity score is below the threshold, a new dialogue is generated and presented to clarify the user's intention, as indicated by arrow Q14. That is, new dialogues are generated and presented, and information required for generating content (additional user-input information) is acquired until the clarity score reaches or exceeds the threshold.

[0047] Thereafter, when the clarity score reaches or exceeds the threshold, that is, when sufficient information necessary for generating the content has been obtained (when the user's intention has become clear), the game device "AAA" generates the content, i.e., edits the video of the content material, as shown by arrow Q15.

[0048] For example, a prompt (instruction) is generated based on the initially input user input information and one or more additional user input information inputs, such that video editing that reflects the user's intentions is performed. The generated prompt and content material are then input into a trained model such as a generation AI, video editing is performed, and content is generated. At this time, gameplay context information obtained from the content material, controller input history, etc. is also used as appropriate in generating the content.

[0049] Once the content is generated, it is presented to the user as shown by arrow Q16. The presentation of the content may be by thumbnail display only, by playback of the content, or by some other method, and is not particularly limited.

[0050] In addition, the user can input instructions to modify or regenerate the content as needed, and the game device will save, modify, or regenerate the content in accordance with the user's input.

[0051] <Configuration Example of Information Processing System> FIG. 2 is a diagram showing a configuration example of an embodiment of an information processing system to which the present technology is applied.

[0052] The information processing system 11 shown in FIG. 2 includes an information processing device 21 and a display device 22 .

[0053] The information processing device 21 is formed of a device such as a game console, a smartphone, a tablet, a PC, etc. The information processing device 21 performs various processes related to game play in response to user operations and generates content related to game play in response to user input.

[0054] The display device 22 is connected to the information processing device 21 by wire, and presents to the user game images, sounds, etc. supplied from the information processing device 21. Note that the display device 22 may also be connected to the information processing device 21 wirelessly.

[0055] The information processing device 21 includes an input unit 31, a recording unit 32, a communication unit 33, and a control unit .

[0056] The input unit 31 supplies a signal corresponding to an operation input or a voice input by the user to the control unit 34. The input unit 31 has a controller 41 and a camera 42.

[0057] The controller 41 has, for example, buttons and sticks, and supplies signals to the control unit 34 in response to operations by the user.

[0058] The controller 41 is also provided with an arbitrary sensor 51 such as a pressure sensor, a pulse wave sensor, a sweat sensor, or an acceleration sensor, and a microphone 52 that picks up ambient sounds.

[0059] The sensor 51 measures information such as the pressure (grip pressure) when the user grips the controller 41, the user's heart rate, the amount of sweat (sweating state), etc. The microphone 52 picks up sounds around the information processing device 21, thereby picking up voices uttered by the user (hereinafter also referred to as user voice). For example, data of the user voice input at a predetermined timing is used as the above-mentioned user input information or additional user input information.

[0060] The controller 41 appropriately supplies information such as grip pressure, heart rate, and sweating state measured by the sensor 51 to the control unit 34 as vital information about the user, and supplies the user's voice obtained by the microphone 52 to the control unit 34.

[0061] The sensor 51 and the microphone 52 do not necessarily have to be provided in the controller 41. That is, for example, the sensor 51 and the microphone 52 may be provided separately from the controller 41 as components of the input unit 31.

[0062] The camera 42 captures, for example, the surroundings of the information processing device 21 as a subject, and supplies the resulting image (hereinafter also referred to as external image) to the control unit 34. For example, the external image captured by the camera 42 may be used to extract information about the user, such as the user's facial expression or gaze, or may be used for VR (Virtual Reality) display or AR (Augmented Reality) display of a game screen or the like on the display device 22.

[0063] Note that a part or all of the input unit 31 may be provided outside the information processing device 21. Furthermore, the input unit 31 may be any device that allows a user to input information, such as a keyboard or a touch panel superimposed on the display device 22.

[0064] The recording unit 32 is made up of a memory and records various data such as various programs for games and the like, various trained models used to generate content and dialogue, and data such as templates and background music used to generate content. The recording unit 32 records data supplied from the control unit 34 and supplies the recorded data to the control unit 34.

[0065] The communication unit 33 communicates with external devices as necessary. For example, the communication unit 33 transmits various requests and the like supplied from the control unit 34 to an external server, and receives various data, such as game programs, sent from the server and supplies the data to the control unit 34.

[0066] The control unit 34 controls the overall operation of the information processing device 21. For example, the control unit 34 executes processes related to game play and content generation based on signals (information) supplied from the input unit 31 and various data recorded in the recording unit 32.

[0067] For example, the control unit 34 executes a program to implement a rendering unit 61, a clarity evaluation unit 62, a dialogue generation unit 63, and a content generation unit 64.

[0068] The rendering unit 61 performs rendering processing of the game's video and audio based on a signal corresponding to the user's operation supplied from the controller 41 (hereinafter also referred to as a controller signal) and data related to the game recorded in the recording unit 32.

[0069] The clarity evaluation unit 62 calculates a clarity score based on the user input information and additional user input information supplied from the input unit 31, and evaluates the clarity of the user's intention regarding the generation of content by performing threshold processing on the clarity score.

[0070] The dialogue generation unit 63 generates a dialogue for inputting additional user information according to the result of the evaluation by the clarity evaluation unit 62, i.e., the result of threshold processing on the clarity score. For example, a model such as a large-scale language model (LLM) recorded in the recording unit 32 is used as appropriate to generate the dialogue.

[0071] The content generation unit 64 uses data such as game video and audio (hereinafter also referred to as game play data) obtained by the rendering unit 61 as content material and generates content based on the results of extracting the user's intentions from the user input information.

[0072] For example, content generation utilizes user input information input by the user and content materials, as well as trained models, data such as templates and background music, and gameplay context information, all of which are appropriately recorded in the recording unit 32. As an example, the context information may include at least one of a history of controller operations performed by the user while playing a game (a history of user operations related to content materials), information related to the user's emotions, and information related to the user's vital signs.

[0073] It is assumed here that video data of the video of the content and audio data of the audio of the content are generated as data for presenting the content (hereinafter also referred to as content data).

[0074] The display device 22 is, for example, a display, a head mounted display (HMD), AR glasses, etc. The display device 22 has a display unit 71 and a speaker 72.

[0075] The display unit 71 displays game images based on image data as game play data supplied from the control unit 34, and displays content images based on image data as content data supplied from the control unit 34.

[0076] The speaker 72 outputs game sounds based on audio data as game play data supplied from the control unit 34, and outputs content sounds based on audio data as content data supplied from the control unit 34.

[0077] The display device 22, more specifically the display unit 71 and the speaker 72, may be provided in the information processing device 21. For example, if the display unit 71 and the speaker 72 are provided in the information processing device 21, the information processing device 21 itself will function as the information processing system 11 by itself. That is, the information processing system 11 will be composed only of the information processing device 21. In this way, the information processing system 11 may be composed of only a single device (subject) such as the information processing device 21, or may be configured by combining multiple different devices (subjects) such as the information processing device 21, the display device 22, and a server (not shown).

[0078] <Explanation of Content Generation Processing> The operation of the information processing system 11 will be explained.

[0079] For example, when the user operates the controller 41 to instruct the start of a predetermined game and the commencement of play, the control unit 34 performs processing in accordance with the user's instruction.

[0080] That is, the control unit 34 executes the program recorded in the recording unit 32 in response to the controller signal supplied from the input unit 31 (controller 41), and starts the game.

[0081] When the game starts, the rendering unit 61 sequentially performs rendering processing based on the data related to the game recorded in the recording unit 32, generates video data and audio data of the game, and the control unit 34 supplies the video data and audio data of the game to the display device 22.

[0082] In the display device 22 , the display unit 71 displays the game images based on the image data from the control unit 34 , and the speaker 72 outputs the game sounds based on the audio data from the control unit 34 .

[0083] When the game starts, the user plays the game by appropriately operating the controller 41. When the user performs an operation to play the game, the rendering unit 61 generates video data and audio data of the game that reflects the user's operation, based on the controller signal supplied from the input unit 31 in response to the operation.

[0084] The control unit 34 uses the game play data, consisting of the video data and audio data of the game, generated by the rendering unit 61 from the time the game is started until the time the game is ended as material data that will serve as the material for the content, and supplies this material data (game play data) to the recording unit 32 to record it.

[0085] The control unit 34 also generates context information during game play as appropriate and supplies it to the recording unit 32 for recording.

[0086] Specifically, for example, the control unit 34 generates controller operation information indicating the content of the user's controller operation at each time from the controller signal supplied from the controller 41 at any time during game play.

[0087] For example, the control unit 34 generates vital information of the user based on the user's grip pressure, heart rate, sweating state, etc., which are supplied from the controller 41 as needed during game play. The control unit 34 sets information indicating changes over time in any one or more pieces of information among the user's grip pressure, heart rate, sweating state, etc. during game play as the user's vital information. Note that the vital information may also include information indicating changes over time in other information, such as the user's brain waves.

[0088] Furthermore, for example, the control unit 34 detects the gaze direction and facial expression of the user (game player) who is included as a subject in the external image, based on the external image supplied from the camera 42 at any time during game play. By identifying changes in the user's gaze direction and facial expression over time, the user's concentration level can be identified.

[0089] The control unit 34 generates context information including the thus obtained controller operation information, vital information, information indicating temporal changes in the user's concentration level, etc. Note that the control unit 34 may generate information indicating temporal changes in the user's emotions by performing a recognition process on an external video image or the like, and generate context information including information indicating the temporal changes in the emotions (information about the user's emotions).

[0090] Suppose that after the user has finished playing a game, he or she speaks a predetermined call phrase and commands the generation of content. In response, the controller 41 supplies the control unit 34 with audio data of the user's voice picked up by the microphone 52.

[0091] The control unit 34 performs an analysis process on the voice data supplied from the controller 41, and when it recognizes that the user voice contains a call phrase, it executes a process according to the content of the instruction (command) from the user. For example, if the content of the instruction from the user is to generate content, the content generation process shown in FIG. 3 is executed.

[0092] The content generation process performed by the information processing system 11 will be described below with reference to the flowchart of FIG.

[0093] In step S11, the control unit 34 acquires user input information. For example, the control unit 34 acquires, as the user input information, voice data of a user voice that includes an instruction from the user to generate content.

[0094] For example, in the example shown in FIG. 1 , voice data including the user's utterance "Hey AAA, Create a nice highlight of my game play!" is acquired as the user input information. Note that text information obtained by converting the user's voice into text may also be acquired as the user input information. Furthermore, instead of the user's voice, text information (prompt) entered by the user via a keyboard or the like may also be acquired as the user input information.

[0095] Furthermore, although the input of each piece of user input information, such as the user input information acquired in step S11, the additional user input information described below, and the user input information related to content modification, may be performed in any manner, the following description will be given assuming that the user input information is input using at least the user's speech. For example, the user may input the user input information using only speech, or may input the user input information using a combination of speech and controller operation using the controller 41 or the like.

[0096] In step S12, the clarity evaluation unit 62 performs an analysis process on the user input information (user voice) acquired in step S11, and calculates a clarity score indicating the degree of clarity of the user's intention regarding the content generation instructions indicated by the user input information.

[0097] In step S13, the clarity evaluation unit 62 determines whether the clarity score calculated in step S12 is equal to or greater than a predetermined threshold value. In other words, the clarity of the extraction result of the user's intention is evaluated.

[0098] If it is determined in step S13 that the clarity score is not equal to or greater than the threshold, that is, that the clarity score is less than the threshold, then the process proceeds to step S14.

[0099] In this case, the user input information does not contain enough information necessary for generating the content, that is, the user's intention is evaluated as unclear (abstract). Therefore, in the subsequent steps S14 to S17, processing is performed to acquire additional user input information in order to clarify the user's intention.

[0100] In step S14, the dialogue generation unit 63 identifies the type of additional information required to generate the content based on the result of the analysis process by the clarity evaluation unit 62.

[0101] For example, a plurality of elements are predefined as types of information (hereinafter also referred to as elements) required for generating content. Elements relate to the content, such as the type (category) of content (e.g., highlight footage) and the length of the content. Note that the types of information (elements) required for generating content may differ depending on the type of content.

[0102] The dialogue generation unit 63 identifies, among the multiple elements, elements for which information was not extracted with sufficient accuracy from the user input information, i.e., elements for which the user's intention is evaluated as unclear because the user input information does not contain information with sufficient accuracy.The dialogue generation unit 63 then uses the identified elements as a result of identifying the type of additional information required for generating content.

[0103] In step S15, the dialogue generation unit 63 generates a dialogue for obtaining additional information (additional user input information) required for generating content, more specifically, data for presenting the dialogue (hereinafter also referred to as dialogue data), based on the identification result in step S14. Here, dialogue data for a dialogue that prompts the user to speak (input) the specific content of the element identified in step S14 is generated.

[0104] For example, the dialogue data may be audio data for outputting the dialogue voice, video data for presenting the dialogue in text form, or data consisting of these audio data and video data. In the following, for simplicity of explanation, the dialogue data is assumed to be audio data of the dialogue voice.

[0105] In step S16, the control unit 34 presents the dialogue to the user (game player) by supplying the dialogue data obtained in step S15 to the display device 22. The speaker 72 of the display device 22 presents the dialogue to the user by outputting (playing) the voice of the dialogue based on the dialogue data (voice data) supplied from the control unit 34.

[0106] When the dialogue is presented to the user, the user speaks in response to the content of the dialogue, i.e., the content of the question. In this case, the user speaks content related to the element identified in step S14. When the user speaks, the microphone 52 picks up the user's voice, and the resulting voice data is supplied from the controller 41 to the control unit 34.

[0107] In step S17, the control unit 34 acquires, as additional user input information (additional user input information), voice data of the user's utterance in response to the dialogue presented in step S16. Note that text information obtained by converting the user's voice into text may also be acquired as the additional user input information.

[0108] After the process of step S17 is performed, the process returns to step S12, and the above-described process is repeated.

[0109] In this case, in step S12, the clarity score is calculated using not only the user input information acquired in step S11 but also the additional user input information acquired in step S17. That is, a new clarity score is calculated based on the user input information acquired so far, and the clarity of the user's intention is evaluated.

[0110] Furthermore, in step S14, which is performed from the second time onwards, the type of information (element) required is identified based on the results of the analysis process of the user input information and each additional user input information obtained in step S17 performed so far.

[0111] The above-described processes from step S14 to step S17 are repeated until the clarity score is equal to or greater than the threshold in step S13, i.e., until an evaluation is obtained that the user's intention regarding the generation of the content is clear. Note that if it is determined in step S13 that the clarity score is less than the threshold a predetermined number of times, the process may proceed to step S18, or the threshold used in the threshold process may be changed to a lower value.

[0112] If it is determined in step S13 that the clarity score is above the threshold, it means that sufficient information necessary for generating the content has been obtained, i.e., the user's intention is evaluated as clear (specific), and processing then proceeds to step S18.

[0113] In step S18, the content generation unit 64 generates content (content data) based on the results of the analysis process on the user input information and additional user input information, the content material data (game play data) recorded in the recording unit 32, the trained model, data such as templates and background music, and context information.

[0114] In step S19, the control unit 34 supplies the content data generated in step S18 to the display device 22, thereby presenting the content generation results.

[0115] For example, the display device 22 presents the result of content generation by playing back the video and audio of the content based on the content data supplied from the control unit 34. That is, the display unit 71 displays the video of the content based on the video data as the content data, and the speaker 72 outputs the audio of the content based on the audio data as the content data.

[0116] For example, as a result of generating the content, a thumbnail image of the content may be presented (displayed), and when the thumbnail image is specified and an instruction to play the content is given, the content may be played (presented).

[0117] When the content generation result is presented, the user speaks or operates the controller 41 to input whether or not the generated content should be corrected. Then, voice data and controller signals corresponding to the user's speech or operation are supplied from the controller 41 to the control unit 34. Note that the input regarding whether or not to correct here also includes the user re-inputting the user input information. In other words, the user re-entering the same input as in step S11 is also considered to be an input regarding correction.

[0118] In step S20, the control unit 34 determines whether or not there is any content modification based on the voice data of the user's speech and the controller signal supplied from the controller 41.

[0119] If it is determined in step S20 that there is a correction, the process then proceeds to step S21.

[0120] In step S21, the control unit 34 acquires user input information regarding content modification input by the user (hereinafter, also referred to as "modified user input information"). Note that the above-described utterance regarding whether or not the content is modified may itself be acquired as the user input information regarding content modification.

[0121] When the user instructs to correct the content, the user speaks a request for the correction, i.e., a speech about the content to be corrected. In other words, the user inputs information about the content to be corrected (correction user input information) by speaking.

[0122] For example, the user may use speech to request (demand) the addition of desired information, to request a change to a template, or to request a change to part of the content. Furthermore, if the control unit 34 presents a sample of templates or dramatic effects that can be used to generate the content along with the content generation result, the user may request a modification of the content by appropriately uttering a speech specifying a desired sample.

[0123] When the user makes an utterance regarding content modification, the utterance is picked up by the microphone 52 as user voice, and the resulting voice data is output from the controller 41 to the control unit 34. The control unit 34 acquires the voice data output from the controller 41 as modified user input information. Note that text information obtained by converting the user voice into text may also be acquired as modified user input information.

[0124] Once the corrected user input information is obtained, the process then returns to step S12, and the above-described process is repeated.

[0125] In this case, in step S12, the clarity score is calculated using not only the user input information acquired in step S11 and the additional user input information acquired as needed, but also the corrected user input information acquired in step S21. That is, the clarity score is calculated based on the user input information that has been input so far, and the clarity of the user's intention regarding the content is evaluated again.

[0126] When the evaluation indicates that the user's intention regarding the content is unclear, i.e., when the clarity score is determined to be below a threshold, a dialogue is generated to obtain additional user input information regarding content modifications and is presented to the user.

[0127] Furthermore, when it is determined that the user's intention regarding the content is clear, that is, when it is determined that the clarity score is above a threshold, the subsequent step S18 generates revised content, that is, content that reflects the user's intention to revise.

[0128] When content is modified, elements are added, for example, in accordance with the modified user input information, and the added elements are taken into consideration when calculating the clarity score and generating (modifying) the content.

[0129] If it is determined in step S20 that no corrections have been made, the control unit 34 performs necessary processing, such as supplying the content data of the generated content to the recording unit 32 for recording, or to the communication unit 33 for transmission to an external device. Then, when the necessary processing is completed, the content generation processing ends.

[0130] In this way, the information processing system 11 conducts a dialogue according to the clarity score and generates content. In this way, the user (player) can create content that is close to his or her own intentions simply by answering questions asked by the information processing system 11, i.e., by performing the simple task of responding to the dialogue by speaking. In other words, content that is close to the user's intentions can be obtained more easily.

[0131] <Specific Example of Content Generation> Next, a more specific example of the processing performed in each step of the content generation processing described with reference to FIG. 3 will be described.

[0132] First, a specific example of calculation of the clarity score performed in step S12 of FIG. 3 will be described with reference to FIG.

[0133] In this example, an analysis process is performed on the user input information acquired in step S11 of FIG. 3, that is, the user's utterance.

[0134] For example, the clarity evaluation unit 62 performs analysis processing such as calculating keyword abstraction, calculating class classification probability, and analyzing intent using NLU (Natural Language Understanding), and identifies the user input information, i.e., the content of the user's utterance.

[0135] That is, for example, in the analysis process, important words and phrases contained in the user-input information are extracted, and the probability of the degree of abstraction of the extracted words and phrases is calculated to understand the content of the utterance (user intention).

[0136] In addition, based on the results of the analysis process, the clarity evaluation unit 62 calculates, for each of multiple elements (Video Element), as shown on the right side of the figure, a score indicating the degree of fulfillment of information about the element and the clarity of the user's intention for each element.

[0137] When generating content, the type of information, i.e., the type of element information, that is required may be determined in advance or may be determined for each type of content.

[0138] In this example, the elements include the content type ("Video Type"), caption hint ("Video Caption Hint"), scene extraction hint ("Scene Extraction Hint"), content length ("Video Length Info"), background music selection hint ("BGM Selection Hint"), and template and visual effect hint ("Visual Template / Effect Hint"). Specific examples of the content type ("Video Type") include highlight footage and tutorial footage.

[0139] The fulfillment indicates the degree to which information about the corresponding element has been obtained from the user-input information, in other words, the degree to which the information about the element (user intention) has been clarified. This fulfillment is determined by the clarity evaluation unit 62 based on the results of an analysis process of the user-input information.

[0140] In this example, the degree of satisfaction is expressed in three levels: "Fulfilled," which indicates that sufficient information has been obtained (high satisfaction); "Partially," which indicates that some, but not all, information has been obtained (medium satisfaction); and "None," which indicates that none or almost none of the necessary information has been obtained (low satisfaction).

[0141] In this example, a required level is also defined for each element, indicating the accuracy (degree of clarity) of the information about the element contained in the user input information required for that element, i.e., the degree of clarity of the user's intention regarding the element. The required level of an element can also be said to indicate the importance of that element. For example, the required level is defined in advance for each element.

[0142] Here, the required level is expressed in three stages: "High" indicating high accuracy, "Middle" indicating medium accuracy, and "Low" indicating low accuracy.

[0143] The score of each element indicates whether sufficiently clear information was obtained about the element, in other words, the evaluation result of the degree of clarity of the information about the element (user intention), with a higher score indicating that clearer information was obtained. Conversely, the score of an element indicates the degree of abstraction of the user's intention extracted from the user input information for that element, with a lower score indicating a higher degree of abstraction.

[0144] The score of an element is calculated based on at least one of the fulfillment level and the required level of the element. For example, the score of an element can be determined by a combination of the fulfillment level and the required level of the element. Specifically, for example, a table may be prepared in advance in which combinations of fulfillment level and required level are associated with score values ​​determined for each combination, and the clarity evaluation unit 62 may refer to the table to determine the score for each element.

[0145] For example, it is conceivable that the higher the degree of fulfillment, the larger the score value, and the higher the requirement level, the less likely the score value to increase relative to the degree of fulfillment.

[0146] In the example of FIG. 4, the element "Video Type" has a high requirement level of "High," but a high fulfillment level of "Fulfilled," resulting in a large score of "4." In other words, it can be seen that for the element "Video Type," sufficiently clear (specific) information has been obtained regarding the type (category) of content to be generated. In other words, it can be seen that the user's intention is clear.

[0147] In contrast, the element "Video Caption Hint" has a high requirement level of "High," but a low satisfaction level of "None," resulting in a small score of "0." In other words, it can be seen that despite the high importance of the element "Video Caption Hint," almost no or no information has been obtained that can serve as a hint for generating captions to be displayed in the content. In other words, it can be seen that the user's intention is abstract.

[0148] For example, the clarity evaluation unit 62 calculates a score for each element (evaluates the clarity of the user's intention) and calculates the sum of the scores for each element as the clarity score for the user input information. Therefore, the lower the score for each element and the higher the level of abstraction of the user's intention, the lower (smaller) the clarity score will be.

[0149] Alternatively, for example, an element whose fulfillment level is "None" may always be given a score of "0" regardless of the required level.

[0150] Note that the clarity score may be calculated by any method, such as weighted addition or average of the scores of each element, as long as it is calculated based on the scores of each element. Alternatively, the score of each element itself may be used as the clarity score, threshold processing may be performed for each element score, and a determination (evaluation) of whether to generate a dialogue, i.e., whether the user's intention is clear, may be made based on the results of the threshold processing for each element. Specifically, for example, if there is even one "0" among the scores of each element, it may be considered that the clarity score is determined to be less than the threshold in step S13 of FIG. 3.

[0151] In any case, in this example, the clarity evaluation unit 62 evaluates the degree of clarity of the user's intention for each element related to the content, and based on the evaluation results for each element, it determines the overall evaluation result as to whether the user's intention regarding the content is clear, i.e., whether or not to generate a dialogue.

[0152] As described above, by performing the analysis process and calculating the clarity score, it is possible to determine whether the elements required for content generation, that is, for editing content material (video editing), are satisfied.

[0153] This makes it possible to verify (determine) whether sufficient information has been obtained from the user when generating the content, that is, whether the user's intention regarding the generation of the content is clear, in step S13 of FIG.

[0154] For example, if the verification (determination) results in insufficient information (user intention is unclear), such as insufficient important information being obtained, i.e., if it is determined in step S13 of Fig. 3 that the amount of information is less than the threshold, the process moves to a dialogue for extracting the user's intention. In such a case, a dialogue for clarifying the user's intention is generated and presented to the user.

[0155] By interacting based on the clarity score, the user does not need to think in advance about what information is needed to generate content, but can communicate his or her intentions through interaction with the information processing system 11 and obtain video editing results that are close to his or her expectations. In other words, the user can more easily obtain content that is close to his or her intentions without needing advanced skills in using video editing tools, etc.

[0156] In the example described with reference to Figure 4, the user may be able to set the required level (Required Level) himself, or the user may be able to specify (select) the elements required to generate the content from a plurality of elements (Video Elements) prepared in advance.

[0157] In addition, the user may be allowed to set the weight of the score of each element when calculating the clarity score, or to set a threshold value for threshold processing of the clarity score. In such a case, for example, when the user sets the level of completeness of the content, the clarity evaluation unit 62 may set a threshold value according to the setting result.

[0158] The decision as to whether to engage in a dialogue to obtain additional user input information, i.e., whether to engage in a dialogue to clarify the user's intention, may be based on the evaluation results of the degree of clarity (degree of abstraction) of the user's intention for the user input information, and may be made using other determination methods, not limited to threshold processing on the clarity score.

[0159] As an example, if the clarity score is equal to or greater than a threshold value but the fulfillment level of a specific element, such as an element with a "High" requirement level or an element set by the user as essential, is "Partially" or "None," a dialogue for clarifying the user's intention may be presented. In such a case, step S13 of FIG. 3 determines that the clarity score is less than the threshold value, and then step S15 presents a dialogue that asks a question about the specific element with a fulfillment level of "Partially" or "None." In this example, the dialogues are repeatedly generated and presented until the fulfillment level of the specific element becomes "Fulfilled."

[0160] Furthermore, for example, if the clarity score is equal to or greater than a threshold, but the number of dialogues presented so far is less than a predetermined number, and there is an element for which the fulfillment level is "Partially" or "None," a dialogue for clarifying the user's intention may be presented. In such a case, it may be possible to present a dialogue that asks a question about an element for which the fulfillment level is "Partially" or "None," or an element for which the requirement level is high and the score value is low.

[0161] Next, an example of the processing performed in steps S14 and S15 in FIG. 3, that is, the generation of a dialogue, will be described with reference to FIG.

[0162] For example, the dialogue generation unit 63 generates a dialogue that prompts input (utterance) of information about elements whose scores for each element are below a predetermined value, that is, elements whose evaluation of the clarity of the user's intention is below a predetermined evaluation.

[0163] More specifically, for example, it is assumed that an analysis result such as that shown by arrow Q31 in FIG. 5 is obtained as a result of the analysis process on the user input information.

[0164] In this example, the scores for the elements "Video Caption Hint" and "Scene Extraction Hint" in particular are "0," indicating that the information for these elements is insufficient. In other words, it is clear that the evaluation results indicate that the user's intention is not clear for these elements. Therefore, for example, in step S14 of FIG. 3, the elements "Video Caption Hint" and "Scene Extraction Hint" are obtained as identification results as types of information necessary for generating content.

[0165] In this case, for example, the dialogue generation unit 63 generates a prompt (command) to generate a dialogue that asks a question that will prompt the user to speak about the element "Video Caption Hint" and the element "Scene Extraction Hint" that have low scores (scores below the predetermined value "0"), as shown by arrow Q32.

[0166] The dialogue generation unit 63 generates a dialogue by inputting the generated prompt into a model such as a large-scale language model, thereby generating a conversational question that elicits (asks) information from the user about the element "Video Caption Hint" and the element "Scene Extraction Hint" that are necessary for generating content, i.e., an intention extraction dialogue.

[0167] Therefore, in subsequent step S16 in Figure 3, the generated dialogue is presented to the user, as indicated by arrow Q33. In this example, the user plays a fighting game, and the dialogue asks the user what was most memorable about the game.

[0168] When the dialogue is presented to the user, the user then makes an utterance in response (answer) to the presented dialogue, as indicated by arrow Q34, and in step S17 of FIG. 3, this utterance is acquired as additional user input information.

[0169] Then, the clarity evaluation unit 62 performs analysis processing and calculation of clarity scores again for the additional user input information thus obtained, as well as for the already acquired user input information and additional user input information, and performs the processing of step S13 in Fig. 3 based on the obtained clarity scores. In this way, by calculating the clarity scores again using the already acquired user input information (scoring the degree of clarity), it is possible to reflect the user's intention in the content without missing anything.

[0170] When calculating the clarity score using newly acquired additional user input information, information obtained from the most recently acquired additional user input information may be prioritized (weighted) over additional user input information already acquired before the acquisition of the additional user input information or information obtained from user input information. In other words, when weighting each piece of information to be used in calculating the clarity score, the weight of the information may be lowered the earlier the information was acquired.

[0171] 5, a dialogue is generated based on the scored clarity of each element, so the user can input the information required for content generation (video editing) simply by responding to the dialogue. In particular, in this case, the user does not need to be particularly aware of what elements are required for content generation, and can easily generate content simply by responding to the dialogue in natural language.

[0172] If it is determined that the clarity score is equal to or greater than the threshold, the process of step S18 in FIG. 3 is then performed, and content is generated.

[0173] Before describing a specific example of content generation, the use of context information when generating content will be described with reference to Fig. 6. In particular, scene extraction based on the generation of controller operation information as context information will be described with reference to Fig. 6.

[0174] For example, the content generation unit 64 acquires (generates) activity level information shown on the left side of Fig. 6 from controller operation information, which is a controller operation history during game play by the user. Note that in the left part of the figure, the horizontal direction indicates time (hours), and the vertical direction indicates whether or not a controller operation has occurred.

[0175] In this example, a rectangular wave (square wave) graph indicating whether or not the user is operating the controller 41 at each time point is generated as activity level information.

[0176] In particular, in the graph shown on the left side of the figure, the upward convex portions indicate periods during which the user operated the controller, i.e., active periods, and the downward convex portions indicate periods during which the user did not operate the controller, i.e., inactive periods. Note that, although an example will be described here in which the activity level indicated by the activity level information is expressed as two values, i.e., the presence or absence of an operation, the activity level may also be expressed as multiple values ​​(multiple stages) or as a numerical value such as a score.

[0177] The content generation unit 64 extracts (specifies) scenes at each time based on the activity level information and the known flow of the game, i.e., the general pattern of scene changes during game play, and generates scene extraction information indicating the extraction results. The scene extraction information includes information indicating the scene type at each time and information indicating the time of a turning point (key point) at which the scene changes.

[0178] Here, the portion indicated by arrow Q51 contains scene extraction information, that is, the scene extraction results (scene type) at each time.

[0179] This example shows a case where a user plays a fighting game. For example, period T11 is a break period until the second round of the fight, and period T12 is a scene from the second round of the fight. For example, when focusing on period T11, the period immediately before period T11 is active for a relatively long time, and the period T11 thereafter is inactive, so period T11 is identified as a break period until the second round.

[0180] By generating activity level information from such controller operation information and performing scene extraction, scene extraction information indicating the scene at each time of game play can be obtained, and this scene extraction information can be used to generate content.

[0181] It should be noted that scene extraction information can also be considered a type of context information. Furthermore, scene extraction information may be generated when controller operation information is acquired, and context information including the scene extraction information may be recorded in the recording unit 32.

[0182] When generating content, scene extraction information and controller operation information can be used in combination to obtain better, more dynamic content.

[0183] For example, while scene extraction information alone can identify scene types and scene changes, it cannot detect the timing of player excitement. In contrast, using scene extraction information and controller operation information makes it possible to identify scenes that the player is excited about. This makes it possible to generate content with a good balance, for example, by appropriately arranging scenes with few controller operations and scenes with many.

[0184] In addition, vital information and content materials may also be used for scene extraction. For example, scenes that excite the user can be identified using vital information such as grip pressure and heart rate, or the user's level of concentration obtained from changes in gaze direction and facial expression over time. Scenes can also be identified by performing video analysis on game footage as content materials.

[0185] A specific example of content generation will be described with reference to Fig. 7. Content generation uses the results of the analysis process on one or more pieces of user input information obtained in step S12 of Fig. 3, context information, content materials, data such as templates and background music, and the scene extraction information described with reference to Fig. 6.

[0186] 7, the left side of the drawing shows the results of the analysis process for one or more pieces of user input information. The user input information here may include, in addition to the information acquired in step S11 of FIG. 3, additional user input information acquired in step S17 and user input information related to corrections acquired in step S21, as appropriate.

[0187] In the example of FIG. 7, information on each element (Video Element) extracted or generated from user-input information during analysis processing is shown as a value.

[0188] The value may be extracted or generated during the analysis process, or may be extracted or generated based on the final result of the analysis process. For example, the value may be generated (estimated) based on the user's intention regarding the element extracted from the user input information.

[0189] 7, for example, the value of the element "Video Type" is "Highlight," indicating that the type of content is a highlight video. Also, for example, the value of the element "Video Caption Hint" is "Finally beat it," indicating that the phrase "Finally beat it" has been obtained as a hint for the caption to be embedded in the content.

[0190] The content generation unit 64 generates content based on the value of each element, context information, content material, data such as templates and background music, and scene extraction information.

[0191] As an example, the content generation unit 64 constructs or generates elements necessary for generating content based on each piece of information such as these values, and generates content by combining the obtained components.

[0192] Specifically, the similarity between the scene extraction results based on the scene extraction information, i.e., each extracted scene, and the value "KO, Special Move" of the "Scene Extraction Hint" element related to scene extraction hints is calculated, and scenes with high similarity are selected as scenes to be used in generating content. In this example, for example, KO scenes or scenes where special moves are executed are selected.

[0193] Furthermore, for example, the similarity between the type of each template indicated by the metadata of each of the multiple templates recorded in the recording unit 32 and the value "Solid" of the element "Visual Template / Effect Hint" related to the template is calculated, and a template type with a high degree of similarity is selected as the template to be used for generating content.

[0194] Similarly, for example, based on the value "High Tempo" of the element "BGM Selection Hint" which relates to hints for selecting BGM, BGM to be used in generating content is selected from multiple BGMs recorded in the recording unit 32, or BGM to be used in generating content is generated.

[0195] Furthermore, based on the value "Finally beat it" of the element "Video Caption Hint" relating to the caption hint, for example, a title of the content and one or more captions to be embedded in the content are generated.

[0196] The content generation unit 64 combines the components such as captions, scenes, background music, templates, and production effects selected and generated in this manner to generate highlight footage as content, which is indicated by the value "Highlight" of the content type element "Video Type."

[0197] In this case, for example, the selection and generation of each component may be performed using different models, and prompts generated based on the resulting components may be input to a model such as a generation AI to generate a single piece of content. Alternatively, for example, prompts generated from the values ​​of each element may be input to a single model, and the model may select and generate each component and generate content by combining the components. Furthermore, content may be generated by any method, such as a method that does not use a model such as a generation AI.

[0198] Once the content has been generated, the generated content is presented to the user in step S19 of Fig. 3. The user inputs whether or not the content needs to be modified as appropriate.

[0199] In this case, for example, as shown by arrow Q71 in FIG. 8, multiple patterns of content may be presented to the user.

[0200] In this example, in step S18 of FIG. 3, different templates are used to generate a plurality of patterns of content, and thumbnail images of these contents are presented to the user.

[0201] For example, the user may select a desired one of the thumbnails presented and instruct the display device 22 to play back the content corresponding to the selected thumbnail. In response to a signal from the controller 41 corresponding to the user's operation, the control unit 34 then supplies the data of the content designated by the user (content data) to the display device 22, causing the display device 22 to play back (present) the content.

[0202] When the user plays one or more pieces of content, he or she can select (designate) the desired piece of content, as indicated by arrow Q72, and input whether or not the content needs to be revised. In this case, the user can use at least one of the following methods to input whether or not the content needs to be revised: operation of controller 41 or voice input to microphone 52.

[0203] For example, if the user inputs that no modifications are to be made, it is determined in step S20 of FIG. 3 that no modifications are to be made, and the generated content is recorded or output as is.

[0204] In contrast, when a user inputs information indicating that corrections are required, the user inputs information including information regarding the corrections to be made, such as the parts to be corrected and the type of corrections to be made, in addition to the information indicating that corrections are required.

[0205] Specifically, for example, the user may make utterances including requests for modifications, such as making the template cooler, making the scene in which the special move is executed longer, etc. Then, in step S21 of Fig. 3, such utterances by the user are acquired by the microphone 52 as modified user input information.

[0206] Then, in the subsequent step S12, as indicated by arrow Q73, the newly acquired corrected user input information is taken into account in addition to the previously acquired user input information, and the clarity score is calculated again. That is, the newly entered corrected user input information is also used in calculating the clarity score, so that the content generation intention that reflects the user's desire for revisions is extracted. As a result, the user's desire for revisions (revision intentions) are reflected in the newly generated content, and content that is even closer to the user's intentions can be obtained.

[0207] In addition to presenting the user with multiple patterns of content, the user may be presented with multiple patterns of templates and sample effects, and allowed to select the template or effect of his or her choice.

[0208] In such a case, the control unit 34 causes the display unit 71 to display, for example, a screen (UI (User Interface)) shown in FIG.

[0209] In the example of FIG. 9, a column for displaying thumbnail images of content is provided on the left side of the drawing, and a column for displaying games, content, etc. is provided on the right side of the drawing.

[0210] In particular, the area indicated by arrow Q81 displays thumbnail images of content generated according to the extracted user's intentions, particularly thumbnail images of content with effect selected based on the user's intentions, as the content generation results. Furthermore, for example, the area indicated by arrow Q82 displays multiple sample images (thumbnail images) of effect effects other than the effect selected based on the user's intentions. Note that multiple template samples may also be displayed in the area indicated by arrow Q82.

[0211] When a user wants to try out a different type of content, they select the desired one from the samples displayed in the area indicated by arrow Q82. Then, for example, content corresponding to the sample selected by the user is generated and played (displayed) on the right side of the display screen.

[0212] In this case, in step S21 of FIG. 3, information relating to the dramatic effect corresponding to the sample designated by the user is acquired by the control unit 34 as modified user input information relating to the modification of the content.

[0213] The user can display different patterns of content as needed, designate content for saving, or instruct further modifications to the desired content. In this way, the operation of selecting a preferred sample from among multiple presented samples of dramatic effects, etc., can also be considered to be an operation of the user instructing modification of content. In other words, the operation of designating a sample can also be considered to be an input of modified user input information. This input operation can be a controller operation or a voice input operation using the user's speech.

[0214] Note that differences in content patterns are not limited to dramatic effects, but can also be anything, such as different templates or captions. Furthermore, the area indicated by arrow Q82 may display thumbnail images of content generated with different patterns, rather than samples of dramatic effects or templates different from those indicated by arrow Q81.

[0215] When generating content, the user input information, additional user input information, and modified user input information may be input by voice, by operating a controller, or by a combination of voice and controller. Alternatively, multimodal input may be performed by combining different modalities such as voice, controller operation, and gestures.

[0216] 10 shows an example in which user input information is input by combining voice input and controller operation input. In this example, the display unit 71 displays the display screen shown on the left side of the figure.

[0217] On this display screen, the area indicated by arrow Q91 shows footage (scenes) of the game played by the user at a specific time, and the area indicated by arrow Q92 shows a timeline with thumbnail images of each scene in the game played by the user. Furthermore, the area indicated by arrow Q93 shows information entered by the user, i.e., the content of the user's utterances.

[0218] 10, the user operates the controller 41 to align (position) the pointer displayed on the display screen with a desired position on the timeline, thereby specifying the scene at that position. For example, the number "1" displayed on the timeline represents the earliest scene specified by the user (hereinafter also referred to as the first scene), or more specifically, the earliest scene for which information has been input by the user.

[0219] When the user selects the first scene, the image of that first scene is displayed in the area indicated by arrow Q91. That is, the control unit 34 supplies the game play data, which has been designated as content material, to the display device 22, thereby displaying the first scene.

[0220] While watching the first scene, the user may utter, for example, their impressions of playing that scene, such as "I made a comeback here," or an explanation of the scene, such as "An ultra-rare item," or "My first match against a level 4 CPU." This utterance is then acquired by the microphone 52 as user input information and supplied to the control unit 34 (clarity evaluation unit 62).

[0221] In this case, the user has input information about a desired scene by combining controller operation and voice input.

[0222] When information about the first scene (user-input information) is entered, that information is displayed in text in the area indicated by arrow Q93. This allows the user to check what they have entered and to correct the information they have entered as necessary.

[0223] In PC and mobile environments, pointer and touch operations can be utilized, and by utilizing the advantages of such operations in combination with natural language (speech), it is possible to realize the input operation described with reference to FIG. 10 . In particular, in this example, the user can input information associated with a scene. The information input in this manner can be used as important key information during automatic video editing, enabling scene extraction and video editing that closely matches the user's expectations (intentions).

[0224] Furthermore, for example, the control unit 34 obtains controller operation information indicating the details of the user's controller operation as context information, and therefore information regarding the controller operation may also be displayed (added) in the content. In other words, the details of the user's operation regarding gameplay may be displayed in the video of the content.

[0225] In such a case, as shown in Figure 11, in the video of a certain playback time of the content, the part indicated by arrow Q101 displays a video of one scene during game play and a caption. Also, in the part indicated by arrow Q102 in the content video, an image showing the details of the user's controller operation in the scene displayed in the part indicated by arrow Q101 is displayed. In this way, for example, when a user watches the content, they can understand how to perform techniques and the timing of operations in each scene.

[0226] This method of displaying simultaneously a game scene and the controller operation in that scene is particularly useful when creating tutorial videos or the like as content.

[0227] The information processing device 21 generates a dialogue prompting the user to input additional user input information according to the clarity of the user's intention extracted from the user input information, in other words, the degree of abstraction of the user's intention. In this case, the user may be able to set in advance the required clarity. In other words, the required level of each element may be determined according to a setting operation by the user.

[0228] In such a case, the required degree of clarity (hereinafter also referred to as required clarity) can be set, for example, as shown in Fig. 12. In Fig. 12, a setting example in a PC environment is shown on the left side, and a setting example in a mobile environment is shown on the right side.

[0229] For example, in the example shown on the left side of Figure 12, a slider bar indicated by arrow Q121 is displayed on a setting screen, and the user sets the degree of clarity of the request by operating the knob (slider) on the slider bar. In this example, the further to the right the knob is moved in the figure, the higher the degree of clarity of the request, i.e., the more specific information is required.

[0230] Based on the required clarity set by the user in this manner, the control unit 34 (clarity evaluation unit 62) can determine the required level (Required Level) of each element (Video Element) shown in Figure 4 and determine the threshold value for the clarity score.

[0231] In the example shown on the right side of the figure, as indicated by arrow Q122, multiple levels of requirement clarity are displayed on the settings screen along with radio buttons, allowing the user to select the desired level of requirement clarity.

[0232] In this example, the degree of request clarity can be set in five stages: "Very High," "High," "Moderate," "Low," and "Very Low."

[0233] Alternatively, as shown in FIG. 13, the required clarity may be set for each element (Video Element).

[0234] In the example of Figure 13, slider bars are displayed for setting the required clarity for each of the four elements: "Video Type," "Caption Hint," "Screen Extraction Hint," and "Video Length Hint." For example, the portion indicated by arrow Q131 displays a slider bar for setting the required clarity for "Video Type."

[0235] In this example, based on the required clarity set for each element, the required level (Required Level) for each element (Video Element) can be determined, and the weight of the score for each element to be used when calculating the clarity score can be determined.

[0236] In addition, the clarity evaluation unit 62 may learn and adjust the required level (Required Level) of each element (Video Element) based on, for example, the history of user modifications to the presented content, i.e., modification user input information.

[0237] In such a case, it is conceivable to adjust the required level (Required Level) for elements that are frequently modified, so that it increases (becomes higher), and to maintain or decrease (become lower) the required level for elements that are rarely modified. In this way, the required level of points (elements) that are highly requested by the user can be accurately adjusted, and video editing can be completed at low cost. In other words, content that meets the user's requirements can be obtained more quickly.

[0238] As a specific example, assume that the user makes an utterance regarding the modification of content, as shown on the left side of FIG.

[0239] In this example, the user has requested modifications related to the template and visual effects, including modifications related to color and appearance. Therefore, as shown on the right side of the figure, the required level of the element "Visual Template / Effect Hint" is changed to a higher level, "High."

[0240] As shown on the left side of the figure, the user has requested modifications related to the BGM, including a modification related to tempo and a modification related to the selection of music (BGM). Therefore, for example, as shown on the right side of the figure, the required level (Required Level) of the element "BGM Selection Hint" is changed to a higher level, "High."

[0241] Furthermore, for example, since no corrections to the element "Video Caption Hint" have been requested by the user in the past, the required level of the element "Video Caption Hint" is changed from the highest "High" to the next highest "Middle."

[0242] In this way, the clarity evaluation unit 62 changes the required level (Required Level), which is the standard for evaluating the clarity of user-input information, based on the user's past revision history, i.e., based on previously acquired revised user-input information, thereby reducing the number of interactions with the user regarding revisions. This reduces the overall number of interactions during content generation, making it possible to generate content that meets user requirements more quickly.

[0243] As described above, the environment of the information processing system 11 may be a headset environment in which the display device 22 is configured from an HMD or the like, or the content that serves as content material or the content generated by the content generation process of Figure 3 may be VR content.

[0244] Furthermore, the content generated by the information processing device 21 is not limited to achievement videos such as highlight videos, but may be any other content.

[0245] For example, the content may be a video that is used to explain a problem that a player (user) is having when playing poorly or when they are unable to complete a game, etc. Furthermore, the content may be, for example, a promotional highlight video that is generated after live streaming of gameplay, etc.

[0246] Furthermore, the content generated by the information processing device 21 may be something other than a game. Specifically, for example, a highlight video of a drive may be generated using video recording information obtained during autonomous driving of a vehicle, i.e., video obtained by filming, as content material. Furthermore, content may be generated using still images or moving images (video) captured by a smartphone or the like as content material.

[0247] Alternatively, for example, in presenting the dialogue in step S16 of FIG. 3, a character such as an avatar concierge may be displayed on the display unit 71, and the dialogue may be presented to the user by the character speaking.

[0248] Furthermore, in the information processing system 11, a system that assumes text input, such as a chatbot, may be used to input user input information and present a dialogue to the user. Furthermore, in addition to voice (utterance) and text, the user may input user input information and the like using any one or more input modalities, such as an input device such as a mouse, gaze, pointing, gesture, etc.

[0249] <Other Configuration Examples of Information Processing System> The information processing system 11 is not limited to the configuration shown in Fig. 2, and may be configured as shown in Fig. 15, for example. Note that in Fig. 15, parts corresponding to those in Fig. 2 are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate.

[0250] The information processing system 11 shown in FIG. 15 includes an information processing device 21 , a display device 22 , and a server 201 .

[0251] In this example, the information processing device 21 has the input unit 31, the recording unit 32, the communication unit 33, and the control unit 34, similar to the case in Fig. 2, but the control unit 34 does not function as the clarity evaluation unit 62, the dialogue generation unit 63, or the content generation unit 64, unlike the case in Fig. 2. Note that, although the rendering unit 61 is on the information processing device 21 side in the example in Fig. 15, this is not limiting, and the rendering unit 61 may be provided on the server 201 side.

[0252] The server 201 acquires user input information and the like from the information processing device 21, generates content, and supplies the generated content to the information processing device 21. The server 201 may be configured by one information processing device or by multiple information processing devices. In other words, the server 201 may have a cloud configuration.

[0253] The server 201 includes a communication unit 211 , a recording unit 212 , and a control unit 213 .

[0254] The communication unit 211 communicates with the information processing device 21 via the network. For example, the communication unit 211 transmits various pieces of information supplied from the control unit 213 to the information processing device 21, receives various requests and information transmitted from the information processing device 21, and supplies the same to the control unit 213.

[0255] The recording unit 212 is made up of a memory and records various data such as various trained models used to generate content and dialogue, templates used to generate content, background music, etc. The recording unit 212 records data supplied from the control unit 213 and supplies the recorded data to the control unit 213.

[0256] The control unit 213 controls the overall operation of the server 201. For example, the control unit 213 executes processing related to content generation based on information supplied from the communication unit 211 and various data recorded in the recording unit 212. In such a case, the control unit 213 executes programs to realize the clarity evaluation unit 62, the dialogue generation unit 63, and the content generation unit 64.

[0257] When the information processing system 11 has the configuration shown in FIG. 15 , the user input information acquired by the control unit 34 of the information processing device 21 is transmitted to the server 201 by the communication unit 33 .

[0258] In the server 201, the communication unit 211 receives user input information from the information processing device 21 and supplies it to the control unit 213, and the control unit 213 performs processing to generate content and dialogue based on the user input information supplied from the communication unit 211.

[0259] That is, for example, the clarity evaluation unit 62 of the control unit 213 performs the same processing as steps S12 and S13 in FIG.

[0260] Furthermore, if it is determined that the clarity score is less than the threshold value, the dialogue generation unit 63 of the server 201 performs processing similar to steps S14 and S15 of Figure 3, and the generated dialogue, or more specifically, data for presenting the dialogue, is transmitted to the information processing device 21 by the communication unit 211.

[0261] Then, in the information processing device 21, the communication unit 33 receives the dialogue data and supplies it to the control unit 34, which then performs the same process as in step S16 of Fig. 3 to present the dialogue to the user. After that, the same process as in step S17 of Fig. 3 is performed to acquire additional user input information, which is then transmitted by the communication unit 33 to the server 201. Then, in the server 201, the communication unit 211 receives the additional user input information, and the server 201 again calculates the clarity score and performs threshold processing, taking the additional user input information into consideration.

[0262] On the other hand, if it is determined that the clarity score is equal to or greater than the threshold value, the content generation unit 64 in the server 201 performs processing similar to step S18 in FIG. 3 to generate content, and the content (content data) is transmitted to the information processing device 21 by the communication unit 211. The content is received by the communication unit 33.

[0263] Then, the control unit 34 performs the same processes as steps S19 and S20 in FIG. 3, and the content received from the server 201 is presented to the user, and the user inputs whether or not to modify the content.

[0264] If the content needs to be corrected, the corrected user input information acquired by the control unit 34 is transmitted to the server 201 by the communication unit 33, and the server 201 receives the corrected user input information transmitted from the information processing device 21 by the communication unit 211. Then, the server 201 and the information processing device 21 then perform the processes from step S12 onwards in Fig. 3. That is, a necessary dialogue is generated as appropriate and presented to the user, or the content is corrected and presented to the user.

[0265] As described above, in the information processing system 11 configured as shown in FIG. 15, content closer to the user's intention can be obtained more easily by processing similar to the content generation processing described with reference to FIG.

[0266] Second Embodiment Description of Content Generation Processing The above describes an example in which the clarity of the user's intention is evaluated for the user input information initially input by the user, a dialogue is conducted based on the evaluation results, and content is generated based on the user input information and additional user input information.

[0267] However, the present invention is not limited to this. First, content may be generated from content material and context information without user input information, and then, if necessary, the content may be modified based on modified user input information entered by the user.

[0268] As a specific example, a highlight video may be generated for a quick replay or the like at an appropriate timing, such as after a battle has ended during gameplay, based only on the content material and context information, without user input, i.e., without user utterances. In this case, for example, the user may be able to give instructions to modify the presented highlight video (content).

[0269] First, when content is generated based only on content material and context information, for example, the information processing system 11 shown in FIG. 2 performs the content generation process shown in FIG.

[0270] The content generation process performed by the information processing system 11 will be described below with reference to the flowchart of FIG.

[0271] It is assumed that, at the time the content generation process is started, the recording unit 32 has recorded therein game play data that will be used as content material, and context information that includes controller operation information and vital information.

[0272] In step S101, the content generation unit 64 generates content (content data) based on the content material data (game play data) recorded in the recording unit 32, learned models, data such as templates and background music, and context information.

[0273] For example, the content generation unit 64 extracts scenes based on the content material and context information, and selects scenes to be used in generating the content, selects templates, background music, and production effects, and generates captions based on the results of the scene extraction, the context information, the game genre, etc. The content generation unit 64 combines the selected and generated components such as captions, scenes, background music, templates, and production effects to generate a highlight video of the content material as content. Note that in step S101, the content may be generated using a model such as a generation AI, as appropriate, or the content may be generated without using a model.

[0274] In step S102, the control unit 34 presents the content generation result by supplying the content data generated in step S101 to the display device 22. For example, in step S102, the same process as step S19 in FIG.

[0275] When the content generation result is presented, the user operates the controller 41, for example, to input whether or not the generated content should be corrected. Then, a controller signal corresponding to the user's operation is supplied from the controller 41 to the control unit 34. Note that the input of whether or not to correct the content may be input by speech (voice).

[0276] In step S103, the control unit 34 determines, based on the controller signal supplied from the controller 41, whether or not the content has been modified.

[0277] If it is determined in step S103 that there is a correction, the process then proceeds to step S104.

[0278] In step S104, the control unit 34 acquires user input information related to content modification (modification user input information). In step S104, the same process as in step S21 in Fig. 3 is performed, and the modification user input information related to content modification input by the user through speech or the like is acquired.

[0279] After the process of step S104 is performed, the processes of steps S105 to S112 are performed, but these processes are the same as the processes of steps S12 to S19 in Fig. 3, so a description thereof will be omitted. After the content generation result is presented in step S112, the process returns to step S103.

[0280] In step S105, which is performed for the first time, the modified user input information acquired in step S104 is analyzed, and a clarity score is calculated based on the analysis results. In this case, the clarity score is calculated taking into consideration whether or not each component element, such as captions, scenes, background music, templates, and dramatic effects, selected and generated during content generation in step S101, has been modified. In other words, in step S105, the clarity of the user's intention regarding the content modification is evaluated for each element, and a clarity score is calculated based on the evaluation results.

[0281] Furthermore, in step S105, which is performed from the second time onwards, if there is additional user input information obtained in step S110 or modified user input information obtained in step S104 from the second time onwards, the clarity score is calculated based on the additional user input information and modified user input information and the modified user input information obtained in step S104 for the first time.

[0282] In step S106, a threshold process based on the clarity score is used to evaluate the clarity of the user's intention regarding the content modification. In step S107, a process similar to step S14 in FIG. 3 is performed to identify the type of information required to generate content that reflects the user's intention to modify the content, i.e., the type of information required to modify the content. Furthermore, in step S111, content that reflects the user's intention to modify is generated. In other words, the content is modified.

[0283] If it is determined in step S103 that no corrections have been made, the control unit 34 performs necessary processing, such as supplying the content data of the generated content to the recording unit 32 for recording, or to the communication unit 33 for transmission to an external device. Then, when the necessary processing is completed, the content generation processing ends.

[0284] In this way, the information processing system 11 generates content based only on the content material and context information, and if there are instructions to modify the content, it generates and presents dialogue as appropriate to clarify the user's intentions, and generates final content that reflects the instructions to modify. In this way, the user can create content that is close to their own intentions simply by answering questions asked by the information processing system 11, i.e., by responding verbally to the dialogue. In other words, content that is close to the user's intentions can be obtained more easily.

[0285] Note that even when the information processing system 11 has the configuration shown in Fig. 15, content can be obtained by performing processing similar to the content generation processing described with reference to Fig. 16 using the information processing system 11. In such a case, the communication unit 211 of the server 201 and the communication unit 33 of the information processing device 21 communicate as appropriate to exchange (transmit / receive) various information such as modified user-input information, additional user-input information, and content data.

[0286] <Example of Computer Configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the computer includes a computer built into dedicated hardware, and a general-purpose personal computer, for example, that can execute various functions by installing various programs.

[0287] FIG. 17 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0288] In the computer, a CPU (Central Processing Unit) 501 , a ROM (Read Only Memory) 502 , and a RAM (Random Access Memory) 503 are interconnected by a bus 504 .

[0289] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.

[0290] The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, etc. The output unit 507 includes a display, a speaker, etc. The recording unit 508 includes a hard disk, a non-volatile memory, etc. The communication unit 509 includes a network interface, etc. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0291] In a computer configured as described above, the CPU 501 loads, for example, a program recorded in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504, and executes the program, thereby performing the above-described series of processes.

[0292] The program executed by the computer (CPU 501) can be provided by being recorded on a removable recording medium 511 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0293] In a computer, a program can be installed in the recording unit 508 via the input / output interface 505 by inserting a removable recording medium 511 into the drive 510. The program can also be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Alternatively, the program can be installed in the ROM 502 or the recording unit 508 in advance.

[0294] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0295] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.

[0296] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.

[0297] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.

[0298] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0299] Furthermore, the present technology can also be configured as follows.

[0300] (1) An information processing system comprising: an evaluation unit that evaluates a degree of clarity of a user's intention regarding content generation, as indicated by user input information; a dialogue generation unit that generates a dialogue to obtain additional user input information when the evaluation indicates that the user's intention is unclear; and a content generation unit that generates the content based on the user input information when the evaluation indicates that the user's intention is clear. (2) The information processing system described in (1), in which at least an utterance by the user is used to input the user input information. (3) The information processing system described in (1) or (2), in which, when the dialogue is presented to the user and additional user input information is input, the evaluation unit evaluates a degree of clarity of the user's intention based on the user input information that has been input so far, and repeatedly generates the dialogue and presents it to the user until the evaluation indicates that the user's intention is clear. (4) The information processing system according to any one of (1) to (3), wherein the evaluation unit calculates a score indicating a degree of clarity of the user's intention based on the user input information, and evaluates the degree of clarity of the user's intention based on the score. (5) When the user inputs the user input information regarding a modification of the content, the evaluation unit evaluates a degree of clarity of the user's intention based on the user input information regarding the modification of the content and the user input information input so far, the dialogue generation unit generates the dialogue when the evaluation indicates that the user's intention is not clear, and the content generation unit generates the content reflecting the user's intention to modify when the evaluation indicates that the user's intention is clear. (6) The information processing system according to any one of (1) to (5), wherein the content generation unit generates the content based on the user input information, material data that will be the material for the content, and context information related to the material for the content.(7) The information processing system according to (6), wherein the context information includes at least any one of the user's operation history regarding material of the content, information regarding the user's emotions, information regarding the user's concentration level, and vital information of the user. (8) The information processing system according to any one of (1) to (7), wherein the evaluation unit evaluates the clarity of the user's intention for each element related to the content and determines an evaluation result as to whether the user's intention is clear overall based on the evaluation result for each element. (9) The information processing system according to (8), wherein the evaluation unit evaluates the clarity of the user's intention for each element based on at least any one of a required level of clarity of information regarding the element and a sufficiency of information regarding the element included in the user input information. (10) The information processing system according to (9), wherein the evaluation unit adjusts the required level of each element based on a history of modifications to the content. (11) The information processing system according to (9), wherein the evaluation unit determines the required level of each of the elements in accordance with a setting operation by the user. (12) The information processing system according to any one of (8) to (11), wherein the dialogue generation unit generates the dialogue prompting input of information about the elements for which the evaluation of the clarity of the user's intention is equal to or lower than a predetermined evaluation. (13) The information processing system according to any one of (1) to (12), wherein the content material includes at least video data of game play, and the video of the content displays the user's operation details related to the game play. (14) The information processing system according to (5), further comprising a control unit that presents the content generation result and a plurality of samples of at least one of templates and dramatic effects that can be used to generate the content, and acquires information about at least one of templates and dramatic effects corresponding to the samples specified by the user as the user input information related to modifying the content. (15) The information processing system according to any one of (1) to (14), wherein the dialogue generation unit generates the dialogue using a large-scale language model.(16) An information processing method, comprising: an information processing system evaluating a degree of clarity of a user's intention regarding content generation, as indicated by user input information; generating a dialogue to obtain additional user input information if the evaluation indicates that the user's intention is unclear; and generating the content based on the user input information if the evaluation indicates that the user's intention is clear. (17) A program that causes a computer to execute processes, including: evaluating a degree of clarity of a user's intention regarding content generation, as indicated by user input information; generating a dialogue to obtain additional user input information if the evaluation indicates that the user's intention is unclear; and generating the content based on the user input information if the evaluation indicates that the user's intention is clear. (18) An information processing system comprising: a content generation unit that generates the content based on material data that is a material for the content; an evaluation unit that, when user input information regarding modification of the content is input by the user, evaluates a degree of clarity of the user's intention regarding modification of the content as indicated by the user input information; and a dialogue generation unit that, when the evaluation indicates that the user's intention is not clear, generates a dialogue to obtain additional user input information, wherein the content generation unit, when the evaluation indicates that the user's intention is clear, generates the content that reflects the user's intention to modify, based on the material data and the user input information. (19) The information processing system described in (18), wherein at least utterances by the user are used to input the user input information. (20) The information processing system described in (18) or (19), wherein, when the dialogue is presented to the user and additional user input information is input, the evaluation unit evaluates the clarity of the user's intention based on the user input information input so far, and the dialogue is repeatedly generated and presented to the user until an evaluation is made that the user's intention is clear.(21) The information processing system according to any one of (18) to (20), wherein the evaluation unit calculates a score indicating a degree of clarity of the user's intention based on the user input information, and evaluates the degree of clarity of the user's intention based on the score. (22) The information processing system according to any one of (18) to (21), wherein the content generation unit generates the content based on the material data and context information related to the material of the content. (23) The information processing system according to (22), wherein the context information includes at least any of the user's operation history related to the material of the content, information related to the user's emotions, information related to the user's concentration level, and vital information of the user. (24) The information processing system according to any one of (18) to (23), wherein the evaluation unit evaluates the degree of clarity of the user's intention for each element related to the content, and determines an evaluation result of whether the user's intention is clear overall based on the evaluation result for each element. (25) The information processing system described in (24), wherein the evaluation unit evaluates the clarity of the user's intention for each element based on at least one of a required level of clarity of information about the element and a degree of sufficiency of information about the element included in the user input information. (26) The information processing system described in (25), wherein the evaluation unit adjusts the required level of each element based on a history of modifications to the content. (27) The information processing system described in (25), wherein the evaluation unit determines the required level of each element in accordance with a setting operation by the user. (28) The information processing system described in any one of (24) to (27), wherein the dialogue generation unit generates the dialogue to prompt input of information about the element for which the evaluation of the clarity of the user's intention is equal to or lower than a predetermined evaluation. (29) The information processing system described in any one of (18) to (28), wherein the material data includes at least video data of gameplay, and the video of the content displays the user's operation details related to the gameplay.(30) The information processing system according to any one of (18) to (29), further comprising a control unit that causes a result of the content generation and a plurality of samples of at least one of a template and a presentation effect that can be used to generate the content to be presented, and acquires information on at least one of a template and a presentation effect that corresponds to the sample specified by the user as the user input information for modifying the content. (31) The information processing system according to any one of (18) to (30), wherein the dialogue generation unit generates the dialogue using a large-scale language model. (32) An information processing method including: an information processing system generating content based on material data that is the raw material of the content; when a user inputs user input information regarding modification of the content, evaluating the degree of clarity of the user's intention regarding modification of the content as indicated by the user input information; when the evaluation indicates that the user's intention is not clear, generating a dialogue to obtain additional user input information; and when the evaluation indicates that the user's intention is clear, generating the content that reflects the user's intention to modify based on the material data and the user input information. (33) A program that causes a computer to execute processes including: generating content based on material data that is the material of the content; when a user inputs user input information regarding modification of the content, evaluating the degree of clarity of the user's intention regarding modification of the content as indicated by the user input information; when the evaluation indicates that the user's intention is unclear, generating a dialogue to obtain additional user input information; and when the evaluation indicates that the user's intention is clear, generating the content that reflects the user's intention to modify based on the material data and the user input information.

[0301] REFERENCE SIGNS LIST 11 Information processing system, 21 Information processing device, 22 Display device, 31 Input unit, 33 Communication unit, 34 Control unit, 61 Rendering unit, 62 Clarity evaluation unit, 63 Dialogue generation unit, 64 Content generation unit, 201 Server, 211 Communication unit, 213 Control unit

Claims

1. An information processing system comprising: an evaluation unit that evaluates the degree of clarity of a user's intention regarding content generation as indicated by user input information; a dialogue generation unit that generates a dialogue to obtain additional user input information if the evaluation determines that the user's intention is unclear; and a content generation unit that generates the content based on the user input information if the evaluation determines that the user's intention is clear.

2. The information processing system according to claim 1, wherein the user input information is input using at least a speech by the user.

3. The information processing system of claim 1, wherein when the dialogue is presented to the user and additional user input information is input, the evaluation unit evaluates the clarity of the user's intention based on the user input information input so far, and the dialogue is repeatedly generated and presented to the user until it is evaluated that the user's intention is clear.

4. The information processing system of claim 1, wherein the evaluation unit calculates a score indicating the clarity of the user's intention based on the user input information, and evaluates the clarity of the user's intention based on the score.

5. The information processing system of claim 1, wherein, when the user inputs the user input information regarding the modification of the content, the evaluation unit evaluates the clarity of the user's intention based on the user input information regarding the modification of the content and the user input information that has been input so far, the dialogue generation unit generates the dialogue when it is evaluated that the user's intention is not clear, and the content generation unit generates the content that reflects the user's intention to modify when it is evaluated that the user's intention is clear.

6. An information processing system according to claim 1, wherein the content generation unit generates the content based on the user input information, material data that is the material for the content, and context information related to the material for the content.

7. An information processing system as described in claim 6, wherein the context information includes at least one of the user's operation history regarding the content material, information regarding the user's emotions, information regarding the user's level of concentration, and the user's vital signs.

8. An information processing system as described in claim 5, further comprising a control unit that presents the content generation results and multiple samples of at least one of templates and presentation effects that can be used to generate the content, and acquires information regarding at least one of templates and presentation effects that correspond to the samples specified by the user as the user input information regarding the modification of the content.

9. The information processing system according to claim 1, wherein the dialogue generation unit generates the dialogue using a large-scale language model.

10. An information processing method including: an information processing system evaluating the degree of clarity of a user's intention regarding content generation as indicated by user input information; if the evaluation indicates that the user's intention is unclear, generating a dialogue to obtain additional user input information; and if the evaluation indicates that the user's intention is clear, generating the content based on the user input information.

11. A program that causes a computer to perform processes including: evaluating the degree of clarity of a user's intention regarding the generation of content as indicated by user input information; generating a dialogue to obtain additional user input information if the evaluation indicates that the user's intention is unclear; and generating the content based on the user input information if the evaluation indicates that the user's intention is clear.

12. An information processing system comprising: a content generation unit that generates content based on material data that is the raw material for the content; an evaluation unit that, when user input information regarding modifications to the content is input by the user, evaluates the degree of clarity of the user's intention regarding modifications to the content as indicated by the user input information; and a dialogue generation unit that, when the evaluation indicates that the user's intention is not clear, generates a dialogue to obtain additional user input information, wherein the content generation unit, when the evaluation indicates that the user's intention is clear, generates the content that reflects the user's intention to modify based on the material data and the user input information.

13. The information processing system according to claim 12, wherein the user input information is input using at least a speech by the user.

14. The information processing system of claim 12, wherein when the dialogue is presented to the user and additional user input information is input, the evaluation unit evaluates the clarity of the user's intention based on the user input information input so far, and the dialogue is repeatedly generated and presented to the user until it is evaluated that the user's intention is clear.

15. An information processing system as described in claim 12, wherein the evaluation unit calculates a score indicating the clarity of the user's intention based on the user input information, and evaluates the clarity of the user's intention based on the score.

16. An information processing system according to claim 12, wherein the content generation unit generates the content based on the material data and context information relating to the material of the content.

17. An information processing system as described in claim 16, wherein the context information includes at least one of the user's operation history regarding the content material, information regarding the user's emotions, information regarding the user's level of concentration, and the user's vital signs.

18. The information processing system according to claim 12, wherein the dialogue generation unit generates the dialogue using a large-scale language model.

19. An information processing method including: an information processing system generating content based on material data that is the raw material of the content; when a user inputs user input information regarding modifications to the content, evaluating the degree of clarity of the user's intention regarding modifications to the content as indicated by the user input information; when the evaluation indicates that the user's intention is not clear, generating a dialogue to obtain additional user input information; and when the evaluation indicates that the user's intention is clear, generating the content that reflects the user's intention to modify based on the material data and the user input information.

20. A program that causes a computer to execute processes including: generating content based on material data that is the raw material for the content; when a user inputs user input information regarding modifications to the content, evaluating the degree of clarity of the user's intention regarding modifications to the content as indicated by the user input information; when it is evaluated that the user's intention is not clear, generating a dialogue to obtain additional user input information; and when it is evaluated that the user's intention is clear, generating the content that reflects the user's intention to modify based on the material data and the user input information.

Citation Information

Patent Citations

  • System and method for providing quiz game capable of presenting quiz created by user

    JP2015222561A

  • Transition between previous conversation contexts by an automated assistant

    JP2021515938A

  • Information processing system, information processing method, and program

    JP7396762B1