Information Processing Apparatus, Information Processing Method, and Program

The information processing apparatus addresses the issue of inappropriate responses in moving image interaction systems by using a machine learning model to generate responses based on user input and moving image content, ensuring contextually relevant interactions.

JP7697159B1Active Publication Date: 2025-06-23KDDI CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025024263
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-23
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

Existing systems that use machine learning models to respond to users viewing moving images often generate inappropriate responses, such as mismatched emotions or unrelated content.

Method used

An information processing apparatus that specifies the content of a moving image, generates two command sentences for a machine learning model to respond to user input and the moving image content, and outputs these responses to an associated information terminal, while also allowing for character-specific responses and interaction arbitration.

Benefits of technology

Enables appropriate and contextually relevant responses to users viewing moving images, improving the accuracy and relevance of interactions by considering both user input and moving image content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697159000001_ABST
    Figure 0007697159000001_ABST
Patent Text Reader

Abstract

Enable appropriate response to a user who is viewing a moving image. 【Solution means】The information processing apparatus 1 includes a specifying unit 132 that specifies the content of the moving image being viewed by the user, a first command sentence for causing a machine learning model to generate a response to an input sentence input by the user viewing the moving image, a second command sentence for causing a machine learning model to generate a response to the content of the moving image, a determining unit 133 that determines the first command sentence and the second command sentence, an input unit 134 that inputs the first command sentence and the second command sentence to the machine learning model, and an output unit 135 that outputs the response output by the machine learning model for the first command sentence and the second command sentence to an information terminal associated with the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program for responding to a user who is viewing a moving image.

Background Art

[0002] Patent Document 1 describes a system that uses a machine learning model to respond to comments input by a user when delivering a moving image to the user.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the system described in Patent Document 1, a machine learning model generates a response based on comments input by a user. Therefore, for example, there are cases where inappropriate responses are made to the moving image being viewed by the user, such as making a response with a happy atmosphere even though the scene of the moving image is sad, or making a response unrelated to the scene of the moving image.

[0005] Therefore, the present invention has been made in view of these points, and an object thereof is to be able to appropriately respond to a user who is viewing a moving image.

Means for Solving the Problems

[0006] The information processing apparatus according to the first aspect of the present invention includes: a specifying unit that specifies the content of a moving image being viewed by a user; a first command sentence for causing a machine learning model to generate a response to an input sentence input by the user viewing the moving image; a second command sentence for causing the machine learning model to generate a response to the content of the moving image; a determining unit that determines the first command sentence and the second command sentence; an input unit that inputs the first command sentence and the second command sentence into the machine learning model; and an output unit that outputs the responses output by the machine learning model for the first command sentence and the second command sentence to an information terminal associated with the user.

[0007] The information processing apparatus further includes a receiving unit that receives from the user a designation of a character appearing in the moving image, and the input unit may input the first command sentence and the second command sentence into the machine learning model corresponding to the character, which is selected from a plurality of the machine learning models pre-stored in a storage unit.

[0008] The machine learning model corresponding to the character may be pre-generated by machine learning at least one of the attributes of the character and the past utterances of the character.

[0009] The output unit may cause an avatar imitating the character to be displayed on the information terminal, and output the responses output by the machine learning model for the first command sentence and the second command sentence in association with the avatar.

[0010] The specifying unit may specify a section in the moving image in which the character appears by image recognition processing, and specify the content of the moving image in the specified section.

[0011] The storage unit pre-stores in association with each of a plurality of the characters character information regarding the character, and the machine learning model corresponding to the character may be configured to output a response using the character information associated with the character in the storage unit.

[0012] The specific unit specifies the viewing status of the moving images of a plurality of users who have viewed the moving images and have specified the character, and the determination unit may determine the second command sentence based on the viewing status of the plurality of users.

[0013] The moving image includes a plurality of scenes that can be selected as viewing targets by the user in response to an operation by the user, and the determination unit may determine the second command sentence based on the scene that the user is viewing among the plurality of scenes.

[0014] The output unit may output the same response to a plurality of the information terminals associated with a plurality of users who have viewed the moving image and have specified the character.

[0015] The determination unit may determine a single command sentence that is the first command sentence for causing the machine learning model to generate a response to the input sentence and is also the second command sentence for causing the machine learning model to generate a response to the content of the moving image.

[0016] The specific unit may specify, as the content of the moving image, an image displayed as the moving image and sound reproduced together with the moving image.

[0017] The determination unit determines whether a negative interaction has occurred among the plurality of users based on the input sentences input by the plurality of users who are viewing the moving image, and determines a third command sentence for causing the machine learning model to generate a response for arbitrating the negative interaction on the condition that it is determined that the negative interaction has occurred. The input unit inputs the third command sentence to the machine learning model, and the output unit may output the response output by the machine learning model for the third command sentence to the information terminals associated with the respective plurality of users.

[0018] The information processing method according to the second aspect of the present invention includes: a step of specifying the content of a moving image being viewed by a user, which is executed by a processor; a first instruction sentence for causing a machine learning model to generate a response to an input sentence input by the user viewing the moving image; a step of determining a second instruction sentence for causing the machine learning model to generate a response to the content of the moving image; a step of inputting the first instruction sentence and the second instruction sentence into the machine learning model; and a step of outputting the responses output by the machine learning model for the first instruction sentence and the second instruction sentence to an information terminal associated with the user.

[0019] The program according to the third aspect of the present invention causes a processor to execute: a step of specifying the content of a moving image being viewed by a user; a step of determining a first instruction sentence for causing a machine learning model to generate a response to an input sentence input by the user viewing the moving image, and a second instruction sentence for causing the machine learning model to generate a response to the content of the moving image; a step of inputting the first instruction sentence and the second instruction sentence into the machine learning model; and a step of outputting the responses output by the machine learning model for the first instruction sentence and the second instruction sentence to an information terminal associated with the user.

Advantages of the Invention

[0020] According to the present invention, there is an effect that an appropriate response can be made to a user viewing a moving image.

Brief Description of the Drawings

[0021]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0022] [Overview of Information Processing System S] FIG. 1 is a schematic diagram of an information processing system S according to the present embodiment. The information processing system S includes an information processing apparatus 1 and a user terminal 2. The information processing system S may include other servers, terminals, and other devices.

[0023] The information processing apparatus 1 is a computer that performs processing for outputting a response to a user who is viewing a moving image. The moving image is, for example, image content in which a plurality of characters appear. The characters appearing in the moving image are, for example, a person, an animal, or a virtual character shown in the moving image. The moving image may include audio reproduced together with the image. The information processing apparatus 1 acquires an input sentence such as a comment input by the user during viewing of the moving image on the user terminal 2, and outputs a response generated based on the input sentence and the content of the moving image to the user terminal 2.

[0024] The user terminal 2 is a computer used by the user. The user is a person who watches a moving image and inputs an input sentence such as a comment on the moving image. The user terminal 2 is, for example, a smartphone, a tablet terminal, or a personal computer. The user terminal 2 has an operation unit such as a touch panel or a keyboard for receiving operations, and a display unit such as a liquid crystal display for displaying information. The user terminal 2 is pre-associated with the user by setting identification information (Identifier: ID) for identifying the user who uses the user terminal 2. The user terminal 2 can communicate with the information processing apparatus 1 via a network.

[0025] An outline of the processing executed by the information processing system S according to the present embodiment will be described below. The information processing apparatus 1 identifies the content of the moving image that the user is watching. The content of the moving image includes, for example, an object shown in the moving image, a situation in the moving image, and the like. Further, the information processing apparatus 1 acquires an input sentence such as a comment input by the user who is watching the moving image.

[0026] The information processing apparatus 1 determines a first command sentence for causing the machine learning model to generate a response to the user's input sentence, and a second command sentence for causing the machine learning model to generate a response to the content of the moving image.

[0027] The information processing apparatus 1 inputs the determined first command sentence and second command sentence into the machine learning model. The machine learning model is configured to output a character string corresponding to the input character string when the character string is input, for example, by executing a known machine learning process. The information processing apparatus 1 outputs the response output by the machine learning model for the first command sentence and the second command sentence to the user terminal 2.

[0028] In this way, the information processing system S determines a command sentence to the machine learning model based on each of the user's input sentence and the content of the moving image. Thereby, the information processing system S can appropriately respond to the user who is watching the moving image in consideration of the input sentence input by the user and the content of the moving image.

[0029] [Configuration of Information Processing System S] FIG. 2 is a block diagram of the information processing system S according to the present embodiment. In FIG. 2, the arrows indicate the main data flow, and there may be data flows other than those shown in FIG. 2. In FIG. 2, each block represents a configuration in terms of functional units, not hardware (device) units. Therefore, the blocks shown in FIG. 2 may be implemented within a single device, or may be divided and implemented in a plurality of devices. The exchange of data between the blocks may be performed via any means such as a data bus, a network, a portable storage medium, etc.

[0030] The information processing apparatus 1 includes a communication unit 11, a storage unit 12, and a control unit 13. The information processing apparatus 1 may be configured by connecting two or more physically separated devices by wire or wirelessly. Also, the information processing apparatus 1 may be configured by a cloud which is a collection of computer resources.

[0031] The communication unit 11 has a communication controller for transmitting and receiving data to and from the user terminal 2 via a network. The communication unit 11 notifies the control unit 13 of the data received from the user terminal 2 via the network. Also, the communication unit 11 transmits the data output from the control unit 13 to the user terminal 2 via the network.

[0032] The storage unit 12 is a storage medium including a ROM (Read Only Memory), a RAM (Random Access Memory), a hard disk drive, an SSD (Solid State Drive), etc. The storage unit 12 stores in advance the program executed by the control unit 13. The storage unit 12 may be provided outside the information processing apparatus 1, and in that case, data may be exchanged with the control unit 13 via a network.

[0033] The control unit 13 includes a reception unit 131, an identification unit 132, a determination unit 133, an input unit 134, and an output unit 135. The control unit 13 is a processor such as a CPU (Central Processing Unit) or an NPU (Neural network Processing Unit), and functions as the reception unit 131, the identification unit 132, the determination unit 133, the input unit 134, and the output unit 135 by executing a program stored in the storage unit 12.

[0034] Hereinafter, the processing executed by the information processing system S will be described in detail. For example, the user selects a moving image that the user wishes to view from among a plurality of moving images. The user terminal 2 transmits information indicating the moving image designated by the user to the information processing apparatus 1. In the information processing apparatus 1, the reception unit 131 stores the information indicating the moving image received from the user terminal 2 in the storage unit 12.

[0035] The reception unit 131 receives from the user a designation of a character appearing in the moving image before or while the user views the moving image. The reception unit 131 causes the user terminal 2 to display a reception screen for receiving the designation of the character.

[0036] FIG. 3 is a schematic diagram of the user terminal 2 displaying the reception screen. The reception screen includes, for example, an area for receiving the designation of a character. The character is, for example, a person or an animal imaged as a subject in a live-action video. Also, the character may be, for example, a virtual character depicted in an animated video.

[0037] The reception unit 131 acquires, for example, information associating a moving image with a plurality of characters appearing in the moving image, which is stored in advance in the storage unit 12. The reception unit 131 displays, on the reception screen, the plurality of characters associated with the moving image that the user plans to view or is viewing in the acquired information. The user selects, from among the plurality of characters displayed on the reception screen, a character for which a response is desired to be received in a delivery screen described later.

[0038] The user performs an operation for specifying a character on the user terminal 2. The user terminal 2 transmits information indicating the character specified by the user to the information processing apparatus 1. In the information processing apparatus 1, the reception unit 131 stores the information indicating the character received from the user terminal 2 in the storage unit 12.

[0039] Although the reception unit 131 accepts the specification of the character on a screen different from the distribution screen in the example of FIG. 3, the specification of the character may be accepted on the distribution screen.

[0040] The moving image specified by the user is stored in advance in the storage unit 12 of the information processing apparatus 1 or in another storage unit that stores the moving image. The information processing apparatus 1 or another storage unit starts distributing the moving image by transmitting the moving image indicated by the information received from the user terminal 2 to the user terminal 2. The user terminal 2 displays a distribution screen A including the moving image received from the information processing apparatus 1 or another storage unit on the display unit.

[0041] FIG. 4 is a schematic diagram of the user terminal 2 displaying the distribution screen A. The distribution screen A includes a first scene display area A1 and a second scene display area A2. The first scene display area A1 is a display area having a larger area than the second scene display area A2.

[0042] The moving image is a live broadcast or live distribution (streaming) moving image, or a pre-recorded moving image. The moving image is, for example, a video of a concert, musical, stage, drama, movie, TV program, press conference, etc. Further, the moving image may be a two-dimensional or three-dimensional animated video.

[0043] In this embodiment, the moving image includes a plurality of scenes that can be selected by the user as the main viewing target according to the user's operation. The moving image includes, for example, an overall scene that images or depicts a predetermined range, and one or more individual scenes that image or depict a range narrower than the overall scene. The overall scene is, for example, a moving image of a scene including a plurality of characters. On the other hand, an individual scene is, for example, a moving image of a scene including any one of the plurality of characters. Each of the one or more individual scenes is pre-associated with the character included in the individual scene.

[0044] An individual scene is generated, for example, by cutting out a part of the overall scene. In this case, the individual scene is generated, for example, by automatically cutting out the area where each of the plurality of characters appears by performing known image recognition processing on the overall scene. Further, the individual scene may be generated by a human visually recognizing the overall scene and performing an operation of cutting out the area where each of the plurality of characters appears.

[0045] Further, an individual scene may be generated by imaging each of the plurality of characters with a camera different from the camera that imaged the overall scene. Further, the individual scene may be created as an independent scene rather than a part of the overall scene.

[0046] Each time point of the overall scene and each time point of the individual scene are pre-associated. When the moving image is played, either the overall scene or the individual scene is selected as the main viewing target at each time point of the moving image. For example, in response to the user performing an operation of selecting either the overall scene or the individual scene as the main viewing target at an arbitrary time point while viewing the moving image, the user terminal 2 starts displaying the selected one in the first scene display area A1 and the unselected one in the second scene display area A2 on the distribution screen A from that time point.

[0047] In addition, when a user performs an operation to select either the overall scene or an individual scene as the main viewing target at any point while watching a video, the user terminal 2 may start displaying the selected scene from that point on without displaying the unselected scene.

[0048] Furthermore, the distribution screen A includes an input area A3, a comment area A4, an avatar area A5, and a response area A6. The input area A3 is an area that accepts input of input text such as comments by a user watching the video. The comment area A4 is an area in which an input text input by a user watching the video is displayed. In addition to an input text input by a user using the user terminal 2, the comment area A4 may also display an input text input by a user other than the user. The avatar area A5 is an area in which an avatar, which will be described later, is displayed. The response area A6 is an area in which a response, which will be described later, is displayed.

[0049] The user performs an operation on the user terminal 2 to input an input sentence into the input area A3. The user terminal 2 transmits the input sentence input by the user watching the video to the information processing device 1. In the information processing device 1, the receiving unit 131 stores the input sentence received from the user terminal 2 in the storage unit 12.

[0050] The identification unit 132 identifies the content of the video being viewed by the user. The identification unit 132 identifies the content of the video, for example, at a predetermined time interval or when the reception unit 131 receives an input sentence from the user.

[0051] 5 is a schematic diagram for explaining the process of the identification unit 132 identifying the contents of the video. The contents of the video include, for example, at least one of an object appearing in the video and a situation in the video. The object appearing in the video is, for example, a captured or depicted character or other object. The situation in the video includes, for example, the action, expression, clothing, etc. of the captured or depicted character or other object. The situation in the video may also include the date, time, location, weather, etc. in the video.

[0052] The specifying unit 132 specifies the content of the moving image by, for example, performing known image recognition processing on the moving image being viewed by the user. Alternatively, the specifying unit 132 may acquire the content of the moving image being viewed by the user from a database in which the moving image and its content are associated in advance.

[0053] The specifying unit 132 may specify the sound reproduced together with the moving image as the content of the moving image. In this case, the specifying unit 132 specifies the sound reproduced together with the moving image by, for example, performing known sound recognition processing on the moving image being viewed by the user.

[0054] The specifying unit 132 specifies the section in which the character designated by the user appears in the moving image, specifies the content of the moving image in the specified section, and does not necessarily specify the content of the moving image outside the section. Thereby, the information processing system S can make a response in consideration of only the content of the portion related to the character designated by the user in the moving image.

[0055] Next, the specifying unit 132 specifies the viewing status of the user's moving image. The viewing status of the moving image includes, for example, the date and time when the user is viewing the moving image, the scene selected by the user as the main viewing target among the plurality of scenes included in the moving image, and the like. Further, the specifying unit 132 may specify the viewing status of the moving images of a plurality of users who have viewed the moving image and who have specified the character specified by the user viewing the moving image (the user viewing the distribution screen A in FIG. 4).

[0056] The storage unit 12 of the information processing apparatus 1 or another storage unit that stores the machine learning model stores in advance a plurality of machine learning models corresponding to a plurality of characters. The machine learning model is configured to output a response corresponding to the input command sentence when the command sentence is input.

[0057] The machine learning model is pre-generated, for example, by machine learning, using at least one of the attributes of each of a plurality of characters and the past utterances of the character as learning data by a known machine learning process (such as DNN (Deep Neural Network)), and is stored in the storage unit 12 of the information processing apparatus 1 or other storage unit in association with the character.

[0058] The attributes of the character are, for example, the character's age, gender, place of residence, place of origin, experience, hobbies, preferences, etc. The past utterances of the character are, for example, the content that the character has uttered in the past on SNS (Social Networking Service), chat, articles, TV programs, radio programs, etc. Thereby, the machine learning model can generate a response according to the characteristics of each character.

[0059] Also, the machine learning model may be configured to output a response using character information about the character. In this case, the storage unit 12 of the information processing apparatus 1 or other storage unit stores in advance a plurality of characters and character information related to each character in association with each other. The character information includes, for example, any information related to the character other than the learning data, such as the appearance information or release information of the character. Thereby, the machine learning model can generate a response with reference to character information other than the learning data.

[0060] When a predetermined response condition is satisfied, the determination unit 133 determines a first command sentence for causing the machine learning model to generate a response to the input sentence input by the user who is viewing the moving image, and a second command sentence for causing the machine learning model to generate a response to the content of the moving image. The first command sentence is a command sentence for causing the machine learning model to generate a response including content (such as an answer to a question) along with the input sentence input by the user. The second command sentence is a command sentence for causing the machine learning model to generate a response including content (such as additional information about the subject shown in the moving image) along with the content of the video.

[0061] FIG. 6(a) and FIG. 6(b) are schematic diagrams for explaining the process in which the determination unit 133 determines the first instruction sentence and the second instruction sentence. The response condition is a condition indicating an opportunity to respond to a user who is watching a moving image. The response condition is, for example, that the reception unit 131 has received an input sentence input by the user. Further, the response condition may be that in the content of the moving image specified by the specifying unit 132, the date and time in the moving image is a predetermined date and time or time zone, the weather in the moving image is a predetermined weather, the location in the moving image is a predetermined location, predetermined characters in the moving image have communicated with each other, the music playing in the moving image is a predetermined name or a predetermined provider, and the like.

[0062] Further, the response condition may be that in the viewing status of one or a plurality of users specified by the specifying unit 132, the date and time when the user is watching the moving image is a predetermined date and time or time zone, the number of views of the moving image by a plurality of users is more or less than a predetermined reference value, and the like. The response condition is not limited to the specific conditions shown here, and may include other conditions related to the moving image or the user.

[0063] When the response condition is satisfied, the determination unit 133 acquires, as an opportunity, a rule for determining an instruction sentence, which is stored in advance in the storage unit 12. The rule for determining an instruction sentence is, for example, a template for each of the first instruction sentence and the second instruction sentence.

[0064] For example, the determination unit 133 determines a first instruction sentence for causing a machine learning model to generate a response to the input sentence by applying the input sentence acquired by the reception unit 131 to the template of the first instruction sentence. For example, the determination unit 133 determines a second instruction sentence for causing a machine learning model to generate a response to the content of the moving image by applying the content of the moving image specified by the specifying unit 132 to the template of the second instruction sentence. Thereby, the information processing system S can automatically output information related to the input sentence and information related to the content of the video using the machine learning model.

[0065] In addition, the determination unit 133 may determine the second command sentence based on the scene being viewed by the user among the plurality of scenes included in the moving image, in addition to the content of the moving image specified by the specifying unit 132. In this case, the determination unit 133 determines different second command sentences according to, for example, the scene being viewed by the user among the plurality of scenes included in the moving image. Thereby, the information processing system S can automatically output a response according to the scene being viewed by the user, such as a comment for guiding a user viewing the entire scene to an individual scene, or a comment for guiding a user viewing an individual scene to the entire scene.

[0066] In addition, the determination unit 133 may determine the second command sentence based on the viewing status of the plurality of users specified by the specifying unit 132, in addition to the content of the moving image specified by the specifying unit 132. In this case, the determination unit 133 calculates, for example, values indicating the viewing status of the plurality of users (the number of views per scene, the number of comments per time, etc.), and determines different second command sentences according to the calculated values. Thereby, the information processing system S can automatically output a response according to the viewing status of the plurality of users, such as a comment for guiding to a scene with a high viewing ratio by the plurality of users among the plurality of scenes included in the moving image, or a comment corresponding to the fact that the number of input sentences input by the plurality of users is large (being lively).

[0067] In the example of FIG. 6(a), the determination unit 133 determines a first command sentence for causing the machine learning model to generate a response to the input sentence and a second command sentence for causing the machine learning model to generate a response to the content of the moving image, respectively. The input unit 134 described later inputs the first command sentence and the second command sentence into the machine learning model at different timings. Thereby, the information processing system S can output a response to the user's input sentence and a response to the content of the video independently.

[0068] In the example of FIG. 6(b), the determination unit 133 determines a single instruction sentence that is the first instruction sentence for causing the machine learning model to generate a response to the input sentence and is also the second instruction sentence for causing the machine learning model to generate a response to the content of the moving image. The input unit 134 described below inputs the single instruction sentence that is the first instruction sentence and also the second instruction sentence into the machine learning model. Thereby, the information processing system S can output a response considering the content of the video with respect to the input sentence of the user.

[0069] The determination unit 133 may execute both the process of determining each of the first instruction sentence and the second instruction sentence illustrated in FIG. 6(a) and the process of determining a single instruction sentence that is the first instruction sentence and also the second instruction sentence illustrated in FIG. 6(b).

[0070] The input unit 134 inputs the first instruction sentence and the second instruction sentence determined by the determination unit 133 into the machine learning model. The input unit 134 selects, for example, a machine learning model corresponding to the character designated by the user received by the reception unit 131 from among a plurality of machine learning models stored in advance in the storage unit 12 of the information processing apparatus 1 or other storage units. The input unit 134 inputs the first instruction sentence and the second instruction sentence into the selected machine learning model. Further, the input unit 134 may input a single instruction sentence that is the first instruction sentence and also the second instruction sentence into the selected machine learning model.

[0071] The output unit 135 acquires the response output by the machine learning model into which the first instruction sentence and the second instruction sentence are input, and outputs the acquired response to the user terminal 2 associated with the user.

[0072] FIGS. 7(a) and 7(b) are schematic diagrams for explaining the process in which the output unit 135 outputs a response. The output unit 135, for example, displays an avatar imitating the character designated by the user in the avatar area A5 on the distribution screen A of FIG. 4. The avatar is a pre-generated still image or moving image. Further, the avatar may be dynamically generated. Information for displaying the avatar is, for example, stored in advance in the storage unit 12 in association with the character.

[0073] The output unit 135 causes the response area A6 to display, for example, in association with the avatar displayed in the avatar area A5, the response output by the machine learning model for the first command sentence and the second command sentence. Thereby, the information processing system S can give the user a feeling of actually conversing with the character designated by the user.

[0074] In the example of FIG. 7(a), the output unit 135 outputs different responses for each user to a plurality of user terminals 2 associated with a plurality of users who have viewed a moving image and designated the same character. That is, the response output according to the command sentence for a certain user's input sentence is displayed only on the user terminal 2 of that user and is not displayed on the user terminals 2 of other users other than that user. Thereby, the information processing system S can provide the response to the user's input sentence only to that user and protect the privacy between users.

[0075] In the example of FIG. 7(b), the output unit 135 outputs the same response to a plurality of user terminals 2 associated with a plurality of users who have viewed a moving image and designated the same character. That is, the response output according to the command sentence for a certain user's input sentence is displayed on the user terminal 2 of that user and on the user terminals 2 of other users who have designated the same character as that user. Thereby, the information processing system S can also provide the response to the user's input sentence to other users and promote conversation among a plurality of users.

[0076] Further, when the response output by the machine learning model includes information on a predetermined product, the output unit 135 may output access information (for example, URI: Uniform Resource Identifier) for accessing a website (such as a product sales page) related to the product to the user terminal 2. In this case, the storage unit 12 stores in advance, in association with each of the plurality of products, the access information related to the product. Thereby, the information processing system S can make it easier for the user to access the information on the product related to the response output by the machine learning model.

[0077] When a plurality of users watching a moving image input input sentences such as comments on the moving image, negative interactions may occur among the plurality of users. Negative interactions include, for example, slander, discriminatory words, obscene words, etc. Therefore, the information processing system S may output a response for arbitrating negative interactions occurring among a plurality of users.

[0078] The determination unit 133 determines whether negative interactions are occurring among a plurality of users based on the input sentences input by the plurality of users watching the moving image. The determination unit 133 determines, for example, that negative interactions are occurring on the condition that any of the keywords of a plurality of negative contents indicated by the information pre-stored in the storage unit 12 is included in the input sentences received by the reception unit 131 from the plurality of users. Further, the determination unit 133 may determine whether negative interactions are occurring by the machine learning model by inputting the input sentences into the machine learning model that has been generated in advance by machine learning using negative interactions as learning data.

[0079] The determination unit 133 determines a third command sentence for causing the machine learning model to generate a response for arbitrating negative interactions on the condition that it is determined that negative interactions are occurring.

[0080] FIG. 8 is a schematic diagram for explaining the process in which the determination unit 133 determines the third command sentence. The determination unit 133 acquires, on the condition that it is determined that negative interactions are occurring, the rules for determining command sentences pre-stored in the storage unit 12. The rules for determining command sentences are, for example, the templates of the third command sentence. The third command sentence is a command sentence for causing the machine learning model to generate a response including content for arbitrating negative interactions.

[0081] The determination unit 133 determines a third instruction sentence for causing the machine learning model to generate a response to the input sentence, for example, by applying the input sentence determined to have a negative interaction to the template of the third instruction sentence.

[0082] The input unit 134 inputs the third instruction sentence determined by the determination unit 133 to the machine learning model. The input unit 134 selects, for example, a machine learning model corresponding to the character specified by the user received by the reception unit 131 from among a plurality of machine learning models stored in advance in the storage unit 12 of the information processing apparatus 1 or other storage units. The input unit 134 inputs the third instruction sentence to the selected machine learning model.

[0083] The output unit 135 acquires the response output by the machine learning model to which the third instruction sentence is input, and outputs the acquired response to the user terminal 2 associated with each of the plurality of users. Thereby, the information processing system S can automatically guide the user not to have a negative interaction by using the machine learning model corresponding to the character specified by the user.

[0084] Furthermore, the output unit 135 may not output the input sentence input by the user who is determined to have a negative interaction to the input sentences input by the user after the determination to the user terminal 2 associated with another user different from the user. Thereby, the information processing system S can isolate the user who has had a negative interaction from other users and make it less likely for problems to occur among users.

[0085] [Flow of the information processing method] FIG. 9 is a diagram showing a flowchart of an exemplary information processing method executed by the information processing system S according to the present embodiment. The reception unit 131 receives from the user a designation of a character appearing in the moving image before or while the user is viewing the moving image (S11). The reception unit 131 receives the input sentence input by the user viewing the moving image (S12).

[0086] The specific unit 132 identifies the content of the moving image that the user is viewing (S13). The specific unit 132 identifies the viewing status of the user's moving image (S14).

[0087] The information processing apparatus 1 determines whether a predetermined response condition, which is a condition indicating an opportunity to respond to the user viewing the moving image, is satisfied. If the response condition is not satisfied (NO in S15), the information processing apparatus 1 returns to step S12 to continue the processing.

[0088] When the response condition is satisfied (YES in S15), the determination unit 133 determines a first command sentence for causing the machine learning model to generate a response to the input sentence input by the user viewing the moving image, and a second command sentence for causing the machine learning model to generate a response to the content of the moving image (S16).

[0089] The input unit 134 selects a machine learning model corresponding to the character designated by the user received by the reception unit 131 from a plurality of machine learning models stored in advance in the storage unit 12 of the information processing apparatus 1 or other storage units. The input unit 134 inputs the first command sentence and the second command sentence determined by the determination unit 133 to the selected machine learning model (S17).

[0090] The output unit 135 acquires the response output by the machine learning model to which the first command sentence and the second command sentence are input, and outputs the acquired response to the user terminal 2 associated with the user (S18).

[0091] The information processing apparatus 1 determines whether a predetermined end condition (such as the user having performed an end operation), which is a condition for ending the processing, is satisfied. If the end condition is not satisfied (NO in S19), the information processing apparatus 1 returns to step S12 to continue the processing. If the end condition is satisfied (YES in S19), the information processing apparatus 1 ends the processing.

[0092] Note that according to the present invention, it becomes possible to contribute to Goal 9, "Build the infrastructure for industry and technological innovation," of the Sustainable Development Goals (SDGs) led by the United Nations.

[0093] As described above, the present invention has been described using embodiments. However, the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist. For example, all or part of the device can be configured by functionally or physically dispersing and integrating it in any unit. Also, new embodiments resulting from any combination of a plurality of embodiments are included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination have the effects of the original embodiments combined.

Explanation of Reference Numerals

[0094] S Information processing system 1 Information processing device 11 Communication unit 12 Storage unit 13 Control unit 131 Reception unit 132 Identification unit 133 Decision unit 134 Input unit 135 Output unit 2 User terminal

Claims

1. A reception unit that receives from a user a designation of a character that appears in a video; an identification unit that identifies the content of the video being viewed by the user and the viewing status of the video of a plurality of users who have viewed the video and designated the character; a determination unit that determines a first command statement for causing a machine learning model to generate a response to an input sentence input by a user who is viewing the video, and determines a second command statement for causing the machine learning model to generate a response to content of the video based on the viewing status of the multiple users; an input unit that inputs the first command statement and the second command statement to the machine learning model corresponding to the character, the machine learning model being selected from a plurality of machine learning models pre-stored in a storage unit; an output unit that outputs a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; An information processing device having the above configuration.

2. A determination unit that determines the content of a moving image being viewed by a user, the moving image including a plurality of scenes that can be selected as a target to be viewed by the user in response to an operation by the user; a determination unit that determines a first command statement for causing a machine learning model to generate a response to an input sentence input by a user who is watching the video, and determines a second command statement for causing the machine learning model to generate a response to content of the video based on a scene among the plurality of scenes that the user is watching; an input unit that inputs the first command statement and the second command statement to the machine learning model; an output unit that outputs a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; An information processing device having the above configuration.

3. A reception unit that receives from a user a designation of a character that appears in a moving image; An identification unit that identifies the content of the video being viewed by the user; A determination unit that determines a first command statement for causing a machine learning model to generate a response to an input sentence input by a user watching the video, and a second command statement for causing the machine learning model to generate a response to content of the video; an input unit that inputs the first command statement and the second command statement to the machine learning model corresponding to the character, the machine learning model being selected from a plurality of machine learning models pre-stored in a storage unit; an output unit that outputs a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; having the output unit outputs the same response to the plurality of information terminals associated with the plurality of users who have viewed the video and designated the character. Information processing device.

4. An identification unit that identifies the content of a video image being viewed by a user; A determination unit that determines a first command statement for causing a machine learning model to generate a response to an input sentence input by a user watching the video, and a second command statement for causing the machine learning model to generate a response to content of the video; an input unit that inputs the first command statement and the second command statement to the machine learning model; an output unit that outputs a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; having the determination unit determines whether or not a negative exchange is occurring between the plurality of users who are watching the video based on the input sentences input by the plurality of users, and, on a condition that it is determined that the negative exchange is occurring, determines a third command statement for causing the machine learning model to generate a response that mediates the negative exchange; The input unit inputs the third command statement to the machine learning model; The output unit outputs a response output by the machine learning model in response to the third command statement to the information terminal associated with each of the multiple users. Information processing device.

5. The machine learning model corresponding to the character is generated in advance by machine learning at least one of the attributes of the character and past utterances of the character. The information processing device according to claim 1 .

6. The output unit displays an avatar that resembles the character on the information terminal, and outputs a response output by the machine learning model to the first command statement and the second command statement in association with the avatar. The information processing device according to claim 1 .

7. the identification unit identifies a section in the video in which the character appears by image recognition processing, and identifies the content of the video in the identified section; The information processing device according to claim 1 .

8. the storage unit stores in advance each of the plurality of characters in association with character information relating to the character; The machine learning model corresponding to the character is configured to output a response using the character information associated with the character in the storage unit. The information processing device according to claim 1 .

9. the determination unit determines a single command statement, the single command statement being the first command statement for causing the machine learning model to generate a response to the input sentence and a second command statement for causing the machine learning model to generate a response to content of the video; The information processing device according to claim 1 .

10. the specifying unit specifies, as the content of the moving image, an image to be displayed as the moving image and a sound to be played together with the moving image; The information processing device according to claim 1 .

11. The processor executes A step of accepting, from a user, a designation of a character appearing in a moving image; Identifying the content of the video being viewed by the user and the viewing status of the video of a plurality of users who have viewed the video and designated the character; determining a first command statement for causing a machine learning model to generate a response to an input sentence input by a user who is viewing the video, and determining a second command statement for causing the machine learning model to generate a response to the content of the video based on the viewing status of the multiple users; inputting the first command statement and the second command statement into the machine learning model corresponding to the character, selected from a plurality of the machine learning models pre-stored in a storage unit; outputting a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; An information processing method comprising the steps of:

12. A processor executes: identifying content of a video being viewed by a user, the video including a plurality of scenes selectable as a subject to be viewed by the user in response to an operation by the user; determining a first command statement for causing a machine learning model to generate a response to an input sentence input by a user viewing the video, and determining a second command statement for causing the machine learning model to generate a response to the content of the video based on a scene being viewed by the user among the plurality of scenes; inputting the first statement and the second statement into the machine learning model; outputting a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; An information processing method comprising the steps of:

13. A processor executes: A step of accepting, from a user, a designation of a character appearing in a moving image; identifying the content of the video being viewed by the user; determining a first command statement for causing a machine learning model to generate a response to an input sentence input by a user watching the video, and a second command statement for causing the machine learning model to generate a response to content of the video; inputting the first command statement and the second command statement into the machine learning model corresponding to the character, selected from a plurality of the machine learning models pre-stored in a storage unit; outputting a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; having In the outputting step, the same response is output to a plurality of the information terminals associated with a plurality of users who have viewed the video and designated the character. Information processing methods.

14. A method for executing a program comprising: Identifying the content of the video being viewed by the user; determining a first command statement for causing a machine learning model to generate a response to an input sentence input by a user watching the video, and a second command statement for causing the machine learning model to generate a response to content of the video; inputting the first statement and the second statement into the machine learning model; outputting a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; determining whether or not a negative exchange is occurring between the plurality of users who are watching the video based on the input sentences input by the plurality of users, and, on condition that it is determined that the negative exchange is occurring, determining a third command sentence for causing the machine learning model to generate a response that mediates the negative exchange; inputting the third statement into the machine learning model; outputting a response output by the machine learning model to the third command statement to the information terminal associated with each of the plurality of users; An information processing method comprising the steps of:

15. The processor: A step of accepting, from a user, a designation of a character appearing in a moving image; Identifying the content of the video being viewed by the user and the viewing status of the video of a plurality of users who have viewed the video and designated the character; determining a first command statement for causing a machine learning model to generate a response to an input sentence input by a user who is viewing the video, and determining a second command statement for causing the machine learning model to generate a response to the content of the video based on the viewing status of the multiple users; inputting the first command statement and the second command statement into the machine learning model corresponding to the character, selected from a plurality of the machine learning models pre-stored in a storage unit; outputting a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; A program to execute.

16. A processor comprising: identifying content of a video being viewed by a user, the video including a plurality of scenes selectable as a subject to be viewed by the user in response to an operation by the user; determining a first command statement for causing a machine learning model to generate a response to an input sentence input by a user watching the video, and a second command statement for causing the machine learning model to generate a response to the content of the video based on a scene being watched by the user among the plurality of scenes; inputting the first statement and the second statement into the machine learning model; outputting a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; A program to execute.

17. A processor comprising: A step of accepting, from a user, a designation of a character appearing in a moving image; identifying the content of the video being viewed by the user; determining a first command statement for causing a machine learning model to generate a response to an input sentence input by a user watching the video, and a second command statement for causing the machine learning model to generate a response to content of the video; inputting the first command statement and the second command statement into the machine learning model corresponding to the character, selected from a plurality of the machine learning models pre-stored in a storage unit; outputting a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; Run the command, In the outputting step, the same response is output to a plurality of the information terminals associated with a plurality of users who have viewed the video and designated the character. program.

18. A processor comprising: Identifying the content of the video being viewed by the user; determining a first command statement for causing a machine learning model to generate a response to an input sentence input by a user watching the video, and a second command statement for causing the machine learning model to generate a response to content of the video; inputting the first statement and the second statement into the machine learning model; outputting a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; determining whether or not a negative exchange is occurring between the plurality of users who are watching the video based on the input sentences input by the plurality of users, and, on condition that it is determined that the negative exchange is occurring, determining a third command sentence for causing the machine learning model to generate a response that mediates the negative exchange; inputting the third statement into the machine learning model; outputting a response output by the machine learning model to the third command statement to the information terminal associated with each of the plurality of users; A program to execute.

Citation Information

Patent Citations

  • Answer selection device, model learning device, method for selecting answer, method for learning model, and program

    JP2019192073A

  • Interactive program, device, and method for expressing sense of listening of character in accordance with user's emotion

    JP2022054326A

  • Persona chatbot control method and system

    JP2022180282A

  • Automatic captioning of audible portions of content on a computing device - Patent Application 20070122997

    JP2022530201A

  • Information distribution apparatus, information distribution method, and program

    JP2023050587A