Information Processing Apparatus, Information Processing Method, and Program

The information processing apparatus addresses inappropriate responses in machine learning systems by specifying image content and generating contextually relevant responses, enhancing user interaction with moving images.

JP7698810B1Active Publication Date: 2025-06-25KDDI CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2025079208
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-25
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

Existing systems generate inappropriate responses to moving images, such as happy responses during sad scenes or responses unrelated to the scene, due to the limitations of machine learning models.

Method used

An information processing apparatus that includes a reception unit for specifying characters in a moving image, a specifying unit for identifying the content of the image, and a determination unit for generating responses using machine learning models based on command sentences that reflect the specified image content, ensuring appropriate responses.

Benefits of technology

Enables appropriate responses to users viewing moving images by considering the content and context of the image, improving user engagement and interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698810000001_ABST
    Figure 0007698810000001_ABST
Patent Text Reader

Abstract

Enable appropriate response to a user who is viewing a moving image. 【Solution means】The information processing apparatus 1 includes a specifying unit 132 that specifies the content of the moving image being viewed by the user, a first command sentence for causing a machine learning model to generate a response to an input sentence input by the user viewing the moving image, a second command sentence for causing the machine learning model to generate a response to the content of the moving image, a determining unit 133 that determines the first command sentence and the second command sentence, an input unit 134 that inputs the first command sentence and the second command sentence to the machine learning model, and an output unit 135 that outputs the response output by the machine learning model for the first command sentence and the second command sentence to an information terminal associated with the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program for responding to a user who is viewing a moving image.

Background Art

[0002] Patent Document 1 describes a system that uses a machine learning model to respond to comments input by a user when delivering a moving image to the user.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the system described in Patent Document 1, a machine learning model generates a response based on comments input by a user. Therefore, for example, there are cases where an inappropriate response is made to the moving image being viewed by the user, such as making a response with a happy atmosphere even though the scene of the moving image is sad, or making a response unrelated to the scene of the moving image.

[0005] Therefore, the present invention has been made in view of these points, and an object thereof is to be able to appropriately respond to a user who is viewing a moving image.

Means for Solving the Problems

[0006] The information processing apparatus according to the first aspect of the present invention includes a reception unit that receives from a user a specification of a character appearing in a moving image, a specification unit that specifies the content of the moving image being viewed by the user, a first command sentence for causing a machine learning model to generate a response to an input sentence input by the user viewing the moving image, and a second command sentence for causing the machine learning model to generate a response to the content of the moving image. A determination unit that determines, an input unit that inputs the first command sentence and the second command sentence to the machine learning model, and an output unit that outputs the response output by the machine learning model in response to the first command sentence and the second command sentence to an information terminal associated with the user. The specifying unit specifies a section in which the character appears in the moving image by image recognition processing, and specifies the content of the moving image in the specified section. The determining unit determines the second command sentence for causing the machine learning model to generate a response that reflects the content of the moving image in the specified section and does not reflect the content of the moving image outside the specified section.

[0007] The input unit may input the first command sentence and the second command sentence to the machine learning model corresponding to the character, which is selected from a plurality of the machine learning models stored in advance in the storage unit.

[0008] The machine learning model corresponding to the character may be generated in advance by machine learning at least one of the attributes of the character and the past utterances of the character.

[0009] The output unit may display an avatar imitating the character on the information terminal and output the response output by the machine learning model in response to the first command sentence and the second command sentence in association with the avatar.

[0010] The memory unit stores in advance, in an associated manner, each of the plurality of characters and character information regarding the character, and the machine learning model corresponding to the character may be configured to output a response using the character information associated with the character in the memory unit.

[0011] The specifying unit specifies the viewing status of the moving images of a plurality of users who have viewed the moving image and have specified the character, and the determining unit may determine the second command sentence based on the viewing status of the plurality of users.

[0012] The moving image includes a plurality of scenes that can be selected as viewing targets by the user in response to an operation by the user, and the determining unit may determine the second command sentence based on the scene that the user is viewing among the plurality of scenes.

[0013] The output unit may output the same response to a plurality of information terminals associated with a plurality of users who have viewed the moving image and have specified the character.

[0014] The determining unit may determine a single command sentence that is the first command sentence for causing the machine learning model to generate a response to the input sentence and is also the second command sentence for causing the machine learning model to generate a response to the content of the moving image.

[0015] The specifying unit may specify, as the content of the moving image, an image displayed as the moving image and sound reproduced together with the moving image.

[0016] The determination unit determines whether a negative interaction has occurred among the plurality of users based on the input sentences input by the plurality of users who are watching the moving image, and determines a third command sentence for causing the machine learning model to generate a response for arbitrating the negative interaction on the condition that it is determined that the negative interaction has occurred. The input unit inputs the third command sentence to the machine learning model, and the output unit may output the response output by the machine learning model for the third command sentence to the information terminals associated with the respective plurality of users.

[0017] The information processing method according to the second aspect of the present invention includes steps executed by a processor: receiving from a user a specification of a character appearing in a moving image; specifying the content of the moving image being watched by the user; determining a first command sentence for causing a machine learning model to generate a response to an input sentence input by a user watching the moving image, and a second command sentence for causing the machine learning model to generate a response to the content of the moving image; inputting the first command sentence and the second command sentence to the machine learning model; and outputting the responses output by the machine learning model for the first command sentence and the second command sentence to an information terminal associated with the user. In the specifying step, an image recognition process is used to specify a section in the moving image in which the character appears, and the content of the moving image in the specified section is specified. In the determining step, the second command sentence for causing the machine learning model to generate a response that reflects the content of the moving image in the specified section and does not reflect the content of the moving image outside the specified section is determined.

[0018] The program according to the third aspect of the present invention causes a processor to execute steps of: receiving from a user a specification of a character appearing in a moving image; specifying the content of the moving image being viewed by the user; determining a first instruction statement for causing a machine learning model to generate a response to an input sentence input by the user viewing the moving image, and a second instruction statement for causing the machine learning model to generate a response to the content of the moving image; inputting the first instruction statement and the second instruction statement into the machine learning model; and outputting the responses output by the machine learning model for the first instruction statement and the second instruction statement to an information terminal associated with the user. In the specifying step, an image recognition process is used to specify a section in the moving image in which the character appears, and the content of the moving image in the specified section is specified. In the determining step, the second instruction statement for causing the machine learning model to generate a response that reflects the content of the moving image in the specified section and does not reflect the content of the moving image outside the specified section is determined.

[0019]

Advantages of the Invention

[0020] According to the present invention, there is an effect that an appropriate response can be made to a user viewing a moving image.

Brief Description of the Drawings

[0021]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiment for Carrying Out the Invention

[0022] [Overview of Information Processing System S] FIG. 1 is a schematic diagram of the information processing system S according to the present embodiment. The information processing system S includes an information processing apparatus 1 and a user terminal 2. The information processing system S may include other devices such as servers and terminals.

[0023] The information processing apparatus 1 is a computer that performs processing for outputting a response to a user who is viewing a moving image. The moving image is, for example, image content in which a plurality of characters appear. The characters appearing in the moving image are, for example, a person, an animal, or a virtual character shown in the moving image. The moving image may include audio reproduced together with the image. The information processing apparatus 1 acquires an input sentence such as a comment input by the user while viewing the moving image on the user terminal 2, and outputs a response generated based on the input sentence and the content of the moving image to the user terminal 2.

[0024] The user terminal 2 is a computer used by the user. The user is a person who watches a moving image and inputs an input sentence such as a comment on the moving image. The user terminal 2 is, for example, a smartphone, a tablet terminal, or a personal computer. The user terminal 2 has an operation unit such as a touch panel or a keyboard for receiving operations, and a display unit such as a liquid crystal display for displaying information. The user terminal 2 is pre-associated with the user by setting identification information (Identifier: ID) for identifying the user who uses the user terminal 2. The user terminal 2 can communicate with the information processing apparatus 1 via a network.

[0025] The outline of the processing executed by the information processing system S according to the present embodiment will be described below. The information processing apparatus 1 identifies the content of the moving image being viewed by the user. The content of the moving image includes, for example, an object shown in the moving image, a situation in the moving image, and the like. Further, the information processing apparatus 1 acquires an input sentence such as a comment input by the user who is watching the moving image.

[0026] The information processing apparatus 1 determines a first command sentence for causing the machine learning model to generate a response to the input sentence of the user, and a second command sentence for causing the machine learning model to generate a response to the content of the moving image.

[0027] The information processing apparatus 1 inputs the determined first command sentence and second command sentence into the machine learning model. The machine learning model is configured to output a character string corresponding to the input character string when the character string is input, for example, by executing a known machine learning process. The information processing apparatus 1 outputs the response output by the machine learning model for the first command sentence and the second command sentence to the user terminal 2.

[0028] In this way, the information processing system S determines a command sentence to the machine learning model based on each of the input sentence of the user and the content of the moving image. Thereby, the information processing system S can appropriately respond to the user who is watching the moving image in consideration of the input sentence input by the user and the content of the moving image.

[0029] [Configuration of Information Processing System S] FIG. 2 is a block diagram of the information processing system S according to the present embodiment. In FIG. 2, the arrows indicate the main data flow, and there may be data flows other than those shown in FIG. 2. In FIG. 2, each block represents a configuration in terms of functional units, not hardware (device) units. Therefore, the blocks shown in FIG. 2 may be implemented within a single device, or may be divided and implemented in a plurality of devices. The exchange of data between the blocks may be performed via any means such as a data bus, a network, a portable storage medium, etc.

[0030] The information processing apparatus 1 includes a communication unit 11, a storage unit 12, and a control unit 13. The information processing apparatus 1 may be configured by connecting two or more physically separated devices by wire or wirelessly. Also, the information processing apparatus 1 may be configured by a cloud which is a collection of computer resources.

[0031] The communication unit 11 has a communication controller for transmitting and receiving data to and from the user terminal 2 via a network. The communication unit 11 notifies the control unit 13 of the data received from the user terminal 2 via the network. Also, the communication unit 11 transmits the data output from the control unit 13 to the user terminal 2 via the network.

[0032] The storage unit 12 is a storage medium including a ROM (Read Only Memory), a RAM (Random Access Memory), a hard disk drive, an SSD (Solid State Drive), etc. The storage unit 12 stores in advance the program executed by the control unit 13. The storage unit 12 may be provided outside the information processing apparatus 1, and in that case, data may be exchanged with the control unit 13 via a network.

[0033] The control unit 13 includes a reception unit 131, an identification unit 132, a determination unit 133, an input unit 134, and an output unit 135. The control unit 13 is a processor such as a CPU (Central Processing Unit) or an NPU (Neural network Processing Unit), and functions as the reception unit 131, the identification unit 132, the determination unit 133, the input unit 134, and the output unit 135 by executing a program stored in the storage unit 12.

[0034] Hereinafter, the processing executed by the information processing system S will be described in detail. For example, the user selects a moving image that the user wishes to view from among a plurality of moving images. The user terminal 2 transmits information indicating the moving image designated by the user to the information processing apparatus 1. In the information processing apparatus 1, the reception unit 131 stores the information indicating the moving image received from the user terminal 2 in the storage unit 12.

[0035] Before or while the user views the moving image, the reception unit 131 receives from the user a designation of a character appearing in the moving image. The reception unit 131 causes the user terminal 2 to display a reception screen for receiving the designation of the character.

[0036] FIG. 3 is a schematic diagram of the user terminal 2 displaying the reception screen. The reception screen includes, for example, an area for receiving the designation of a character. The character is, for example, a person or an animal imaged as a subject in a live-action video. Further, the character may be, for example, a virtual character depicted in an animated video.

[0037] For example, the reception unit 131 acquires information associating the moving image with a plurality of characters appearing in the moving image, which is stored in advance in the storage unit 12. The reception unit 131 displays, on the reception screen, the plurality of characters associated with the moving image that the user plans to view or is viewing in the acquired information. The user selects, from among the plurality of characters displayed on the reception screen, the character for which a response is desired to be received on the delivery screen described later.

[0038] The user performs an operation for specifying a character on the user terminal 2. The user terminal 2 transmits information indicating the character specified by the user to the information processing apparatus 1. In the information processing apparatus 1, the reception unit 131 stores the information indicating the character received from the user terminal 2 in the storage unit 12.

[0039] Although the reception unit 131 receives the specification of the character on a screen different from the distribution screen in the example of FIG. 3, the reception of the specification of the character on the distribution screen may be accepted.

[0040] The moving image specified by the user is pre-stored in the storage unit 12 of the information processing apparatus 1 or in another storage unit that stores the moving image. The information processing apparatus 1 or another storage unit starts the distribution of the moving image by transmitting the moving image indicated by the information received from the user terminal 2 to the user terminal 2. The user terminal 2 displays a distribution screen A including the moving image received from the information processing apparatus 1 or another storage unit on the display unit.

[0041] FIG. 4 is a schematic diagram of the user terminal 2 displaying the distribution screen A. The distribution screen A includes a first scene display area A1 and a second scene display area A2. The first scene display area A1 is a display area having a larger area than the second scene display area A2.

[0042] The moving image is a live broadcast or live distribution (live streaming) moving image, or a pre-recorded moving image. The moving image is, for example, an image of a concert, musical, stage, drama, movie, TV program, press conference, etc. Further, the moving image may be a two-dimensional or three-dimensional animated video.

[0043] In this embodiment, the moving image includes a plurality of scenes that can be selected by the user as the main viewing target according to the user's operation. The moving image includes, for example, an overall scene that images or depicts a predetermined range, and one or more individual scenes that image or depict a range narrower than the overall scene. The overall scene is, for example, a moving image of a scene including a plurality of characters. On the other hand, an individual scene is, for example, a moving image of a scene including any one of the plurality of characters. Each of the one or more individual scenes is pre-associated with the character included in the individual scene.

[0044] An individual scene is generated, for example, by cutting out a part of the overall scene. In this case, the individual scene is generated, for example, by performing known image recognition processing on the overall scene to automatically cut out the area where each of the plurality of characters appears. Also, the individual scene may be generated by a human visually recognizing the overall scene and performing an operation of cutting out the area where each of the plurality of characters appears.

[0045] Also, an individual scene may be generated by imaging each of the plurality of characters with a camera different from the camera that imaged the overall scene. Also, the individual scene may be created as an independent scene rather than a part of the overall scene.

[0046] Each time point of the overall scene and each time point of the individual scene are pre-associated. When the moving image is played back, either the overall scene or the individual scene is selected as the main viewing target at each time point of the moving image. For example, in response to the user performing an operation of selecting either the overall scene or the individual scene as the main viewing target at an arbitrary time point while viewing the moving image, the user terminal 2 starts displaying the selected one in the first scene display area A1 and the unselected one in the second scene display area A2 on the distribution screen A from that time point.

[0047] In addition, when a user performs an operation to select either the overall scene or an individual scene as the main viewing target at any point while watching a video, the user terminal 2 may start displaying the selected scene from that point on without displaying the unselected scene.

[0048] Furthermore, the distribution screen A includes an input area A3, a comment area A4, an avatar area A5, and a response area A6. The input area A3 is an area that accepts input of input text such as comments by a user watching the video. The comment area A4 is an area in which an input text input by a user watching the video is displayed. In addition to an input text input by a user using the user terminal 2, the comment area A4 may also display an input text input by a user other than the user. The avatar area A5 is an area in which an avatar, which will be described later, is displayed. The response area A6 is an area in which a response, which will be described later, is displayed.

[0049] The user performs an operation on the user terminal 2 to input an input sentence into the input area A3. The user terminal 2 transmits the input sentence input by the user watching the video to the information processing device 1. In the information processing device 1, the receiving unit 131 stores the input sentence received from the user terminal 2 in the storage unit 12.

[0050] The identification unit 132 identifies the content of the video being viewed by the user. The identification unit 132 identifies the content of the video, for example, at a predetermined time interval or when the reception unit 131 receives an input sentence from the user.

[0051] 5 is a schematic diagram for explaining the process of the identification unit 132 identifying the contents of the video. The contents of the video include, for example, at least one of an object appearing in the video and a situation in the video. The object appearing in the video is, for example, a captured or depicted character or other object. The situation in the video includes, for example, the action, expression, clothing, etc. of the captured or depicted character or other object. The situation in the video may also include the date, time, location, weather, etc. in the video.

[0052] The specifying unit 132 specifies the content of the moving image by, for example, executing known image recognition processing on the moving image being viewed by the user. Further, the specifying unit 132 may acquire the content of the moving image being viewed by the user from a database in which the moving image and the content of the moving image are associated in advance.

[0053] The specifying unit 132 may specify the sound reproduced together with the moving image as the content of the moving image. In this case, the specifying unit 132 specifies the sound reproduced together with the moving image by, for example, executing known sound recognition processing on the moving image being viewed by the user.

[0054] The specifying unit 132 specifies the section in which the character designated by the user appears in the moving image, specifies the content of the moving image in the specified section, and does not have to specify the content of the moving image outside the section. Thereby, the information processing system S can make a response in consideration of only the content of the portion related to the character designated by the user in the moving image.

[0055] Next, the specifying unit 132 specifies the viewing status of the user's moving image. The viewing status of the moving image includes, for example, the date and time when the user is viewing the moving image, the scene selected by the user as the main viewing target among the plurality of scenes included in the moving image, and the like. Further, the specifying unit 132 may specify the viewing status of the moving images of a plurality of users who have viewed the moving image and who have specified the character specified by the user viewing the moving image (the user viewing the distribution screen A in FIG. 4).

[0056] The storage unit 12 of the information processing apparatus 1 or another storage unit that stores the machine learning model stores in advance a plurality of machine learning models corresponding to a plurality of characters. The machine learning model is configured to output a response corresponding to the input command sentence when the command sentence is input.

[0057] The machine learning model is pre-generated, for example, by performing machine learning on at least one of the attributes of each of a plurality of characters and the past utterances of the character using a known machine learning process (such as DNN (Deep Neural Network)) as learning data, and is stored in the storage unit 12 of the information processing apparatus 1 or another storage unit in association with the character.

[0058] The attributes of the character are, for example, the character's age, gender, place of residence, place of origin, experience, hobbies, preferences, etc. The past utterances of the character are, for example, the content that the character has uttered in the past on SNS (Social Networking Service), chat, articles, TV programs, radio programs, etc. Thus, the machine learning model can generate a response according to the characteristics of each character.

[0059] Also, the machine learning model may be configured to output a response using character information about the character. In this case, the storage unit 12 of the information processing apparatus 1 or another storage unit stores in advance in association each of a plurality of characters and the character information about the character. The character information includes, for example, any information about the character other than the learning data, such as the appearance information or release information of the character. Thus, the machine learning model can generate a response by referring to character information other than the learning data.

[0060] The determination unit 133 determines a first command sentence for causing the machine learning model to generate a response to the input sentence input by the user who is viewing the moving image and a second command sentence for causing the machine learning model to generate a response to the content of the moving image, triggered by the satisfaction of a predetermined response condition. The first command sentence is a command sentence for causing the machine learning model to generate a response including content (such as an answer to a question) along with the input sentence input by the user. The second command sentence is a command sentence for causing the machine learning model to generate a response including content (such as additional information about the subject shown in the moving image) along with the content of the video.

[0061] Figs. 6(a) and 6(b) are schematic diagrams for explaining the process in which the determination unit 133 determines the first instruction sentence and the second instruction sentence. The response condition is a condition indicating an opportunity to respond to a user who is viewing a moving image. The response condition is, for example, that the reception unit 131 has received an input sentence input by the user. Further, the response condition may be that in the content of the moving image specified by the specifying unit 132, the date and time in the moving image is a predetermined date and time or time zone, the weather in the moving image is a predetermined weather, the location in the moving image is a predetermined location, predetermined characters in the moving image have communicated with each other, the music flowing in the moving image is a predetermined name or a predetermined provider, and the like.

[0062] Further, the response condition may be that in the viewing status of one or a plurality of users specified by the specifying unit 132, the date and time when the user is viewing the moving image is a predetermined date and time or time zone, the number of views of the moving image by a plurality of users is more or less than a predetermined reference value, and the like. The response condition is not limited to the specific conditions shown here, and may include other conditions related to the moving image or the user.

[0063] The determination unit 133, taking the fulfillment of the response condition as an opportunity, acquires rules for determining instruction sentences, which are pre-stored in the storage unit 12. The rules for determining instruction sentences are, for example, templates for the first instruction sentence and the second instruction sentence respectively.

[0064] The determination unit 133 determines, for example, a first instruction sentence for causing a machine learning model to generate a response to the input sentence by applying the input sentence acquired by the reception unit 131 to the template of the first instruction sentence. The determination unit 133 determines, for example, a second instruction sentence for causing a machine learning model to generate a response to the content of the moving image by applying the content of the moving image specified by the specifying unit 132 to the template of the second instruction sentence. Thereby, the information processing system S can automatically output information related to the input sentence and information related to the content of the video using the machine learning model.

[0065] Further, in addition to the content of the moving image specified by the specifying unit 132, the determining unit 133 may determine a second command sentence based on the scene that the user is viewing among the plurality of scenes included in the moving image. In this case, for example, the determining unit 133 determines different second command sentences according to the scene that the user is viewing among the plurality of scenes included in the moving image. Thereby, the information processing system S can automatically output a response corresponding to the scene that the user is viewing, such as a comment that guides a user viewing the overall scene to an individual scene, or a comment that guides a user viewing an individual scene to the overall scene.

[0066] Further, in addition to the content of the moving image specified by the specifying unit 132, the determining unit 133 may determine a second command sentence based on the viewing status of the plurality of users specified by the specifying unit 132. In this case, for example, the determining unit 133 calculates values indicating the viewing status of the plurality of users (the number of views per scene, the number of comments per time, etc.), and determines different second command sentences according to the calculated values. Thereby, the information processing system S can automatically output a response corresponding to the viewing status of the plurality of users, such as a comment that guides to a scene with a high ratio of views by the plurality of users among the plurality of scenes included in the moving image, or a comment corresponding to the fact that the number of input sentences input by the plurality of users is large (being lively).

[0067] In the example of FIG. 6(a), the determining unit 133 determines a first command sentence for causing the machine learning model to generate a response to the input sentence and a second command sentence for causing the machine learning model to generate a response to the content of the moving image, respectively. The input unit 134 described later inputs the first command sentence and the second command sentence to the machine learning model at different timings. Thereby, the information processing system S can independently output a response to the user's input sentence and a response to the content of the video.

[0068] In the example of FIG. 6(b), the determination unit 133 determines a single instruction sentence that is the first instruction sentence for causing the machine learning model to generate a response to the input sentence and is also the second instruction sentence for causing the machine learning model to generate a response to the content of the moving image. The input unit 134 described later inputs the single instruction sentence that is the first instruction sentence and also the second instruction sentence into the machine learning model. Thereby, the information processing system S can output a response considering the content of the video with respect to the input sentence of the user.

[0069] The determination unit 133 may execute both the process of determining each of the first instruction sentence and the second instruction sentence illustrated in FIG. 6(a) and the process of determining a single instruction sentence that is the first instruction sentence and also the second instruction sentence illustrated in FIG. 6(b).

[0070] The input unit 134 inputs the first instruction sentence and the second instruction sentence determined by the determination unit 133 into the machine learning model. The input unit 134 selects, for example, a machine learning model corresponding to the character specified by the user received by the reception unit 131 from among a plurality of machine learning models preliminarily stored in the storage unit 12 of the information processing apparatus 1 or other storage units. The input unit 134 inputs the first instruction sentence and the second instruction sentence into the selected machine learning model. Further, the input unit 134 may input a single instruction sentence that is the first instruction sentence and also the second instruction sentence into the selected machine learning model.

[0071] The output unit 135 acquires the response output by the machine learning model into which the first instruction sentence and the second instruction sentence are input, and outputs the acquired response to the user terminal 2 associated with the user.

[0072] FIGS. 7(a) and 7(b) are schematic diagrams for explaining the process in which the output unit 135 outputs a response. The output unit 135, for example, displays an avatar imitating the character specified by the user in the avatar area A5 on the distribution screen A of FIG. 4. The avatar is a pre-generated still image or moving image. Also, the avatar may be dynamically generated. Information for displaying the avatar is, for example, preliminarily stored in the storage unit 12 in association with the character.

[0073] The output unit 135 causes the response area A6 to display the response output by the machine learning model for the first command sentence and the second command sentence, in association with, for example, the avatar displayed in the avatar area A5. Thereby, the information processing system S can give the user a feeling of actually conversing with the character designated by the user.

[0074] In the example of FIG. 7(a), the output unit 135 outputs different responses for each user to a plurality of user terminals 2 associated with a plurality of users who have watched a moving image and designated the same character. That is, the response output according to the command sentence for a certain user's input sentence is displayed only on the user terminal 2 of that user, and is not displayed on the user terminals 2 of other users other than that user. Thereby, the information processing system S can provide the response to the user's input sentence only to that user and protect the privacy between users.

[0075] In the example of FIG. 7(b), the output unit 135 outputs the same response to a plurality of user terminals 2 associated with a plurality of users who have watched a moving image and designated the same character. That is, the response output according to the command sentence for a certain user's input sentence is displayed on the user terminal 2 of that user and the user terminals 2 of other users who have designated the same character as that user. Thereby, the information processing system S can also provide the response to the user's input sentence to other users and promote conversation among a plurality of users.

[0076] Further, when the response output by the machine learning model includes information on a predetermined product, the output unit 135 may output access information (for example, URI: Uniform Resource Identifier) for accessing a website (such as a product sales page) related to the product to the user terminal 2. In this case, the storage unit 12 stores in advance, in association with each of the plurality of products, the access information related to the product. Thereby, the information processing system S can make it easier for the user to access the information on the product related to the response output by the machine learning model.

[0077] When a plurality of users who are watching a moving image input input sentences such as comments on the moving image, negative interactions may occur among the plurality of users. Negative interactions include, for example, slander, discriminatory words, obscene words, and the like. Therefore, the information processing system S may output a response for arbitrating negative interactions that occur among a plurality of users.

[0078] The determination unit 133 determines whether negative interactions are occurring among a plurality of users based on the input sentences input by the plurality of users who are watching the moving image. For example, the determination unit 133 determines that negative interactions are occurring on the condition that any of the keywords of a plurality of negative contents indicated by the information pre-stored in the storage unit 12 is included in the input sentences received by the reception unit 131 from the plurality of users. Further, for example, the determination unit 133 may determine whether negative interactions are occurring by inputting the input sentences into a machine learning model generated in advance by machine learning using negative interactions as learning data.

[0079] The determination unit 133 determines a third command sentence for causing the machine learning model to generate a response for arbitrating negative interactions on the condition that it is determined that negative interactions are occurring.

[0080] FIG. 8 is a schematic diagram for explaining the process in which the determination unit 133 determines the third command sentence. The determination unit 133 acquires, on the condition that it is determined that negative interactions are occurring, the rules for determining command sentences pre-stored in the storage unit 12. The rules for determining command sentences are, for example, the templates of the third command sentence. The third command sentence is a command sentence for causing the machine learning model to generate a response including content for arbitrating negative interactions.

[0081] The determination unit 133 determines a third instruction sentence for causing the machine learning model to generate a response to the input sentence, for example, by applying the input sentence determined to have a negative interaction to the template of the third instruction sentence.

[0082] The input unit 134 inputs the third instruction sentence determined by the determination unit 133 into the machine learning model. For example, the input unit 134 selects a machine learning model corresponding to the character specified by the user received by the reception unit 131 from a plurality of machine learning models pre-stored in the storage unit 12 of the information processing apparatus 1 or other storage units. The input unit 134 inputs the third instruction sentence into the selected machine learning model.

[0083] The output unit 135 acquires the response output by the machine learning model into which the third instruction sentence is input, and outputs the acquired response to the user terminal 2 associated with each of the plurality of users. Thereby, the information processing system S can automatically guide the user not to perform a negative interaction by using the machine learning model corresponding to the character specified by the user.

[0084] Furthermore, the output unit 135 may not output the input sentence input by the user who is determined to have a negative interaction to the input sentence input after the determination to the user terminal 2 associated with another user different from the user. Thereby, the information processing system S can isolate the user who has performed a negative interaction from other users and make it difficult for problems to occur between users.

[0085] [Flow of Information Processing Method] FIG. 9 is a diagram showing a flowchart of an exemplary information processing method executed by the information processing system S according to the present embodiment. The reception unit 131 receives a designation of a character appearing in the moving image from the user before or while the user is viewing the moving image (S11). The reception unit 131 receives the input sentence input by the user viewing the moving image (S12).

[0086] The specific unit 132 identifies the content of the moving image that the user is viewing (S13). The specific unit 132 identifies the viewing status of the moving image by the user (S14).

[0087] The information processing apparatus 1 determines whether a predetermined response condition, which is a condition indicating an opportunity to respond to the user viewing the moving image, is satisfied. When the response condition is not satisfied (NO in S15), the information processing apparatus 1 returns to step S12 and continues the processing.

[0088] When the response condition is satisfied (YES in S15), the determination unit 133 determines a first instruction sentence for causing the machine learning model to generate a response to the input sentence input by the user viewing the moving image, and a second instruction sentence for causing the machine learning model to generate a response to the content of the moving image (S16).

[0089] The input unit 134 selects a machine learning model corresponding to the character specified by the user received by the reception unit 131 from a plurality of machine learning models stored in advance in the storage unit 12 of the information processing apparatus 1 or other storage units. The input unit 134 inputs the first instruction sentence and the second instruction sentence determined by the determination unit 133 into the selected machine learning model (S17).

[0090] The output unit 135 acquires the response output by the machine learning model into which the first instruction sentence and the second instruction sentence are input, and outputs the acquired response to the user terminal 2 associated with the user (S18).

[0091] The information processing apparatus 1 determines whether a predetermined end condition (such as the user having performed an end operation), which is a condition for ending the processing, is satisfied. When the end condition is not satisfied (NO in S19), the information processing apparatus 1 returns to step S12 and continues the processing. When the end condition is satisfied (YES in S19), the information processing apparatus 1 ends the processing.

[0092] Note that according to the present invention, it becomes possible to contribute to Goal 9, "Build the infrastructure for industry and technological innovation," of the Sustainable Development Goals (SDGs) led by the United Nations.

[0093] As described above, the present invention has been described using embodiments. However, the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist. For example, all or part of the device can be configured by functionally or physically dispersing and integrating it in any unit. Also, new embodiments resulting from any combination of a plurality of embodiments are included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination have the effects of the original embodiments combined.

Explanation of Reference Numerals

[0094] S Information processing system 1 Information processing device 11 Communication unit 12 Storage unit 13 Control unit 131 Reception unit 132 Identification unit 133 Decision unit 134 Input unit 135 Output unit 2 User terminal

Claims

1. a reception unit that receives, from a user, a designation of a character appearing in a video; An identification unit that identifies the content of the video being viewed by the user; A determination unit that determines a first command statement for causing a machine learning model to generate a response to an input sentence input by a user watching the video, and a second command statement for causing the machine learning model to generate a response to content of the video; an input unit that inputs the first command statement and the second command statement to the machine learning model; an output unit that outputs a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; having the identification unit identifies a section in the video in which the character appears by image recognition processing, and identifies the content of the video in the identified section; the determination unit determines the second command statement for causing the machine learning model to generate a response that reflects a content of the video in the specified section and does not reflect a content of the video outside the specified section. Information processing device.

2. the input unit inputs the first command statement and the second command statement to the machine learning model corresponding to the character, the machine learning model being selected from a plurality of machine learning models pre-stored in a storage unit; The information processing device according to claim 1 .

3. The machine learning model corresponding to the character is generated in advance by machine learning at least one of the attributes of the character and past utterances of the character.

3. The information processing device according to claim 1 or 2.

4. The output unit displays an avatar that resembles the character on the information terminal, and outputs a response output by the machine learning model to the first command statement and the second command statement in association with the avatar.

3. The information processing device according to claim 1 or 2.

5. the storage unit stores in advance each of the plurality of characters in association with character information relating to the character; The machine learning model corresponding to the character is configured to output a response using the character information associated with the character in the storage unit. The information processing device according to claim 2 .

6. The identification unit identifies viewing situations of the video of a plurality of users who have viewed the video and who have designated the character, the determination unit determines the second command statement based on the viewing states of the plurality of users.

3. The information processing device according to claim 1 or 2.

7. the video includes a plurality of scenes that the user can select as a viewing target in response to an operation by the user, the determination unit determines the second command statement based on a scene currently being viewed by the user among the plurality of scenes.

3. The information processing device according to claim 1 or 2.

8. the output unit outputs the same response to the plurality of information terminals associated with the plurality of users who have viewed the video and designated the character.

3. The information processing device according to claim 1 or 2.

9. the determination unit determines a single command statement, the single command statement being the first command statement for causing the machine learning model to generate a response to the input sentence and a second command statement for causing the machine learning model to generate a response to content of the video; 3. The information processing device according to claim 1 or 2.

10. the specifying unit specifies, as the content of the moving image, an image to be displayed as the moving image and a sound to be played together with the moving image; 3. The information processing device according to claim 1 or 2.

11. the determination unit determines whether or not a negative exchange is occurring between the plurality of users who are watching the video based on the input sentences input by the plurality of users, and, on a condition that it is determined that the negative exchange is occurring, determines a third command statement for causing the machine learning model to generate a response that mediates the negative exchange; The input unit inputs the third command statement to the machine learning model; The output unit outputs a response output by the machine learning model in response to the third command statement to the information terminal associated with each of the multiple users.

3. The information processing device according to claim 1 or 2.

12. The processor executes A step of accepting, from a user, a designation of a character appearing in a moving image; identifying the content of the video being viewed by the user; determining a first command statement for causing a machine learning model to generate a response to an input sentence input by a user watching the video, and a second command statement for causing the machine learning model to generate a response to content of the video; inputting the first statement and the second statement into the machine learning model; outputting a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; having In the step of specifying, a section in the video in which the character appears is specified by image recognition processing, and a content of the video in the specified section is specified; determining the second command statement for causing the machine learning model to generate a response that reflects the content of the video in the specified section and does not reflect the content of the video outside the specified section, in the determining step; Information processing methods.

13. The processor: A step of accepting, from a user, a designation of a character appearing in a moving image; identifying the content of the video being viewed by the user; determining a first command statement for causing a machine learning model to generate a response to an input sentence input by a user watching the video, and a second command statement for causing the machine learning model to generate a response to content of the video; inputting the first statement and the second statement into the machine learning model; outputting a response output by the machine learning model to the first command statement and the second command statement to an information terminal associated with the user; Run the command, In the step of specifying, a section in the video in which the character appears is specified by image recognition processing, and a content of the video in the specified section is specified; determining the second command statement for causing the machine learning model to generate a response that reflects the content of the video in the specified section and does not reflect the content of the video outside the specified section, in the determining step; program.

Citation Information

Patent Citations

  • Answer selection device, model learning device, method for selecting answer, method for learning model, and program

    JP2019192073A

  • Interactive program, device, and method for expressing sense of listening of character in accordance with user's emotion

    JP2022054326A

  • Persona chatbot control method and system

    JP2022180282A

  • Automatic captioning of audible portions of content on a computing device - Patent Application 20070122997

    JP2022530201A

  • Information distribution apparatus, information distribution method, and program

    JP2023050587A