Information processing device, information processing method, and information processing system

The information processing device automatically generates tailored commentary for video content using a virtual commentator, addressing the high cost and labor of manual comment addition, enabling easy and personalized video creation.

JP7797501B2Active Publication Date: 2026-01-13SONY GROUP CORP +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023523954
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-24
Filing Date
2021-12-27
Publication Date
2026-01-13
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

The manual process of adding comments to video content, such as sports and educational content, is costly and labor-intensive, making it difficult for individuals to easily create video content with comments.

Method used

An information processing device and method that automatically generates comments based on the relationship between the content poster and viewer, using a virtual commentator to provide tailored commentary depending on the commentator's position and target audience.

Benefits of technology

Enables the easy creation of video content with personalized commentary, reducing costs and labor, and allowing for targeted commentary styles suited to specific viewers or groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007797501000001
    Figure 0007797501000001
  • Figure 0007797501000002
    Figure 0007797501000002
  • Figure 0007797501000003
    Figure 0007797501000003
Patent Text Reader

Abstract

The present invention more easily creates a video content with a comment attached thereto. An information processing device according to an embodiment of the present invention is provided with: an acquisition unit (10, 110, 120) that acquires information concerning a relationship between a provider of a content and a viewer of the content; and a comment generation unit (40) that generates, on the basis of the information concerning the relationship, a comment to be made by a virtual commentator.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, an information processing method, and an information processing system. [Background technology]

[0002] In recent years, with the emergence of video distribution services such as YouTube (registered trademark), there has been an increase in the distribution of video content such as sports, games, and education with commentary and commentary. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2018-187712 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the editing process of adding comments to video content in the form of audio or text is generally done manually by content creators, and the costs of installing the equipment required for editing and the labor required for editing are high, so it has not been possible for anyone to easily create video content with comments.

[0005] Therefore, the present disclosure proposes an information processing device, an information processing method, and a program that enable video content with comments to be created more easily. [Means for solving the problem]

[0006] In order to solve the above problem, one form of information processing device according to the present disclosure includes an acquisition unit that acquires information regarding the relationship between a content poster and a viewer of the content, and a comment generation unit that generates comments to be uttered by a virtual commentator based on the information regarding the relationship. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 10 is a diagram for explaining changes in comments depending on the position of a virtual commentator according to an embodiment of the present disclosure. [Figure 2A] FIG. 10 is a diagram illustrating an example of a selection screen for video content according to an embodiment of the present disclosure. [Figure 2B] FIG. 10 is a diagram illustrating an example of a selection screen for a virtual commentator according to an embodiment of the present disclosure. [Figure 2C] FIG. 10 is a diagram illustrating an example of a selection screen for selecting the position of a virtual commentator according to an embodiment of the present disclosure. [Figure 2D] FIG. 10 is a diagram illustrating an example of a selection screen for selecting listeners of comments made by a virtual commentator according to an embodiment of the present disclosure. [Figure 2E] FIG. 10 is a diagram illustrating an example of a playback screen of video content with comments according to an embodiment of the present disclosure. [Figure 3A] FIG. 10 is a diagram illustrating an example of adding a comment when posting (uploading) video content according to an embodiment of the present disclosure. [Figure 3B] FIG. 10 is a diagram showing an example of adding a comment when viewing (downloading) video content according to an embodiment of the present disclosure. [Figure 4] FIG. 1 is a block diagram illustrating an example of a system configuration when adding a comment when posting (uploading) video content according to an embodiment of the present disclosure. [Figure 5A] 10 is a flowchart illustrating an example of an operation from the start of an app to setting a virtual commentator according to an embodiment of the present disclosure. [Figure 5B] FIG. 10 is a diagram illustrating an example of a video content management table for managing video content according to an embodiment of the present disclosure. [Figure 5C] FIG. 10 is a diagram showing an example of a character management table for managing virtual commentator characters according to an embodiment of the present disclosure. [Figure 5D] FIG. 10 is a diagram illustrating an example of a position management table for managing the positions of virtual commentators according to an embodiment of the present disclosure. [Figure 5E] FIG. 10 is a diagram illustrating an example of a comment target management table that manages targets (comment targets) on which a virtual commentator makes a comment according to an embodiment of the present disclosure. [Figure 6A] 10 is a flowchart illustrating an example of an event extraction operation according to an embodiment of the present disclosure. [Figure 6B] FIG. 10 is a diagram illustrating an example of a model management table for managing recognition models according to an embodiment of the present disclosure. [Figure 6C] FIG. 10 is a diagram illustrating an example of an event management table for managing events according to an embodiment of the present disclosure. [Figure 6D] FIG. 10 is a diagram illustrating an example of a list of event data according to an embodiment of the present disclosure. [Figure 7A] 10 is a flowchart illustrating an example of a comment generation operation according to an embodiment of the present disclosure. [Figure 7B] 10A and 10B are diagrams illustrating an example of a position comment list and an example of a usage comment history according to an embodiment of the present disclosure. [Figure 7C] FIG. 10 is a diagram illustrating an example of a target comment list according to an embodiment of the present disclosure. [Figure 8A] 10 is a flowchart illustrating an example of a suffix conversion operation according to a modified example of an embodiment of the present disclosure. [Figure 8B] FIG. 10 is a diagram illustrating an example of a hierarchical relationship management table according to an embodiment of the present disclosure. [Figure 9A] 10 is a flowchart illustrating an example of a generating operation and a distributing operation of video content with comments according to an embodiment of the present disclosure. [Figure 9B] 10 is a flowchart illustrating an example of an operational flow of an avatar generation process according to an embodiment of the present disclosure. [Figure 9C] 10 is a flowchart showing an example of an operation flow of an editing and rendering process according to an embodiment of the present disclosure. [Figure 10A] 10A and 10B are diagrams illustrating an example in which a comment target is changed in video content according to an embodiment of the present disclosure. [Figure 10B]10B is a diagram showing an example of a comment made by a virtual commentator at each time in FIG. 10A. FIG. [Figure 11A] 10 is a flowchart illustrating an example of a comment generation operation according to an embodiment of the present disclosure. [Figure 11B] 11B is a flowchart showing a more detailed operation flow of the standing position / target adjustment operation shown in step S220 of FIG. 11A. [Figure 12A] FIG. 10 is a diagram illustrating an example of generating a comment based on an emotion value according to an embodiment of the present disclosure. [Figure 12B] 10 is a table showing the amount of change in emotion value for each standing position for each event according to an embodiment of the present disclosure. [Figure 12C] FIG. 12B is a diagram for explaining the change in emotion value shown in FIG. 12A in association with an event according to an embodiment of the present disclosure. [Figure 13] 10 is a flowchart illustrating an example of an operational flow for generating a comment when the position of a virtual commentator according to an embodiment of the present disclosure is "friend" and the comment target is "player." [Figure 14] FIG. 1 is a block diagram illustrating an example of a system configuration in which comments are added when video content is viewed (downloaded) according to an embodiment of the present disclosure. [Figure 15A] 10 is a flowchart illustrating another example of the operation from the start of an app to setting a virtual commentator according to an embodiment of the present disclosure. [Figure 15B] FIG. 10 is a diagram illustrating an example of a character management table according to an embodiment of the present disclosure. [Figure 16A] FIG. 10 is a block diagram illustrating an example of a system configuration when generating dialogue comments by two virtual commentators for real-time video distribution according to an embodiment of the present disclosure. [Figure 16B] 10 is a flowchart illustrating an example of an operation when generating a dialogue comment by two virtual commentators for a real-time video distribution according to an embodiment of the present disclosure. [Figure 16C]FIG. 10 is a diagram showing an example of an event interval when two virtual commentators generate dialogue comments for a real-time video distribution according to an embodiment of the present disclosure. [Figure 16D] 10 is a flowchart illustrating an example of an operational flow of a comment generation process according to an embodiment of the present disclosure. [Figure 16E] 10 is a flowchart illustrating an example of an operation flow of a speech control process according to an embodiment of the present disclosure. [Figure 16F] 10 is a flowchart showing an example of an operation flow of an editing and rendering process according to an embodiment of the present disclosure. [Figure 17A] 10 is a flowchart illustrating an example of an operation when a comment is generated in response to viewer feedback during real-time video distribution according to an embodiment of the present disclosure. [Figure 17B] 10 is a flowchart illustrating an example of an operational flow of a viewer feedback process according to an embodiment of the present disclosure. [Figure 18] 10 is a flowchart illustrating an example of an operation flow when the number of virtual commentators increases or decreases during live distribution according to an embodiment of the present disclosure. [Figure 19A] FIG. 10 is a diagram illustrating an example of a virtual position when a virtual commentator according to an embodiment of the present disclosure is standing in the position of a "friend watching the game together" and speaking to video content. [Figure 19B] FIG. 10 is a diagram illustrating a case where a virtual commentator according to an embodiment of the present disclosure is positioned as a "friend watching the game together" and is speaking to viewers. [Figure 19C] FIG. 10 is a diagram illustrating a case where a virtual commentator according to an embodiment of the present disclosure is in the position of a "commentator" and is speaking to viewers. [Figure 20A] FIG. 10 is a diagram showing an example in which a caption of an audio comment is placed in a basic position (for example, below the center of the screen) according to an embodiment of the present disclosure. [Figure 20B] 10A and 10B are diagrams illustrating an example in which the position of a caption of an audio comment is adjusted based on gaze information of a poster according to an embodiment of the present disclosure. [Figure 20C]10A and 10B are diagrams illustrating an example in which a standing position comment is generated based on viewer line-of-sight information according to an embodiment of the present disclosure, and the display position of the caption of the audio comment is adjusted. [Figure 21] 10 is a block diagram showing an example of a system configuration when machine learning according to an embodiment of the present disclosure is applied to generate comments for each position of a virtual commentator. [Figure 22] FIG. 1 is a hardware configuration diagram illustrating an example of an information processing device that executes various processes according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, an embodiment of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0009] The present disclosure will be described in the following order. 1. One embodiment 1.1 Comments based on the virtual commentator's position and target audience 1.2 Example of app screen 1.3 Examples of commenters 1.4 Example of system configuration for adding comments when posting (uploading) 1.5 Example of operation flow when adding a comment when posting (uploading) 1.5.1 Example of operation from app launch to virtual commentator setting 1.5.2 Event extraction operation example 1.5.3 Example of comment generation 1.5.3.1 Variations of inflection 1.5.4 Example of generating and distributing video content with comments 1.5.4.1 Example of avatar generation process 1.5.4.2 Editing and rendering process example 1.6 Example of dynamically changing position and comment target 1.6.1 Operational flow example 1.7 Example of generating comments based on sentiment values 1.7.1 Operational flow example 1.8 Example of system configuration for adding comments when viewing (downloading) 1.9 Example of operation flow when adding comments when viewing (downloading) 1.10 Example of generating interactive comments by two virtual commentators for real-time video streaming 1.10.1 Example of comment generation process 1.10.2 Speech Control Processing Example 1.10.3 Editing and rendering process example 1.11 Example of generating comments based on viewer feedback during real-time video streaming 1.11.1 Audience Feedback Examples 1.12 Example of virtual commentators increasing and decreasing during real-time video viewing 1.13 Example of adjusting the display position according to the virtual commentator's position 1.14 Example of comment rules when multiple virtual commentators exist 1.15 Example of commenting using gaze information 1.16 Examples of caption display positions 1.17 Example of applying machine learning to generate comments for each position of a virtual commentator 2. System configuration example 3. Hardware configuration

[0010] 1. One embodiment An information processing device, an information processing method, and an information processing system according to an embodiment of the present disclosure will be described in detail below with reference to the drawings. In this embodiment, optimal comments are automatically generated for video of sports, games, etc., depending on "where the comment is being made and to whom." A typical service for adding comments to video is live broadcasts of sports programs, such as baseball and soccer broadcasts, broadcast on television. However, with the recent emergence of internet-based video distribution services such as YouTube (registered trademark), this has been expanded to include video content such as game play-by-plays and product introductions. As a result, while traditional live television programs were broadcast to an unspecified number of viewers, internet-based video distribution services have given rise to styles of commentary that are targeted to specific viewers, such as posting and streaming to specific groups such as friends and family, or interactively responding to viewer chat during live streaming.

[0011] 1.1 Comments based on the virtual commentator's position and target audience Here, the meaning of "from what position to whom to make a comment" according to this embodiment will be explained using Fig. 1. Fig. 1 illustrates an example in which a virtual commentator C1 makes comments from various positions to video content G0 in which a player (George) U1 is playing a soccer game.

[0012] FIG. 1 (1) shows a case where a virtual commentator C1 makes comments as a player U1 playing a soccer game. FIG. 1 (2) shows a case where the virtual commentator C1 makes comments as a team member of player U1. Note that the team member may be another player playing the same soccer game via a network, or may be an NPC (Non-Player Character). FIG. 1 (3) shows a case where the virtual commentator C1 makes comments as one of the viewers (friends, etc.) watching the soccer game.

[0013] In (1), virtual commentator C1 expresses the enthusiasm of player U1 by saying, for example, "Let's give it a go!". In (2), virtual commentator C1, as a team member of player U1, says, "Let's run together!". In (3), virtual commentator C1, as a viewer, cheers on player U1 by saying, "Go for it! Turn the tables!". As you can see, the appropriate comments for a single video scene vary depending on the position of virtual commentator C1.

[0014] Furthermore, regarding (3), if the target of the comment is the player U1 in (3-1) and the viewer (friend) A1 in (3-2), then in (3-1) the commenter might encourage player U1 by calling his name, saying, "George, keep it up," but in (3-2) the commenter might say, "George, aren't you making a late decision?", leading to a frank conversation among friends.

[0015] Thus, unlike live broadcasts of television programs, when generating comments that take posters and viewers into consideration, it is necessary to clarify the position of the virtual commentator and the target of the comment in order to generate appropriate comments. Furthermore, the term "position" has two meanings: the first is the position where a person stands, and the second is the position or perspective that a person intends to take in human relationships or society. Therefore, this embodiment mainly illustrates a case where comments are generated according to the second position. However, the present invention is not limited to this, and it is also possible to configure the generated comments to change according to the position of the viewer or virtual commentator in the 3D virtual space, etc., with regard to the first position.

[0016] 1.2 Example of app screen Next, an example screen of an app that automatically adds comments to video content (hereinafter simply referred to as video) will be described. In this description, an app that allows a user to play a game while recording it on a smartphone or game console, and then share the gameplay video (video content) with friends along with comments will be exemplified. Note that the app may be an app executed on various information processing devices, such as a smartphone app, an app within a game console, or an app on a personal computer. Furthermore, the app that automatically adds comments to video content may be an app that only has a comment adding function, or may be implemented as one function of a game app, a recording app, a video playback app, a video distribution app, or an SNS (Social Networking Service) app that can handle video content.

[0017] 2A to 2E are diagrams showing example screens of an application that automatically adds comments to video content according to this embodiment. Note that in this example, the input device operated by the user is a touch panel arranged on the screen, but this is not limiting, and other input devices such as a mouse, keyboard, or touchpad may also be used.

[0018] FIG. 2A is a diagram showing an example of a selection screen for video content to which text is to be added. As shown in FIG. 2A, on video content selection screen G10, for example, thumbnails B11 to B13 of video content that are candidates for text addition are presented to the user. The video content to which text is to be added is selected by the user selecting one of thumbnails B11 to B13 displayed on video content selection screen G10. FIG. 2A illustrates an example in which a live video of a soccer game is selected as the video content. Note that, for example, if a video of a game that has been played is recorded, the video content to which text is to be added may be selected by displaying a message inquiring the user as to whether or not to distribute the content and whether or not to add a comment.

[0019] FIG. 2B is a diagram showing an example of a selection screen for a virtual commentator who will comment on the selected video content. As shown in FIG. 2B, the commentator selection screen G20 presents, for example, characters B21 to B23, who are candidates for the virtual commentator who will comment on the video content, to the user. FIG. 2B illustrates a girl character B21, a boy character B22, and a male character B23. Note that, instead of abstract characters such as a girl, a boy, a man, or a woman, more specific characters with personal names may be used. In this case, each virtual commentator can be given a more distinctive personality, and therefore, it is possible to increase the variety of virtual commentators and configure the system so that a virtual commentator that suits the user's preferences can be selected.

[0020] FIG. 2C is a diagram showing an example of a selection screen for selecting the position of the selected virtual commentator. As shown in FIG. 2C, the position selection screen G30 presents the user with, for example, the user (player) B31, team members B32, and friends B33. For example, the user selects "myself" if he / she wants the virtual commentator to comment on his / her behalf, selects "team members" if he / she wants the virtual commentator to comment as a team member playing the game together, or selects "friends" if he / she wants the virtual commentator to comment as a friend watching the game. Note that the position presented to the user on the position selection screen G30 is not limited to the above and may be changed as appropriate depending on the video content selected in FIG. 2A or the virtual commentator selected in FIG. 2B. For example, if educational video content is selected, the user may be given options for the virtual commentator's position, such as "myself (teacher or student)," "classmate," "teacher (if not a teacher)," or "parent."

[0021] FIG. 2D is a diagram showing an example of a selection screen for selecting a comment target, i.e., a listener (comment target) of a comment made by a virtual commentator. As shown in FIG. 2D, the comment target (listener) selection screen G40 presents, for example, the user (player) B41, team members B42, and friends B43 to the user. For example, the user selects "Self" if the user wants the virtual commentator of the selected position to comment on the user (player), selects "Team members" if the user wants the comment on a team member, or selects "Friends" if the user wants the comment on a friend (e.g., a conversation between friends). Note that, as with the position selection screen G30, the position presented to the user on the comment target (listener) selection screen G40 is not limited to the above and may be changed as appropriate depending on the video content selected in FIG. 2A or the virtual commentator selected in FIG. 2B.

[0022] FIG. 2E is a diagram showing an example of a playback screen of video content with comments automatically added to the video content selected in FIG. 2A based on the type, position, and comment target of the virtual commentator selected in FIGS. 2B to 2D. In this embodiment, the selected video content is analyzed, and comments from the selected virtual commentator for each scene in the video are automatically generated based on the selected position and comment target. For example, if a live video of a soccer game is selected, the game live video is analyzed to extract events such as players and ball positions in the game, types of techniques, and the like, and live commentary comments corresponding to each event are generated based on the selected character, their position, and comment target. Then, as in the video content with comments G50 shown in FIG. 2E, the generated live commentary comments T1 are superimposed on the game live video in the form of a voice or text (font) corresponding to the character of the virtual commentator C1. Note that, in addition to analyzing the video content, events may be extracted from the video content by utilizing an API (Application Programming Interface) if the game has a publicly available API.

[0023] 1.3 Examples of commenters Possible entities that may add comments to video content include, for example, the creator of the video content, the poster of the video content, the distributor of the video content, and the viewer of the video content. FIG. 3A is a diagram showing an example of a case where a creator of the video content, the poster of the video content, or the distributor of the video content adds a comment to video content, i.e., a case where a comment is added when the video content is posted (uploaded). FIG. 3B is a diagram showing an example of a case where each viewer of the video content (including the viewer himself / herself) individually adds a comment to the video content, i.e., a case where a comment is added when the video content is viewed (downloaded). In this description, the entity that adds a comment to video content refers to a person who decides to add a comment, including, for example, the position of a virtual commentator and the target of the comment.

[0024] Note that Figures 3A and 3B illustrate an example in which a comment is automatically added to video content G0 in cloud 100 managed by a video distribution service provider that is the distributor of the video content, but this is not limited to this. A comment may be automatically added to video content G0 in an information processing device (personal computer, smartphone, game console, etc.) of user side U100, who is the creator or poster of the video content, or a comment may be automatically added to video content in an information processing device (personal computer, smartphone, game console, etc.) of viewer side A100 of the video content.

[0025] 3A, when a video content creator, a video content contributor, or a video content contributor adds a comment to video content G0 using user terminal M1, the same video content creator and / or contributor (user U1 in this example) and video content viewers A11-A13 will watch video content G1 with the same comment added by the same virtual commentator. However, for example, if the added comment is text, the position and size at which the comment is displayed may differ depending on the viewing environment of each of video content viewers A11-A13.

[0026] On the other hand, as shown in FIG. 3B, when each viewer (including the viewer) U1, A11 to A13 of a video content individually adds a comment to a common video content G0, each viewer (including the viewer) U1, A11 to A13 selects a virtual commentator and selects their position and comment target. In this case, the comments added to the video content G1 to G4 viewed by each viewer (including the viewer) U1, A11 to A13 are not necessarily the same. In this way, when the viewer individually adds a comment, each viewer can select a virtual commentator, their position, and comment target that suit their own preferences and situation. This allows each viewer to watch the video content with comments that suit their own preferences and situation.

[0027] In addition, the video content G1 to which comments have been added by the creator of the video content, the poster of the video content, or the distributor of the video content may be configured so that viewers (including the viewers themselves) U1, A11 to A13 can add additional comments based on the virtual commentator they selected when watching, their position, and the subject of the comment.

[0028] 1.4 Example of system configuration for adding comments when posting (uploading) First, the case where a comment is added when video content is posted (uploaded) (see FIG. 3A) will be described below. FIG. 4 is a block diagram showing an example of a system configuration when a comment is added when video content is posted (uploaded). Note that this description illustrates the case where an audio comment is added to video content.

[0029] As shown in Figure 4, the system configuration of the information processing system 1 when adding a comment at the time of posting (uploading) includes a user terminal 10, a user information storage unit 20, an event extraction unit 30, a comment generation unit 40, an utterance control unit 50, an avatar generation unit 60, an editing / rendering unit 70, and a distribution unit 80.

[0030] The user terminal 10 is an information processing device on the user U1 side, such as a smartphone, a personal computer, or a game console, and executes applications such as generating (recording, etc.) video content and automatically adding comments.

[0031] The user information storage unit 20 is, for example, a database managed by a video distribution service provider, and stores information about the user U1 himself, information about other users (including viewers) related to the user U1, and information about the relationship between the user U1 and other users (viewers, etc.). In this description, this information is collectively referred to as user information.

[0032] Here, the information regarding the relationship between user U1 and other users (viewers, etc.) may include, for example, at least one of the following: the degree of intimacy between user U1 who is the poster and other users (viewers, etc.) (such as account follow relationships), the relationship between user U1 and other users (viewers, etc.) in the video content, and history information of other users (viewers, etc.) regarding video content posted by user U1 in the past. However, the information is not limited to this and may include various information regarding the relationship between user U1 and other users (viewers, etc.).

[0033] The event extraction unit 30 includes, for example, an image analysis unit 31 and an audio analysis unit 32, and extracts events from the video content by analyzing the images and audio of the video content generated by the user terminal 10. The extraction of events will be described in detail later.

[0034] The comment generation unit 40 includes a position / target control unit 41, and generates a comment (text data) for the event extracted by the event extraction unit 30 based on the virtual commentator selected by the user terminal 10, their position, and the comment target. That is, the comment generation unit 40 generates a comment to be uttered by the virtual commentator based on information regarding the relationship between user U1 and other users (viewers, etc.). The generation of comments will be described in detail later.

[0035] The speech control unit 50 converts the comment (text data) generated by the comment generating unit 40 into voice data, for example, by using TTS (Text-To-Speech).

[0036] The avatar generation unit 60 generates an avatar of the selected virtual commentator based on the virtual commentator selected on the user terminal 10, for example, and then generates movement data for moving the avatar based on the position and comment target of the selected virtual commentator and the voice data generated by the speech control unit 50.

[0037] The editing and rendering unit 70 renders the moving avatar generated by the avatar generation unit 60 and superimposes the resulting video data (hereinafter also referred to as avatar animation) on the video data of the video content. The editing and rendering unit 70 also superimposes the audio data generated by the speech control unit 50 on the audio data of the video content. The editing and rendering unit 70 may, for example, digest the video content by editing the video content based on the events extracted by the event extraction unit 30. This digesting of the video content may be performed before or after the avatar animation and audio data are superimposed.

[0038] The distribution unit 80 distributes the video content with comments generated by the editing and rendering unit 70 to the viewer's terminal via a predetermined network 90. ​​The predetermined network may be any of various networks, such as the Internet, a LAN (Local Area Network), a WAN (Wide Area Network), or a mobile communication system (including 4G (4th Generation Mobile Communication System), 4G-LTE (Long Term Evolution), 5G (5th Generation Mobile Communication System), etc.).

[0039] 1.5 Example of operation flow when adding a comment when posting (uploading) Figures 5A to 9C are diagrams showing an example of an operational flow when adding a comment when posting (uploading) video content (see Figure 3A). The sequence of this operational flow is shown in Figures 5A, 6A, 7A, and 9A. This description also illustrates an example of adding an audio comment to video content.

[0040] 1.5.1 Example of operation from app launch to virtual commentator setting First, the operation from when user U1 starts an application on user terminal 10 to when a virtual commentator is set will be described. FIG. 5A is a flowchart showing an example of the operation from when user U1 starts an application on user terminal 10 to when a virtual commentator is set. Also, FIGS. 5B to 5E are diagrams showing examples of management tables managed by a provider of an automatic comment-adding service via an application. FIG. 5B shows an example of a video content management table for managing video content. FIG. 5C shows an example of a character management table for managing the characters of virtual commentators. FIG. 5D shows an example of a position management table for managing the positions of virtual commentators. FIG. 5E shows an example of a comment target management table for managing targets (comment targets) on which the virtual commentators make comments. In this description, it is assumed that the provider of the video distribution service and the provider of the automatic comment-adding service are the same.

[0041] 5A, when user U1, who is a video poster, launches an app on user terminal 10, the app on user terminal 10 (hereinafter simply referred to as user terminal 10) acquires user information about user U1 from user information storage unit 20 (step S101). That is, in this example, user terminal 10 can function as an acquisition unit that acquires information about the relationship between user U1, who is a video content poster, and viewers of the video content. The user information may include, as described above, information about user U1 himself, information about other users (including viewers) related to user U1, and information about the relationship between user U1 and other users, as well as history information such as information about video content already uploaded by user U1 and information about video content to which user U1 has previously commented using the app and its genre.

[0042] When the user information is acquired, the user terminal 10 acquires a list of video content (also called a video list) such as live-action footage shot by the user U1 and gameplay videos of games, etc., based on the acquired user information, and creates a video content selection screen G10 shown in Fig. 2A using the acquired video list and displays this to the user U1 (step S102). The video list on the video content selection screen G10 may prioritize video content in genres that are likely to be commented on, based on history information included in the user information.

[0043] When the user U1 selects a video content to which a comment is to be added based on the video content selection screen G10 displayed on the user terminal 10 (step S103), the user terminal 10 acquires the genre (video genre) of the selected video content from the meta information (hereinafter referred to as video information) added to the selected video content. Note that the video information may be tag information such as the title or genre of the video content, the name of a game, or the name of a dish. Next, the user terminal 10 acquires the genre ID of the selected video content by referring to the video content management table shown in FIG. 5B using the acquired video genre. Next, the user terminal 10 acquires a list of virtual commentators for the selected video content and their priorities by referring to the character management table shown in FIG. 5C using the acquired genre ID. Then, the user terminal 10 creates a list of virtual commentators selectable by the user U1 based on the acquired list of virtual commentators and their priorities, and creates and displays the commentator selection screen G20 shown in FIG. 2B using the created list of virtual commentators (step S104).

[0044] For example, if user U1 selects video content of a soccer game, the genre ID will be G03, and if user U1 selects video content of cooking, the genre ID will be G04. The user terminal 10 refers to the character management table shown in Fig. 5C using the genre ID identified based on the video content management table shown in Fig. 5B to obtain a commentator list for each genre ID and the priority of each virtual commentator, creates a commentator selection screen G20 that lists icons (options) of the virtual commentators according to the priorities, and displays this to user U1.

[0045] Here, the priority may be, for example, the order in which virtual commentators are prioritized for each video genre. For example, in the example shown in Fig. 5C, for the video genre "game" (genre ID = G03), the "girl" character with character ID = C01 has the highest priority, the "boy" character with character ID = C02 has the second highest priority, and the "male" character with character ID = C03 has the third highest priority. Therefore, the characters are displayed in that order on the commentator selection screen G20.

[0046] Various methods may be used for setting the priority, such as a rule-based method of setting characters suitable for each video genre (e.g., a character who likes sports), or a method of setting the priority of each character based on the user's preference history, etc.

[0047] When user U1 selects a virtual commentator character using the commentator selection screen G20 (step S105), the user terminal 10 acquires a list of positions that can be set for the virtual commentator by referring to the position management table shown in FIG. 5D using the genre ID, and generates the position selection screen G30 shown in FIG. 2C using the acquired list of positions and displays it to user U1 (step S106). For example, when genre ID=G03, position ID=S01 is a "player," S02 is a "team member," and S03 is a "friend." Also, when genre ID=G04, position ID=S01 is a "presenter" who explains the dish, S02 is a "guest" who samples the food, and S03 is a "spectator."

[0048] When user U1 selects the position of the virtual commentator using the position selection screen G30 (step S107), the user terminal 10 acquires a list of comment targets to be commented on by referring to the comment target management table shown in Fig. 5E using the genre ID, and generates the comment target selection screen G40 shown in Fig. 2D using the acquired list of comment targets and displays it to user U1 (step S108). For example, when genre ID=G03, target ID=T01 is "player", T02 is "team member", and T03 is "friend".

[0049] In this way, the user terminal 10 sets the position of the virtual commentator and the comment target of the comment made by the virtual commentator, and notifies the event extraction unit 30 of the setting contents. Note that, in steps S103, S105, and S107, the case where the user U1 selects each item has been exemplified, but this is not limiting, and for example, the user terminal 10 may be configured to automatically select each item based on the service definition, user preferences, etc.

[0050] 1.5.2 Event extraction operation example Once the virtual commentator has been set as described above, the event extraction unit 30 then executes an operation to extract events from the video content selected by user U1. FIG. 6A is a flowchart showing an example of the event extraction operation executed by the event extraction unit 30. FIGS. 6B and 6C are diagrams showing examples of management tables managed by the event extraction unit 30. FIG. 6B shows an example of a model management table that manages the recognition models used for event extraction for each video genre, and FIG. 6C shows an example of an event management table for managing events extracted from video content for each recognition model. FIG. 6D is a diagram showing an example of a list of event data extracted from video content in chronological order by the recognition models.

[0051] As shown in FIG. 6A, when the event extraction unit 30 acquires the video genre (genre ID) identified in step S104 of FIG. 5A from the user terminal 10, the event extraction unit 30 uses the genre ID to refer to the model management table shown in FIG. 6B to identify a recognition model ID for identifying a recognition model associated with each video genre, and then acquires a recognition model to be used for event extraction using the identified recognition model ID (step S111). Note that a recognition model for each genre may be constructed in advance using, for example, a rule base or machine learning and stored in the event extraction unit 30. For example, when the genre ID is G03, the event extraction unit 30 selects a recognition model with a model ID of M03. Note that a recognition model that can be used generally within each video genre may be a recognition model that covers the entire video genre, or may be a recognition model that covers a range that is more detailed than the classification by video genre. For example, for the video genre "Sports 1," recognition models specialized for more detailed classifications, such as "for baseball," "for rugby," and "for soccer," may be prepared.

[0052] Next, the event extraction unit 30 extracts events from the video content by inputting the video content into the acquired recognition model (step S112). More specifically, the event extraction unit 30 extracts events by analyzing video features, the movements of people and the ball recognized by the recognition model, feature points extracted from image data such as on-screen data display (e.g., scores), and keywords recognized from audio data. Events according to this embodiment may be defined not only by a broad genre such as "sports," but also by smaller genres such as "baseball" or "soccer," or by specific game titles. For example, in the case of a soccer game, event ID=E001 may be "goal," E002 may be "shot," and representative techniques, scores, and shots may be defined as events. Furthermore, the extracted events may include parameters indicating whether the event is for the player's team or the opposing team. Furthermore, information such as the name and uniform number of the player who caused each event, the ball position, and the total points may be acquired as part of the event.

[0053] After extracting the event in the above manner, the event extraction unit 30 generates event data together with a time code indicating the time at which the event occurred in the video content, and adds the generated event data to the event data list shown in FIG. 6D (step S113).

[0054] Then, the event extractor 30 repeatedly executes the above-described event extraction operation until the end of the video content (NO in step S114).

[0055] 1.5.3 Example of comment generation Once the event data list is created from the video content as described above, the comment generation unit 40 then performs an operation of generating comments (text data) for each event based on the virtual commentator selected by user U1, their position, and the comment target. FIG. 7A is a flowchart showing an example of the comment generation operation performed by the comment generation unit 40. FIG. 7B is a diagram showing an example of a management table managed by the comment generation unit 40, showing an example of a list of comments for each virtual commentator's position for each event (position comment list) and an example of a used comment history that manages the history of comments used in the past at each position for each event. FIG. 7C is a diagram showing an example of a list of comments for each event (target comment list) generated by the comment generation unit 40.

[0056] 7A, when the comment generating unit 40 acquires the event data list created in step S113 of Fig. 6A from the event extracting unit 30, the comment generating unit 40 acquires event data in chronological order from the event data list (step S121). The acquisition of event data in chronological order may be performed based on a time code, for example.

[0057] Here, when events occur one after another, it may be difficult to add comments to all of the events. In such a case, as in step S122, the comment generation unit 40 may filter the event data list. For example, the comment generation unit 40 may determine the time interval between events based on the time codes in the event data list, and for events that are not separated by a predetermined time interval (e.g., 15 seconds) or more (NO in step S122), the comment generation unit 40 may return to step S121 without generating a comment and acquire the event with the next time code. In this case, priorities may be assigned to the events (e.g., prioritizing goals over fouls), and events with low priorities may be preferentially excluded. However, the present invention is not limited to this. Events may be assigned to multiple virtual commentators without filtering, and each virtual commentator may comment on the assigned event.

[0058] Next, the comment generation unit 40 acquires a comment list corresponding to the virtual commentator's position for each event by referencing the position comment list shown in FIG. 7B based on the event ID in the event data acquired in step S121 and the position set for the virtual commentator (step S123). For example, in the case of event ID=E001 (goal (opponent team)), the position comment list shown in FIG. 7B lists one or more variations of comments that a person in that position would think when a goal is scored. The position comment list may be created based on rules or using machine learning, etc. Furthermore, each comment may be designed to be universally usable across events. For example, each comment may be tagged with a tag that succinctly describes its content, such as a "disappointed" tag for a comment that says "I'm depressed" and an "encouraging" tag for a comment that says "I can turn the game around." Tags may be used, for example, to add variations to comments, and multiple tags may be attached to a single comment. Note that in FIG. 7B, text that can be replaced depending on the event is enclosed in "<>". For example, comment W1001 can be reused by inserting the event name into <event>, such as "I was depressed when the opposing team scored a goal" or "I was depressed when the opposing team shot."

[0059] Furthermore, the comment generating unit 40 refers to the comment usage history for the comment list acquired in step S123, and extracts n comments (n is an integer equal to or greater than 1) that have been used in the past for the same user U1 (step S124). The comments that have been used in the past for the same user U1 may be comments that have been used for video content different from the video content currently being processed, or comments that have been used for the same video content, or may include both of these.

[0060] Next, the comment generation unit 40 selects one of the comments (hereinafter, a comment according to a position is referred to as a position comment) obtained by excluding the n comments obtained in step S124 from the comment list obtained in step S123 (step S125). In this step S125, for example, the position comment may be selected randomly using pseudo-random numbers or the like, or the position comment may be selected according to the order of the comment list obtained by excluding the past n comments. In this way, by controlling the repetition or frequent appearance of the same comment based on the history of comments used in the past, it is possible to prevent user U1 and viewers from getting bored with the virtual commentator's comments. Note that step S124 may be configured to vectorize the event and the comment, respectively, and select the closest candidate.

[0061] Next, the comment generation unit 40 analyzes the position comment selected in step S125 using morphological analysis or the like (step S126). Subsequently, the comment generation unit 40 omits the event name included in the position comment (step S127). This is because, if it is assumed that the viewer recognizes the occurrence of an event from the video, omitting the event name allows for a comment with a better tempo. However, if the virtual commentator is a character that requires accurate remarks, such as an announcer, the omission of the event name (step S127) may be skipped.

[0062] Next, the comment generating unit 40 adds an exclamation to the position comment (step S128). For example, when the team's team shoots, an exclamation expressing anticipation such as "It's here!" is added, and when the opposing team scores, an exclamation expressing disappointment such as "Oh no," is added, thereby making it possible to encourage empathy from the comment target.

[0063] Next, the comment generating unit 40 acquires proper nouns and pronouns such as player names, team names, and nicknames from the video content and the user information storage unit 20, and adds them to the position comment (step S129). This makes it possible to create a sense of intimacy with the comment subject.

[0064] Next, the comment generation unit 40 converts the ending of the position comment into one suitable for calling out to the target (step S130). For example, the "position comment" "Even though there was an <event>, I'm going to cheer for you" is converted into "Everyone, let's cheer for George!" This makes it possible to call out to the comment target in a more natural way. Note that in a language like English where expressions of relationships between users (for example, the presence or absence of "please") appear somewhere other than at the end of a sentence, the process of step S130 may be performed to convert something other than the ending.

[0065] In this way, in steps S127 to S130, the comment generating unit 40 modifies the generated position comment based on information about the relationship between the user U1 and other users (viewers, etc.).

[0066] Next, the comment generating unit 40 adds an event ID, a parameter, and a time code to the position comment, and registers it in the target comment list shown in Fig. 7C (step S131). Note that the added time code may be the time code included in the corresponding event data.

[0067] Then, the comment generating unit 40 repeatedly executes this operation (NO in step S132) until the generation of comments for all event data is completed (YES in step S132).

[0068] In this example, an event data list is created from the entire video content, and then a position comment is generated for each event, but this is not limited to this, and for example, the system may be configured to generate a position comment each time an event is extracted from the video content.

[0069] 1.5.3.1 Variations of inflection Here, a modified example of the suffix conversion shown in step S130 of FIG. 7A will be described. In the suffix conversion of step S130, for example, by converting the suffix of the position comment based on the relationship or hierarchical relationship between user U1 (player) and viewers in the community, it becomes possible to make the comment subject feel more familiar, or to use language that suits the relationship between user U1 (player) and viewers or the relationship between virtual commentators in the case where there are multiple virtual commentators. FIG. 8A is a flowchart showing an example of the suffix conversion operation according to this modified example. FIG. 8B is a diagram showing an example of a management table managed by the comment generating unit 40, and shows an example of a hierarchical relationship management table that manages the hierarchical relationships between virtual commentators. Note that, while the conversion of suffixes in a language in which expressions of relationships between users appear at the end of a sentence are described here, in a language in which expressions of relationships between users (e.g., the presence or absence of "please") appear in a position other than the end of a sentence, such as English, the processing of step S130 may be performed to convert parts other than the suffix.

[0070] As shown in FIG. 8A, the comment generation unit 40 first checks the attributes of the selected position and the comment target from the characteristics of the service used by user U1 and the user information, and determines whether the attributes of the selected position and the comment target are the same (step S1301). The same attributes refer to cases where the service is for a limited age group of viewers (such as high school girls only), where the occupation is limited to doctors or teachers only, or where the user information indicates a community that likes a specific game. If the selected position and the comment target have the same attributes (YES in step S1301), the comment generation unit 40 proceeds to step S1302. On the other hand, if the selected position and the comment target have different attributes (NO in step S1301), the comment generation unit 40 proceeds to step S1304.

[0071] In step S1302, the comment generating unit 40 acquires a dictionary according to the attribute identified in step S1301. This is because, when the selected position and the comment target have the same attribute, it is desirable to use trendy words and technical terms frequently used within the community in the comment, and therefore it is desirable to use a dictionary according to the attribute. Note that the dictionary may be youth slang, trendy words, technical terms, or a corpus collected from the community's SNS (Social Networking Service). The dictionary may be stored in advance in the comment generating unit 40, or may be obtained by searching information on a network such as the Internet in real time.

[0072] Next, the comment generating unit 40 replaces the words in the position comment with those in the community language if there are any equivalents (step S1303).

[0073] Next, the comment generating unit 40 checks the hierarchical relationship between the selected position and the comment target by referring to the hierarchical relationship management table shown in Fig. 8B (step S1304). In the hierarchical relationship management table, the hierarchical relationship is defined as a parameter for position setting, and for example, in a cooking video content whose video genre is "hobbies," the presenter's hierarchical level is "2 (middle)," the guest's level is the highest at "1," and the audience's level is "2," the same as the presenter's.

[0074] If the virtual commentator's position is higher than the comment target (step S1304 "position>target"), the comment generation unit 40 converts the ending of the position comment to a commanding ending (step S1305), if the position and the comment target are at the same level (step S1304 "position=target"), the comment generation unit 40 converts the ending of the position comment to a friendly ending (step S1306), and if the position is lower than the comment target (step S1304 "position<target"), the comment generation unit 40 converts the ending of the position comment to a polite ending (step S1307). Thereafter, the comment generation unit 40 returns to the operation shown in FIG. 7A.

[0075] 1.5.4 Example of generating and distributing video content with comments Next, an example of an operation for generating video content with comments using the position comments generated as described above and an operation for distributing the generated video content with comments will be described. FIG. 9A is a flowchart showing an example of an operation for generating and distributing video content with comments. Note that the comment data and avatar animation may be distributed to user U1 or viewers as data separate from the video content. In that case, for example, step S160 may be omitted.

[0076] As shown in FIG. 9A, when the speech control unit 50 obtains the target comment list created in step S131 of FIG. 7A from the comment generation unit 40 (step S141), it converts the position comment into audio data (also called comment audio) by performing a voice synthesis process using TTS on the text data of each position comment in the target comment list (step S142), and saves the comment audio of each position comment thus generated (step S143).

[0077] Furthermore, the avatar generation unit 60 executes an avatar generation process for generating an avatar animation of the virtual commentator (step S150).

[0078] Next, the editing and rendering unit 70 executes an editing and rendering process to generate video content with comments from the video content, comment voice, and avatar animation (step S160).

[0079] The video content with comments thus generated is distributed from the distribution unit 80 to the user U1 and viewers via the predetermined network 90 (step S171), after which the operation ends.

[0080] 1.5.4.1 Example of avatar generation process The avatar generation process shown in step S150 of Fig. 9A will now be described in more detail. Fig. 9B is a flowchart showing an example of the operational flow of the avatar generation process.

[0081] 9B, the avatar generation unit 60 selects a corresponding avatar from the virtual commentator character selected by the user U1 in step S105 of Fig. 5A (step S1501). The avatar generation unit 60 also acquires the comment voice generated by the speech control unit 50 (step S1502).

[0082] Then, the avatar generation unit 60 creates an animation that moves the avatar selected in step S1501 in accordance with the speech section of the comment voice (step S1503), and returns to the operation shown in Fig. 9A. Note that in step S1503, a way of opening the mouth in accordance with the speech may be learned to generate a realistic avatar animation.

[0083] 1.5.4.2 Editing and rendering process example Next, the editing and rendering process shown in step S160 of Fig. 9A will be described in more detail. Fig. 9C is a flowchart showing an example of the operational flow of the editing and rendering process.

[0084] As shown in FIG. 9C, the editing / rendering unit 70 first acquires the video content selected by the user U1 in step S103 of FIG. 5A, the comment audio saved in step S143 of FIG. 9A, and the avatar animation created in step S150 of FIG. 9A (step S1601).

[0085] Next, the editing / rendering unit 70 arranges the comment voice and avatar animation in the video content in accordance with the time code (step S1602), and generates video content with comments by rendering the video content in which the comment voice and avatar animation are arranged (step S1603). Note that text data of the position comments may be arranged as subtitles in the video content.

[0086] The video content with comments thus generated may be distributed in advance to user U1 (player) via distribution unit 80 for confirmation by user U1 before being distributed to viewers (step S164). When confirmation (distribution approval) is received from user U1 in response via a communication unit or the like (not shown) (step S165), the operation returns to the operation shown in Fig. 9A.

[0087] 1.6 Example of dynamically changing position and comment target The position selected in step S107 of Fig. 5A and the comment target selected in step S109 may change during the video content. Fig. 10A is a diagram for explaining an example of changing the comment target during the video content. Fig. 10A shows a case where the position of the virtual commentator and / or the comment target are dynamically changed based on the ball possession rate of the player's team over time during a soccer game. Fig. 10B is a diagram showing an example of a comment made by the virtual commentator at each time in Fig. 10A.

[0088] In FIG. 10A, (a) shows a graph of the team's ball possession rate, (b) shows the transition of the virtual commentator's position, and (c) shows the transition of the comment target. The ball possession rate can be obtained, for example, from image analysis of video content or an API published by the game. The virtual commentator's initial position is set to "friends," and the initial comment target is also set to "friends."

[0089] At time T1, the virtual commentator's team has scored a goal and has a high ball possession rate, so the target of the comment remains "friends." Therefore, as shown in Figure 10B, at time T1, the virtual commentator comments to the friends, "Yay! You're doing great today."

[0090] At time T2, the ball possession rate falls below a preset threshold (for example, 50%). At the timing (time T2) when the ball possession rate falls below the threshold, the position / target control unit 41 of the comment generation unit 40 makes a comment to the friend, saying, "Hey, isn't this getting serious? I'll go cheer you up a bit," and then switches the comment target from "friends" to "team members." Therefore, the comment generation unit 40 operates to issue comments to the team members while the comment target is set to "team members."

[0091] When the team concedes a goal at time T3, the virtual commentator makes an encouraging comment to the team members, saying, "Team A (team name), don't worry! Your movements aren't bad."

[0092] After that, when the ball possession rate exceeds the threshold at time T5, the position / target control unit 41 changes the comment target back to "friends" from "team members." Therefore, the comment generating unit 40 operates to generate comments to be sent to friends after time T5.

[0093] 1.6.1 Operational flow example Fig. 11A is a flowchart showing an example of a comment generation operation executed by the comment generation unit 40 when the standing position and comment target are dynamically changed. Fig. 11A shows an example of an operation when a step of dynamically changing the standing position and comment target is incorporated into the operation shown in Fig. 7A. In this example, it is assumed that the events extracted by the event extraction unit 30 include an occurrence of the ball possession rate exceeding a threshold and an occurrence of the ball possession rate falling below a threshold.

[0094] As shown in FIG. 11A, when the position and comment target are dynamically changed, the position and target control unit 41 of the comment generation unit 40 acquires event data in a predetermined order from the event data list in the operation shown in FIG. 7A (step S121), and then executes a process to adjust the position and / or comment target of the virtual commentator (step S220).

[0095] Fig. 11B is a flowchart showing a more detailed operation flow of the position / target adjustment operation shown in step S220 of Fig. 11A. As shown in Fig. 11B, in step S220 of Fig. 11A, the position / target control unit 41 first sets a threshold value for the ball possession rate (step S2201). Next, the position / target control unit 41 acquires the ball possession rate of the player's team based on, for example, the event ID and parameters in the event data acquired in step S121 (step S2202).

[0096] Next, if the ball possession rate of the player's team acquired in step S2202 is below the threshold set in step S2201 ("Below" in step S2203), the position / target control unit 41 switches the comment target to "Team Members" and returns to the operation shown in Fig. 11A. On the other hand, if the ball possession rate of the player's team is above the threshold ("Above" in step S2203), the position / target control unit 41 switches the comment target to "Friends" and returns to the operation shown in Fig. 11A. Thereafter, the comment generating unit 40 performs the same operations as those from step S122 onwards in Fig. 7A.

[0097] While Figures 10A to 11B primarily illustrate examples in which the comment target is changed, it is also possible to change the position of the virtual commentator, or both the position and the comment target. For example, in an active learning workshop video, based on the activity level (e.g., amount of conversation) of each team, the position of the team member with the lowest activity level could be set as the virtual commentator, and the virtual commentator could then make comments to the team members in order to energize that team. This type of change in position and / or comment target is effective not only for archived content, but also for real-time online classes.

[0098] 1.7 Example of generating comments based on sentiment values Comment generation section 40 according to this embodiment may vary the comments it generates not only in accordance with events and positions, but also, for example, emotional values. Fig. 12A is a diagram illustrating an example of generating comments based on emotional values. Fig. 12B is a table showing the amount of change in emotional value for each position for each event, and Fig. 12C is a diagram illustrating the change in emotional value shown in Fig. 12A in association with the event.

[0099] There are various models for classifying emotions, such as the Russell model and the Plutchik model, but this embodiment deals with a simple emotional level of positive / negative. In this embodiment, the emotional value is mapped to a value between 0 and 1, with 1 being the most positive and 0 being the most negative.

[0100] As shown in Figure 12B, for example, if the player's position is "Player," the emotional value increases by 0.15 (becomes positive) when the player's team scores, and decreases by 0.15 (becomes negative) when the opposing team scores. Also, if the player's position is "Friend," the emotional value increases by 0.1 when the player's team scores, but even if the opposing team scores, the game becomes more interesting, so the emotional value increases by 0.05. Also, if either team scores three or more goals in a row, the game becomes monotonous and uninteresting, so the emotional value decreases by 0.1 when the player's position is "Friend."

[0101] The amount of change in emotion value shown in FIG. 12B may be defined on a rule-based basis from general emotions, or may be a value acquired by sensing user U1 or the like.

[0102] 12A and 12C show an example in which the player's team commits a foul at time T2, the opposing team scores goals at times T3 and T4, the opposing team commits a foul at time T4, and the player's team scores goals at times T6 to T9. At the starting point, time T0, the initial emotion values ​​of the "player" and "friend" are set to 0.5.

[0103] As shown in Figures 12A and 12C, as a result of changing the emotion values ​​at each standing position due to goals or fouls by either the player's own team or the opposing team, it can be seen that there are times when the emotion values ​​match and times when there are differences.

[0104] 1.7.1 Operational flow example The operational flow of comment generation when the virtual commentator's position is "friend" and the comment target is "player" is shown in Fig. 13. Note that in Fig. 13, the same operations as those in Fig. 7A are denoted by the same reference numerals, and detailed explanations thereof will be omitted.

[0105] As shown in FIG. 13, when generating a comment based on an emotion value, in the operation shown in FIG. 7A, comment generation unit 40 acquires event data from the event data list in a predetermined order (step S121), and then acquires the emotion value a of the "player" who is the comment target included in the event data (step S301), as well as the emotion value b of the "friend" who is the position of the virtual commentator (step S302).

[0106] Next, the comment generation unit 40 determines whether the emotional value a of the comment target is lower than the threshold value for a negative state (defined as 0.3, for example) (step S303), and if it is lower (YES in step S303), sets "encouragement" as a search tag for the position comment list shown in FIG. 7B (step S305), and proceeds to step S122.

[0107] If the emotion value a of the comment target is equal to or greater than the threshold for the negative state (NO in step S303), the comment generation unit 40 determines whether the absolute value (|ab|) of the difference in emotion value between the comment target and the virtual commentator is greater than a threshold (defined as 0.3, for example) (step S304). If the absolute value of the difference in emotion value is greater than the threshold (YES in step S304), the comment generation unit 40 sets "sympathy" and "sarcasm" as search tags (step S306) and proceeds to step S122.

[0108] On the other hand, if the absolute value of the difference in emotion value is equal to or less than the threshold value (NO in step S304), the comment generating section 40 sets "sympathy" as the search tag (step S307), and proceeds to step S122.

[0109] Next, after executing steps S122 to S124, the comment generating unit 40 selects one of the position comments to which the search tag set in step S305, S306 or S307 is attached from the position comments obtained by excluding the n comments obtained in step S124 from the comment list obtained in step S123 (step S325). In this step S325, as in step S125, the position comment may be selected randomly using, for example, pseudo-random numbers, or the position comment may be selected according to the order of the comment list obtained by excluding the past n comments.

[0110] Thereafter, the comment generating unit 40 performs the same operations as steps S126 to S132 shown in FIG. 7A to create a target comment list.

[0111] In this way, by limiting the selection range of position comments based on emotional values, in other words, by narrowing down the candidate position comments using search tags set based on emotional values, it is possible to generate comments that are more in line with the emotions of players and friends. For example, by setting the search tag "encouragement," it is possible to generate a comment that encourages a player who is feeling negative and discouraged, saying, "Don't worry, it's only just beginning! You can make up for it!". Also, by setting the search tag "empathy," it is possible to generate a comment that sympathizes with a discouraged player, saying, "That's tough, I know!". Furthermore, by setting both "empathy" and "sarcastic" as search tags, it is possible to generate comments that express complex psychological states. For example, a friend who is feeling bored and bored with a player who is scoring consecutive goals can praise a player who is very positive, saying, "Wow, that's amazing! Is it because of the difference in PC specs?" while sarcastic, suggesting that the difference is not in skill but in equipment. In this way, by generating comments based on emotional relationships, it is possible to generate more human-like comments.

[0112] The definition of the emotional value for the event shown in Fig. 12B may be changed. For example, a beginner may feel happy even if he or she scores with a simple technique, but an expert may feel happy only when he or she scores with a difficult technique. Therefore, an emotion estimation model may be created in which the emotional value (increase rate) for a simple technique gradually decreases and the emotional value for a complex technique gradually increases, by acquiring the user's proficiency in the technique from information such as the number of hours of use and the number of trophies in the game.

[0113] 1.8 Example of system configuration for adding comments when viewing (downloading) Next, the case where a comment is added when viewing (downloading) video content (see FIG. 3B) will be described below.

[0114] When adding a comment when viewing (downloading) video content, a comment that is more suited to the viewer's preferences and circumstances can be generated compared to adding a comment when posting (uploading). The viewer may select the video content to view by, for example, launching an app on their device and using a video content selection screen G10 (see FIG. 2A) that is displayed thereby. As with adding a comment when posting (uploading) video content, it is also possible to configure the system so that the viewer selects the character, position, and comment target of the virtual commentator (see FIGS. 2B to 2D). In this description, however, an example is given in which the degree of intimacy between the viewer and the video poster (user U1 in this description) is recognized from information about the viewer and the video poster, and the character, position, and comment target of the virtual commentator are automatically set.

[0115] 14 is a block diagram showing an example of a system configuration for adding a comment when viewing (downloading) video content. Note that, in this description, a case where an audio comment is added to video content is illustrated as an example, similar to FIG. 4.

[0116] As shown in Figure 14, an example system configuration of the information processing system 1 when adding comments when viewing (downloading) video content is similar to the example system configuration when adding comments when posting (uploading) video content described using Figure 4, except that a setting unit 120 is added, the distribution unit 80 is omitted, and the user terminal 10 is replaced with the user terminal 110 of the viewer A10.

[0117] Like the user terminal 10, the user terminal 110 is an information processing device on the viewer A10 side, such as a smartphone, personal computer, or game console, and executes an application for automatically playing video content and adding comments. The user terminal 110 also includes a communication unit (not shown), and acquires user information stored in the user information storage unit 20 and video content to which comments are to be added via the communication unit.

[0118] When the setting unit 120 receives a selection of video content to which a comment is to be added, notified from the user terminal 110, the setting unit 120 sets the position of the virtual commentator and the comment target based on the video information of the selected video content and the user information of the user U1 and the viewer A10 acquired from the user information storage unit 20. That is, in this example, the setting unit 120 can function as an acquisition unit that acquires information about the relationship between the user U1 who posted the video content and the viewer of the video content. The setting unit 120 also sets the position of the virtual commentator and the comment target of the comment to be made by the virtual commentator, and notifies the event extraction unit 30 of the setting contents. Note that, as described with reference to FIGS. 2B to 2D and 5A, etc., if the viewer A10 manually sets the character, position, and comment target of the virtual commentator, the setting unit 120 may be omitted.

[0119] Other configurations may be similar to the system configuration example of the information processing system 1 shown in FIG. 4, and therefore detailed description thereof will be omitted here.

[0120] 1.9 Example of operation flow when adding comments when viewing (downloading) The operational flow for adding a comment when viewing (downloading) a video may basically be the same as the operational flow for adding a comment when posting (uploading) video content described above with reference to Figures 5A to 9C. However, when adding a comment when viewing (downloading), steps S164 and S165 in Figure 9C are omitted, and step S171 in Figure 9A is replaced with playing the video content with the comment.

[0121] Furthermore, in this description, as described above, the operation shown in FIG. 5A is replaced with the operation shown in FIG. 15A to illustrate a case where the character, position, and comment target of a virtual commentator are automatically set by recognizing the degree of intimacy from information between viewer A10 and a video poster (user U1 in this description). Also, FIG. 15B is a diagram showing an example of a management table managed by a provider of an automatic comment-adding service via an app, and shows an example of a character management table for managing the character of a virtual commentator. Note that, in this description, as described above, it is assumed that the provider of the video distribution service and the provider of the automatic comment-adding service are the same.

[0122] 15A, when user (viewer) A10 starts an app on user terminal 110, the app on user terminal 110 (hereinafter simply referred to as user terminal 110) acquires user information about viewer A10 from user information storage unit 20 (step S401). Like the user information of user U1, the user information may include information about viewer A10 himself, information about other users (including user U1) related to viewer A10, and information about the relationship between viewer A10 and other users, as well as history information such as the viewing history of viewer A10 and information about the video genres, game genres, etc. preferred by viewer A10.

[0123] Upon acquiring the user information, the user terminal 110 acquires a list (video list) of video content viewable by the viewer A10 based on the acquired user information, creates a video content selection screen G10 shown in FIG. 2A using the acquired video list, and displays this to the viewer A10 (step S402). The video list on the video content selection screen G10 may prioritize videos in genres that are likely to be commented on, based on history information included in the user information. Note that if the commenting function is implemented as a function of a game app, a video playback app, a video distribution app, or a social networking service (SNS) app that can handle video content, the video content to which text is to be added may be selected by displaying a screen that asks the user whether or not to add a comment to the video content being played or displayed, instead of the video content selection screen G10.

[0124] When viewer A10 selects video content to which a comment is to be added based on the video content selection screen G10 displayed on user terminal 110 (step S403), user terminal 110 notifies setting unit 120 of the selected video content. Setting unit 120 acquires the genre (video genre) of the selected video content from meta information (video information) assigned to the video content notified by user terminal 110, and acquires the genre ID of the selected video content by referring to the video content management table shown in FIG. 5B using the acquired video genre. Next, setting unit 120 acquires history information of viewer A10 from the user information acquired in step S401, and matches the acquired history information with character information of the virtual commentator managed in the character management table shown in FIG. 15B to automatically select a virtual commentator that matches the video content and the preferences of viewer A10 (step S404). For example, if viewer A10 likes anime, setting unit 120 selects an anime character (character Y) with character ID = C11.

[0125] Next, the setting unit 120 acquires the user information of the user U1 who uploaded the video content selected in step S403 to the cloud 100 (see FIG. 3B) (step S405).

[0126] Next, the setting unit 120 acquires the intimacy level with the viewer A10 from the name, service ID, etc. of user U1 included in the user information of user U1 (step S406). The intimacy level may be set based on information such as whether or not the viewer A10 and user U1 have played online games together, whether or not the viewer A10 are registered as friends, whether or not the viewer U1 has previously viewed video content posted by user U1, or whether or not the viewer A10 has interacted with user U1 on social media or the like. For example, if the viewer A10 and user U1 are registered as friends who play online games together or have interacted with user U1 on social media within the past month, the intimacy level may be set to the highest level of 3. If the viewer A10 has not interacted with user U1 directly but has previously viewed video content uploaded by user U1 three or more times, the intimacy level may be set to the second highest level of 2. Otherwise, the intimacy level may be set to the lowest level of 1.

[0127] The setting unit 120 sets a position and a comment target defined in a rule base or the like based on the intimacy acquired in step S406. Specifically, if the intimacy between the viewer and the poster is level 3 ('3' in step S406), the setting unit 120 determines that the viewer and the poster are close, sets a team member in the virtual commentator's position, and sets the poster (player = user U1) as the comment target (step S407). Also, if the intimacy between the viewer and the poster is level 2 ('2' in step S406), the setting unit 120 determines that the viewer and the poster are not as close as the team members, sets a friend in the virtual commentator's position, and sets the poster (player = user U1) as the comment target (step S408). On the other hand, if the intimacy between the viewer and the poster is level 1 ('1' in step S406), the setting unit 120 selects a wait-and-see style by setting a friend in the position of the virtual commentator and setting a friend as the comment target (step S408).

[0128] In addition, the viewer A10's level of awareness of the video content may be obtained from the viewer A10's viewing history, and if it is the viewer's first time watching the video content, the comment may prioritize beginner-level explanations and basic rules, and as the viewer becomes more familiar with the content, more advanced comments may be included. Furthermore, the comment may be changed according to the sensed emotional arousal level of the viewer A10. For example, if the viewer A10 is relaxed, a comment or comment voice with calm content and tone may be generated, and if the viewer A10 is excited, a comment or comment voice with more stimulating content and tone may be generated. Alternatively, the sensed pleasantness and unpleasantness of the viewer A10's emotions may be learned, and the results may be used to generate the comment or comment voice.

[0129] As described above, comments can be generated according to the situation each time the game is viewed. For example, in the first viewing of a game in which the player suffered a crushing defeat, the virtual commentator in the "player" position might generate a comment to the commentee "friend" saying, "Oh, that's frustrating! I'll practice more!" However, the second viewing six months later could detect the player's improvement by obtaining the number of wins and trophies acquired since then from the player's activity information, and generate a comment that includes a reflection, such as, "I practiced hard after that disappointment, and that's why I'm where I am today."

[0130] 1.10 Example of generating interactive comments by two virtual commentators for real-time video streaming Next, we will explain an example of generating interactive comments by two virtual commentators for real-time video streaming (hereinafter also referred to as live streaming). In this explanation, the way the two virtual commentators interact is assumed to be, for example, when an event occurs, the commentator first comments on "what event happened," and then the analyst provides an explanation of that.

[0131] Fig. 16A is a block diagram showing an example of a system configuration when two virtual commentators generate exchange comments for real-time video distribution. As shown in Fig. 16A, the system configuration example of the information processing system 1 when two virtual commentators generate exchange comments for real-time video distribution has a configuration in which the setting unit 120 in the system configuration example shown in Fig. 14 is added to the system configuration example shown in Fig. 4, for example.

[0132] Next, an example of operation when two virtual commentators generate a dialogue for a real-time video stream will be described. Fig. 16B is a flowchart showing an example of operation when two virtual commentators generate a dialogue for a real-time video stream.

[0133] 16B, in this operation, when user U1 starts an app on the user terminal 10 and starts preparations for live video streaming, a request for live streaming is notified to the setting unit 120. Upon receiving this request, the setting unit 120 first acquires user information about user U1 from the user information storage unit 20 (step S501).

[0134] Meanwhile, a selection screen for selecting a video genre to be live-streamed is displayed on the user terminal 10. When the user U1 selects a video genre in accordance with this selection screen (step S502), the selected video genre is notified from the user terminal 10 to the setting unit 120. Note that the selection screen may be in various formats, such as an icon format or a menu format.

[0135] Next, the setting unit 120 automatically selects the characters and positions of the two virtual commentators based on, for example, the video genre notified from the user terminal 10 and the user information of the user U1 acquired from the user information storage unit 20 (step S503). In this example, since it is assumed that "Sports 1" in Fig. 5B is selected as the video genre, the setting unit 120 automatically selects two virtual commentators whose positions are "commentator" and "analyst."

[0136] Similarly, the setting unit 120 automatically selects a comment target based on the video genre, user information, etc. (step S504). In this example where the video genre is assumed to be "Sports 1," for example, "viewers" are automatically set as the comment target.

[0137] Next, similar to step S111 in FIG. 6A, the event extraction unit 30 acquires a recognition model based on the video genre of the video content (step S505).

[0138] Once preparations for live streaming are complete in this manner, user U1 then starts live streaming of video content by operating the user terminal 10 or the imaging equipment connected thereto (step S506). Once live streaming starts, captured or imported video data (hereinafter referred to as video content) is sequentially output from the user terminal 10. At this time, the video content may be transmitted by streaming.

[0139] The video content transmitted from the user terminal 10 is input to the event extraction unit 30 directly or via the setting unit 120. As in step S112 of Fig. 6A, the event extraction unit 30 inputs the video content into the recognition model acquired in step S505, thereby extracting events from the video content (step S507). The event data generated in this way is sequentially input to the comment generation unit 40.

[0140] In response to this, the comment generating unit 40 periodically checks whether or not event data has been input (step S508). If no event data has been input (NO in step S508), the operation proceeds to step S540. On the other hand, if event data has been input (YES in step S508), the comment generating unit 40 determines whether or not a predetermined time (e.g., 30 seconds) has elapsed since the time when the most recent event data was input (step S509), and if the predetermined time or more has elapsed (YES in step S509), the operation proceeds to step S520. Note that, in step S122 in FIG. 7A and the like, an example of the minimum event interval was 15 seconds, but since this example is a dialogue between two virtual commentators, a longer interval of 30 seconds is exemplified as shown in FIG. 16C.

[0141] On the other hand, if the next event data is input before the predetermined time has elapsed (NO in step S509), the comment generation unit 40 determines whether the event data is for a high-priority event (step S510). If it is for a high-priority event (YES in step S510), the comment generation unit 40 notifies the editing / rendering unit 70 of a request to interrupt or stop the speech currently being made or about to be made by one of the virtual commentators (speech interruption / cancellation request) in order to interrupt the speech of the previous event (step S511), and proceeds to step S520. If the event data is not for a high-priority event (NO in step S510), the comment generation unit 40 discards the input event data, and the operation proceeds to step S540. For example, if the video content is a video of a soccer game play, an event such as "passing the ball" has a low priority, while an event such as "goal" has a high priority. The priority of each event may be set, for example, in parameters in the event management table shown in FIG. 6C.

[0142] In step S520, the comment generating unit 40 generates a position comment to be uttered by one of the virtual commentators based on the input event data.

[0143] In step S530, the speech control unit 50 converts the text data of the standing position comment generated by the comment generating unit 40 into voice data.

[0144] In step S540, the avatar generation unit 60 generates an avatar animation in which the avatar of the virtual commentator moves in accordance with the voice data generated by the speech control unit 50. Note that an example of the operation of the avatar generation process may be the same as the example of the operation described above with reference to FIG. 9B.

[0145] In step S550, the editing and rendering unit 70 generates video content with comments from the selected video content, audio data, and avatar animation.

[0146] Once the video content with comments has been generated in this manner, the generated video content with comments is live-distributed from the distribution unit 80 via the predetermined network 90 (step S512).

[0147] Thereafter, for example, a control unit (not shown) in the cloud 100 determines whether or not to end the distribution (step S513), and if not to end (NO in step S513), the operation returns to step S507. On the other hand, if to end (YES in step S513), the operation ends.

[0148] 1.10.1 Example of comment generation process Here, the comment generation process shown in step S520 of Fig. 16B will be described in more detail. Fig. 16D is a flowchart showing an example of the operational flow of the comment generation process.

[0149] As shown in FIG. 16D, the comment generation unit 40 first acquires a comment list of the position of the "commentator" from the position comment list illustrated in FIG. 7B for the comment data sequentially input from the event extraction unit 30 (step S5201).

[0150] Next, the comment generation unit 40 refers to the comment usage history for the comment list acquired in step S5201, and selects one of the position comments from the comment list acquired in step S5201, excluding the past n comments (step S5202). In this step S5202, similar to step S125 in FIG. 7A or 11A, the position comment may be selected randomly using, for example, pseudo-random numbers, or may be selected according to the order of the comment list excluding the past n comments. Alternatively, one of the comment lists excluding the position comments used during the current live distribution may be selected.

[0151] Similarly, the comment generation unit 40 first acquires a comment list of the "commentator" position for the comment data sequentially input from the event extraction unit 30 (step S5203), and selects one of the position comments excluding the past n comments based on the comment usage history for the acquired comment list (step S5204).

[0152] Next, similar to steps S126 to S131 in FIG. 7A, the comment generation unit 40 analyzes the position comment selected in steps S5202 and S5204 using morphological analysis or the like (step S5206), omits the event name included in the position comment (step S5207), adds an exclamation to the position comment (step S5208), adds proper nouns, pronouns, etc. to the position comment (step S5209), converts the ending to one suitable for addressing the target (step S5210), and then adds an event ID, parameters, and a time code to the position comment, registers it in the target comment list (step S5211), and returns to the operation shown in FIG. 16B.

[0153] 1.10.2 Speech Control Processing Example The speech control process shown in step S540 in Fig. 16B will now be described in more detail. Fig. 16E is a flowchart showing an example of the operation flow of the speech control process.

[0154] As shown in FIG. 16E, when the speech control unit 50 acquires a target comment list from the comment generation unit 40 (step S5301), it extracts comments with the position of "commentator" and "analyzer" separately from the acquired target comment list (step S5302).

[0155] Next, the speech control unit 50 converts each position comment into voice data (comment voice) by performing voice synthesis processing using TTS on the text data of each position comment of the "commentator" and the "analyzer" (step S5303).

[0156] Next, the speech control unit 50 acquires the speech times of the comment voice of the "commentator" and the comment voice of the "analyzer" (step S5304).

[0157] Next, in order to create timing for the exchange between the commentator and the analyst, the speech control unit 50 sets the commentator's speech start time (time code) to a time after the commentator's speech start time (time code) so that the creator's virtual commentator starts speaking after the commentator's virtual commentator finishes speaking, as shown in Figure 16C (step S5305), and updates the target comment list with the updated commentator's speech start time (step S5306).

[0158] Thereafter, the speech control unit 50 saves the audio files of the commentary voice of the "commentator" and the commentary voice of the "analyzer" (step S5307), and returns to the operation shown in FIG. 16B.

[0159] 1.10.3 Editing and rendering process example The editing and rendering process shown in step S550 of Fig. 16B will now be described in more detail. Fig. 16F is a flowchart showing an example of the operational flow of the editing and rendering process.

[0160] 16F, the editing / rendering unit 70 first determines whether or not a speech interruption request has been notified from the comment generation unit 40 (step S5501), and if not (NO in step S5501), proceeds to step S5503. On the other hand, if a request has been notified (YES in step S5501), the editing / rendering unit 70 discards the comment audio and avatar animation stored in the buffer for the next editing / rendering (step S5502), and proceeds to step S5503. Note that if video content with comments before distribution is stored in the distribution unit 80 or the editing / rendering unit 70, this video content with comments may be discarded.

[0161] In step S5503, editing / rendering unit 70 acquires comment audio from speech control unit 50 and acquires avatar animation from avatar generation unit 60. Next, similar to steps S1602 to S1603 in Fig. 9C, editing / rendering unit 70 arranges comment audio and avatar animation in the video content in accordance with its time code (step S5504), and generates video content with a comment by rendering the video content in which the comment audio and avatar animation have been arranged (step S5505). Then, this operation returns to Fig. 16B.

[0162] 1.11 Example of generating comments based on viewer feedback during real-time video streaming Next, an example of generating comments in response to viewer feedback during real-time video distribution will be described. Fig. 17A is a flowchart showing an example of operation when generating comments in response to viewer feedback during real-time video distribution. Note that this explanation is based on the example of operation described above using Fig. 16B, and is specified to incorporate comments generated based on viewer feedback to fill in the gaps when an event is not detected for a predetermined time (90 seconds in this example) or more during event recognition during live distribution.

[0163] As shown in Figure 17A, this operation is similar to the operation shown in Figure 16B, and step S601 is executed when it is determined in step S508 that no event data has been input (NO in step S508), or when it is determined in step S510 that the event data is not that of a high-priority event (NO in step S510).

[0164] In step S601, the comment generating unit 40 determines whether or not a predetermined time (e.g., 90 seconds) has elapsed since the time when the most recent event data was input. If the predetermined time has not elapsed (NO in step S601), the comment generating unit 40 proceeds directly to step S530. On the other hand, if the predetermined time has elapsed (YES in step S601), the comment generating unit 40 executes viewer feedback (step S610), and then proceeds to step S530.

[0165] 1.11.1 Audience Feedback Examples In this example, viewer feedback is assumed to be a function that allows each viewer (which may include the poster (the player himself / herself)) to send comments via chat on the distribution service, etc. Fig. 17B is a flowchart showing an example of the operation flow of viewer feedback processing.

[0166] As shown in FIG. 17B, the comment generation unit 40 first acquires feedback such as chat messages sent from each viewer within a predetermined time period in the past (90 seconds in this example) (step S6101). Next, the comment generation unit 40 recognizes keywords, content, etc. of the acquired feedback using a recognizer based on machine learning or the like, and, based on the recognition results, tags the acquired feedback with tags such as "support," "disappointment," and "jeer" (step S6102) and recognizes the feedback target (step S6103). For example, if the feedback from a viewer is "Don't worry, do your best," the comment generation unit 40 recognizes that the tag to be assigned is "support" and the target (feedback target) is the player; if the feedback is "Yay, I'm so happy," the tag is "joy" and the viewer himself; and if the feedback is "Why don't you stop watching now?" the tag is "disappointment" and the target is another viewer.

[0167] Next, the comment generating unit 40 determines whether there is a feedback target that matches the target set as the comment target (the user, a team member, a friend, etc.) (step S6104), and if there is no matching target (NO in step S6104), returns to the operation shown in Fig. 17A. On the other hand, if there is a matching target (YES in step S6104), the comment generating unit 40 extracts feedback whose feedback target matches the comment target from the acquired feedback, and identifies the most common tag among the tags attached to the extracted feedback (step S6105).

[0168] Next, the comment generation unit 40 acquires a comment list for which the position is "commentator" from the position comment list shown in Fig. 7B, for example, and extracts comments tagged with the tag specified in step S6105 from the acquired comment list to create a list (comment list) of those comments (step S6106). By narrowing down the comments of the virtual commentator to comments with the same tag as feedback from viewers, it becomes possible to have the virtual commentator utter comments that resonate with viewers.

[0169] Next, the comment generation unit 40 performs operations similar to steps S5204 to S5209 in FIG. 16D to generate a position comment that the virtual commentator will actually utter (steps S6107 to S6113), and then adds the current time code of the video being live streamed to this position comment and registers it in the target comment list (step S6114).

[0170] Next, in order to control the event intervals in steps S510 and S601, the comment generating unit 40 counts the position comment registered in the target comment list in step S6114 as one of the events (step S6115), and returns to the operation shown in FIG. 17A.

[0171] 1.12 Example of virtual commentators increasing and decreasing during real-time video viewing 16A to 16F show an example in which two virtual commentators generate interactive comments for a live broadcast. In contrast, this example describes an example in which the number of virtual commentators increases or decreases while watching a live broadcast (hereinafter referred to as live viewing). Here, the increase or decrease in the number of virtual commentators refers to an environment in which players and viewers change in real time, such as an online competitive game. Each viewer's user terminal displays not only the virtual commentators set for that viewer but also the virtual commentators set for other viewers. Therefore, the number of virtual commentators increases or decreases, or the virtual commentators are replaced, depending on the increase or decrease or replacement of other viewers. This configuration can create a sense of excitement in the match and generate a dialogue between the virtual commentators. Furthermore, when displaying not only the viewer's virtual commentators but also the player's virtual commentators, the number of virtual commentators may increase or decrease, or the virtual commentators may be replaced, depending on the increase or decrease or replacement of players.

[0172] 18 is a flowchart showing an example of an operation flow when the number of virtual commentators increases or decreases during live distribution. Note that when a virtual commentator is replaced, it can be treated as a case where a decrease and an increase in the number of virtual commentators occur simultaneously, so detailed explanation will be omitted here.

[0173] 18, in this operation, first, when viewer A10 starts an application on the user terminal 110 and starts live viewing of a video, a request for live viewing is notified to the setting unit 120. Upon receiving this request, the setting unit 120 first obtains user information about viewer A10 from the user information storage unit 20 (step S701).

[0174] Meanwhile, a selection screen for selecting a video genre to be viewed live is displayed on the user terminal 110. When the viewer A10 selects a video genre in accordance with this selection screen (step S702), the selected video genre is notified from the user terminal 110 to the setting unit 120. Note that the selection screen may be in various formats, such as an icon format or a menu format.

[0175] Next, the setting unit 120 automatically selects a virtual commentator character based on, for example, the video genre notified by the user terminal 110 and the user information of the viewer A10 acquired from the user information storage unit 20, and sets the character's position as "viewer" (step S703).

[0176] Next, similar to step S111 in FIG. 6A, the event extraction unit 30 acquires a recognition model based on the video genre of the video content (step S704).

[0177] Once preparations for live viewing are complete in this manner, distribution of live video begins from the distribution unit 80 to the user terminal 110 (step S705). When live viewing begins, live video (video content) is sequentially output from the distribution unit 80 to the user terminal 110. At this time, the video content may be transmitted by streaming.

[0178] The number of viewers participating in the live distribution, user information of the players and viewers, etc. are managed, for example, by the setting unit 120. The setting unit 120 manages the number of viewers currently watching the live distribution and determines whether the number of viewers has increased or decreased (step S706). If there has been no increase or decrease (NO in step S706), the video proceeds to step S709.

[0179] If the number of viewers increases or decreases (YES in step S706), the setting unit 120 adjusts the display of the virtual commentators (step S707). Specifically, when a new viewer joins, the setting unit 120 adjusts the settings so that the virtual commentator of the newly joined viewer is additionally displayed in the live video (video content) distributed to viewer A10. Furthermore, when a viewer leaves the live broadcast, the setting unit 120 adjusts the settings so that the virtual commentator of the viewer who left is not displayed in the live video for viewer A10. At this time, a specific animation may be executed, such as the virtual commentator of the viewer who left entering or exiting through a door that appears on the screen. For example, if viewer A10 is cheering for team A, a virtual commentator for the opposing team B may appear, and the two teams may engage in a cheering battle, which can liven up the competition. Furthermore, if a friend of viewer A10 joins the live broadcast, the friend's virtual commentator may appear, which can liven up the cheering and the competition.

[0180] Next, the setting unit 120 adjusts the comment targets of the virtual commentators related to the viewer A10 based on the increase or decrease in the number of virtual commentators (step S708). For example, if the number of virtual commentators increases, the increased virtual commentators are added to the comment targets, and if the number of virtual commentators decreases, the decreased virtual commentators are deleted from the comment targets. Then, the setting unit 120 sequentially inputs the adjusted virtual commentators and their comment targets to the event extraction unit 30 and the avatar generation unit 60.

[0181] Next, the comment generation unit 40 acquires the position comment generated about another viewer's virtual commentator (hereinafter, for convenience of explanation, referred to as virtual commentator B) (step S709), and determines whether the target (listener) of this position comment is viewer A10's virtual commentator (hereinafter, for convenience of explanation, referred to as virtual commentator A) (step S710).

[0182] If virtual commentator A is the target (YES in step S710), the comment generation unit 40 regards the target (listener) of the comment made by virtual commentator A as virtual commentator B who spoke to him (step S711), generates a positional comment for virtual commentator B (step S712), and proceeds to step S715. On the other hand, if virtual commentator A is not the target (NO in step S710), the comment generation unit 40 regards the target of the comment as the viewer (step S713), generates a positional comment for the viewer on an event basis (step S714), and proceeds to step S715.

[0183] In steps S715 to S718, similar to steps S530 to S550 and S512 in FIG. 16B, speech control processing (S715), avatar generation processing (S716), and editing / rendering processing (S717) are executed, and the video content with comments generated thereby is live-streamed to viewer A10 and other viewers (step S718).

[0184] Thereafter, for example, a control unit (not shown) in the cloud 100 determines whether or not to end the distribution (step S719), and if not (NO in step S719), this operation returns to step S706. On the other hand, if it is to be ended (YES in step S719), this operation ends.

[0185] In this example, if no event occurs for a predetermined period (e.g., 90 seconds), virtual commentator A may actively speak to other commentators, for example by generating and inserting a comment addressing another virtual commentator B. Regarding the timing of speeches by multiple virtual commentators, the example shown in FIG. 16B shows a rule that the virtual commentator who is the announcer speaks first, followed by the virtual commentator who is the analyst. However, in this example, to promote excitement, multiple virtual commentators may be allowed to speak simultaneously. In this case, measures may be taken, such as relatively lowering the volume of the voices of virtual commentators other than virtual commentator A or separating the sound locations, to make it easier for viewers to hear what each virtual commentator is saying.

[0186] 1.13 Example of adjusting the display position according to the virtual commentator's position Furthermore, the display position of the virtual commentator on the user terminal 10 / 110 may be adjusted according to the position of the virtual commentator. In other words, a virtual position (hereinafter referred to as a virtual position) may be set for the virtual commentator according to the position of the virtual commentator.

[0187] For example, if the video content is a two-dimensional image, the virtual position may be within or outside the area of ​​that image. Furthermore, if the video content is a three-dimensional image, the virtual position may be within a 3D space or within the area of ​​a two-dimensional image superimposed on the three-dimensional image. Furthermore, the virtual commentator does not have to be visually displayed. In this case, the presence of the virtual commentator may be conveyed to the viewer (including the user) by, for example, sound source localization when rendering the voice of the virtual commentator.

[0188] For example, if the virtual commentator is standing as a "friend watching the game together," the virtual commentator's virtual position is assumed to be next to the viewer. In this case, the virtual position can be set to be on either the left or right side of the viewer, at the same distance from the video content (specifically, the user terminal 10 / 110) as the viewer, facing basically toward the content, and facing the viewer when speaking to the viewer.

[0189] Furthermore, if the virtual commentator is the person in question, they may be positioned next to the viewer to create the feeling that they are experiencing the content together, or they may be positioned as a presenter facing the viewer from the side of the content.

[0190] Once such a virtual position has been determined, it is possible to give viewers a sense of the virtual position, including its location, orientation, and changes over time, by drawing the virtual commentator in 2D / 3D or localizing the sound source of the virtual commentator in three-dimensional space. For example, it is possible to express the difference in sound between speaking directly to the viewer and speaking in a direction 90 degrees away from the viewer, the difference in sound between speaking from a distance and speaking close to the viewer's ear, and the appearance of the virtual commentator approaching / moving away.

[0191] Such adjustment of the display position (virtual position) according to the virtual commentator's standing position can be realized, for example, by the avatar generation unit 60 adjusting the display position of the generated avatar animation in the video content according to the virtual commentator's standing position. Also, adjustment of the virtual commentator's orientation according to the virtual commentator's standing position can be realized, for example, by the avatar generation unit 60 adjusting the orientation of the virtual commentator when generating the avatar animation according to the virtual commentator's standing position.

[0192] Figures 19A to 19C are figures for explaining examples of adjusting the display position depending on the position of the virtual commentator, where Figure 19A shows an example of the virtual position when the virtual commentator is in the position of a "friend watching together" and speaking to the video content, Figure 19B shows a case where the virtual commentator is in the position of a "friend watching together" and speaking to the viewers, and Figure 19C shows a case where the virtual commentator is in the position of a "commentator" and speaking to the viewers.

[0193] As shown in Figure 19A, for example, when the virtual commentator is positioned as a "friend watching together" and speaking to the video content, the virtual commentator C1 may be positioned closer to the viewer (front side) in the video content G50, and may face the event in the video content G50 (in this example, the character controlled by user U1).

[0194] Also, as shown in Figure 19B, for example, if the virtual commentator is standing in the position of a "friend watching the game together" and speaking to the viewer, the virtual commentator C1 may be positioned on the side closer to the viewer (near side) in the video content G50, and may be facing the viewer.

[0195] As shown in Figure 19C, for example, if the virtual commentator is positioned as a "commentator" and speaking to the viewers, the virtual commentator C1 may be positioned on the side of the video content G50 that is far from the viewers (at the back), and may face the event or the viewers.

[0196] In this way, by controlling the display position and orientation of the virtual commentator depending on the position of the virtual commentator and the subject of the comment, it is possible to improve the sense of realism of the content experience for the viewer.

[0197] 1.14 Example of comment rules when multiple virtual commentators exist Furthermore, it is expected that the content of the virtual commentator's comments will refer to a specific location within a video frame, such as "This player's movement is amazing." However, since the area that humans can focus on is limited relative to the video frame, if the viewer's gaze moves significantly, it may be difficult for the viewer to understand. To address this issue, for example, if there are multiple virtual commentators, it is possible to set a rule that if one virtual commentator comments on a specific location within a video frame, another virtual commentator will not comment on any location that does not fit within the central field of view of that location within a certain time period.

[0198] 1.15 Example of commenting using gaze information In addition, the point of gaze on the screen may be detected from the poster's gaze information and a virtual commentator may be made to comment on that part, thereby guiding the viewer's attention in the direction desired by the poster, or conversely, the point of gaze on the screen may be detected from the viewer's gaze information and a virtual commentator may be made to make a comment that is tailored to that part.

[0199] 1.16 Examples of caption display positions Furthermore, captions for audio comments made by virtual commentators may be superimposed on the video content. In this case, for example, by adjusting the display position of the captions based on gaze information of the poster or viewer, it is possible to reduce the viewer's gaze movement and improve comprehension. Specifically, by displaying the captions near the viewer or the area in the video content that the viewer is gazing at, it is possible to visually link the comment with the area that the comment targets, thereby reducing the viewer's gaze movement and improving comprehension.

[0200] Figure 20A shows an example where the caption for the audio comment is placed in a basic position (e.g., bottom center of the screen), Figure 20B shows an example where the position of the caption for the audio comment is adjusted based on the poster's gaze information, and Figure 20C shows an example where a standing position comment is generated based on the viewer's gaze information and the display position of the caption for the audio comment is adjusted.

[0201] As shown in Figure 20A, if captions are displayed in a predetermined position without considering where the poster or viewer is looking, it will be difficult to easily identify which player on the screen the virtual commentator is commenting on, reducing clarity.

[0202] In contrast, as shown in Figure 20B, for example, if the display position of the caption is adjusted based on the gaze information of the poster (or viewer), it becomes possible to display the caption near the subject to which the virtual commentator is making a comment, making it possible for the viewer to easily recognize which part of the screen (the player in the example shown) the virtual commentator is commenting on.

[0203] 20C, by generating a comment based on where the viewer is looking on the screen (i.e., on the video content), it becomes possible for the virtual commentator to make timely comments about the event the viewer is watching. Also, by adjusting the display position of the caption of the audio comment based on the viewer's line of sight, it becomes possible for the viewer to easily recognize which part of the screen (a player in the example shown) the virtual commentator is commenting on.

[0204] Note that the gaze information of the poster or viewer may be acquired using, for example, a gaze detection sensor or a camera provided in the user terminal 10 / 110. In other words, the user terminal 10 / 110 may also function as an acquisition unit that acquires the gaze information of the poster or viewer.

[0205] 1.17 Example of applying machine learning to generate comments for each position of a virtual commentator Next, we will explain an example of applying machine learning to generate comments for each position of a virtual commentator. Recently, video streaming has become commonplace, with the number of streams of game videos in particular experiencing a rapid increase. On the main streaming platforms, comments by game players, viewers, and commentators at e-SPORTS (registered trademark) tournaments are exchanged in real time during video streaming, and this trend is expected to continue to grow, even in categories other than gaming, as a way to enjoy video interactively.

[0206] This diversification of commenters corresponds to the positions of virtual commentators. For example, in a game commentary, the comments from the "player" position in Figure 1 are comments from the player himself, the comments from the "team members" position are comments from other players playing with the player, and the comments from the "viewers (friends)" position can be thought of as comments from viewers who are not directly involved in the gameplay. Each can be acquired independently based on differences in microphone and comment ID. In addition, comments from various positions such as "enemy (opponent)," "commentator," "analyst," and "host" can also be acquired independently.

[0207] In other words, the number of video streams with comments from various virtual commentator positions is rapidly increasing, and as a result, a large number of comments from various positions are being acquired independently for each position. This has made it possible to create a large-scale comment data set for each virtual commentator position (for example, the position comment list in Figure 7B).

[0208] Although the volume of comments in large-scale comment datasets is smaller than that of general-purpose language models based on other ultra-large datasets, by applying optimization techniques such as fine-tuning and N-Shot Learning to general-purpose language models trained on datasets for each position, it is possible to construct language models for each position that use both the general-purpose language model and large-scale position datasets.

[0209] Fig. 21 is a block diagram showing an example of a system configuration when machine learning is applied to generate comments for each virtual commentator's position. As shown in Fig. 21, the content of statements made by players 1 to 3, commentators, and analysts who are game participants is converted into text data by speech recognition 211 to 213, 215, and 216, and stored in position comment groups 221 to 226 for each position. In addition, for chat 204 by viewers, the input text data is stored in viewer (friend) position comment group 224.

[0210] The general-purpose language models 231-236 prepared for the players 1-3, the viewer chat 204, the commentators and the analysts are trained using the comment data sets registered in the position comment groups 221-226 of the respective positions as training data, thereby creating the language models (position comment lists) 241-246 for the respective positions.

[0211] The comment generation unit 40 in the information processing system 1 according to this embodiment can generate appropriate position comments according to the positions of various virtual commentators by using the language models (position comment lists) 241 to 246 for each position.

[0212] Although the above explanations have focused on examples in which comment audio and avatar animation are placed in video content, the processes explained based on each flowchart may be performed for generating and adding text (caption) comments, or for both audio comments and text comments. Also, only text comments and / or audio comments may be added to video content without adding avatar animation.

[0213] 2. System configuration example At least some of the setting unit 120, event extraction unit 30, comment generation unit 40, speech control unit 50, avatar generation unit 60, editing and rendering unit 70, and distribution unit 80 according to the above-described embodiments may be realized in the user terminal 10 / 110, with the remainder realized in one or more information processing devices such as a cloud server on a network, or all of them may be realized in a cloud server on a network, etc. For example, the setting unit 120 may be realized in the user terminal 10 / 110, and the event extraction unit 30, comment generation unit 40, speech control unit 50, avatar generation unit 60, editing and rendering unit 70, and distribution unit 80 may be realized in a cloud server on a network, etc.

[0214] 3. Hardware configuration One or more information processing devices that execute at least one of the user terminal 10 / 110, setting unit 120, event extraction unit 30, comment generation unit 40, speech control unit 50, avatar generation unit 60, editing / rendering unit 70, and distribution unit 80 according to the above-described embodiments can be realized by a computer 1000 configured as shown in Fig. 22. Fig. 22 is a hardware configuration diagram showing an example of the computer 1000 that realizes the functions of the user terminal 10 / 110 and the information processing device.

[0215] 22, a computer 1000 includes a CPU 1001, a ROM (Read Only Memory) 1002, a RAM (Random Access Memory) 1003, a sensor input unit 1101, an operation unit 1102, a display unit 1103, an audio output unit 1104, a storage unit 1105, and a communication unit 1106. The CPU 1001, the ROM 1002, and the RAM 1003 are connected to one another via an internal bus 1010. The sensor input unit 1101, the operation unit 1102, the display unit 1103, the audio output unit 1104, the storage unit 1105, and the communication unit 1106 are connected to the internal bus 1010 via an input / output interface 1100.

[0216] The CPU 1001 operates and controls each unit based on programs stored in the ROM 1002 or the storage unit 1105. For example, the CPU 1001 loads the programs stored in the ROM 1002 or the storage unit 1105 into the RAM 1003 and executes processing corresponding to the various programs.

[0217] The ROM 1002 stores boot programs such as a BIOS (Basic Input Output System) that is executed by the CPU 1001 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .

[0218] The storage unit 1105 is a computer-readable recording medium that non-temporarily records programs executed by the CPU 1001 and data used by such programs, etc. Specifically, the storage unit 1105 is a recording medium that records programs for executing each operation according to the present disclosure.

[0219] The communication unit 1106 is an interface for connecting the computer 1000 to an external network (e.g., the Internet). For example, the CPU 1001 receives data from other devices and transmits data generated by the CPU 1001 to other devices via the communication unit 1106.

[0220] The sensor input unit 1101 includes, for example, a gaze detection sensor or camera that detects the gaze of the poster, viewer, etc., and generates gaze information of the poster, viewer, etc. from the acquired sensor information. In addition, if the user terminal 10 / 110 is, for example, a game console, the sensor input unit 1101 may include an IMU (Inertial Measurement Unit), microphone, camera, etc. that are provided on the game console or its controller.

[0221] The operation unit 1102 is an input device such as a keyboard, mouse, touchpad, touch panel, or controller, which allows a contributor or viewer to input operation information.

[0222] The display unit 1103 is a display that displays game screens, video content, etc. The display unit 1103 can display various selection screens, for example, as shown in Figures 2A to 2D.

[0223] The audio output unit 1104 is configured by, for example, a speaker, and outputs the audio of games and video content, audio comments made by virtual commentators within the video content, and the like.

[0224] For example, when the computer 1000 functions as one or more of the user terminal 10 / 110, setting unit 120, event extraction unit 30, comment generation unit 40, speech control unit 50, avatar generation unit 60, editing / rendering unit 70, and distribution unit 80 according to the above-described embodiments, the CPU 1001 of the computer 1000 executes a program loaded onto the RAM 1003 to realize the functions of the corresponding units. The storage unit 1105 stores programs according to the present disclosure. The CPU 1001 reads and executes the programs from the storage unit 1105. Alternatively, the CPU 1001 may acquire these programs from another device on a network via the communication unit 1106.

[0225] Although the embodiments of the present disclosure have been described above, the technical scope of the present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure. Furthermore, components of different embodiments and modifications may be combined as appropriate.

[0226] Furthermore, the effects of each embodiment described in this specification are merely examples and are not intended to be limiting, and other effects may also be obtained.

[0227] Furthermore, each of the above-described embodiments may be used alone or in combination with other embodiments.

[0228] The present technology can also be configured as follows. (1) an acquisition unit that acquires information about a relationship between a contributor of content and a viewer of the content; a comment generating unit that generates a comment to be uttered by a virtual commentator based on the information regarding the relationship; An information processing device. (2) The information about the relationship includes at least one of the degree of intimacy between the poster and the viewer, the relationship between the poster and the viewer in the content, and history information of the viewer regarding content posted in the past by the poster. The information processing device according to (1) above. (3) The information about the relationship includes at least one of the degree of intimacy between the poster and the viewer, the relationship between the poster and the viewer in the content, and history information of the viewer regarding content posted in the past by the poster. The information processing device according to (1) or (2). (4) the acquisition unit sets a position of the virtual commentator; The comment generating unit generates the comment according to the position. The information processing device according to any one of (1) to (3) above. (5) The comment generating unit generates the comment based on a position comment list that lists candidate comments to be uttered by the virtual commentator for each position. The information processing device according to (4) above. (6) The comment generating unit generates the comment based on a comment list obtained by excluding a predetermined number of comments from the position comment list among the comments that the virtual commentator has previously made. The information processing device according to (5) above. (7) The acquisition unit allows the poster or the viewer to select the position of the virtual commentator. The information processing device according to any one of (4) to (6) above. (8) The acquisition unit automatically sets the position of the virtual commentator based on the information about the relationship. The information processing device according to any one of (4) to (6) above. (9) The acquisition unit sets a target of a comment to be made by the virtual commentator, The comment generating unit generates the comment according to a target of the comment. The information processing device according to any one of (1) to (8) above. (10) The acquisition unit allows the poster or the viewer to select a target of the comment. The information processing device according to (9) above. (11) The acquisition unit automatically sets a target of the comment based on the information about the relationship. The information processing device according to (9) above. (12) The comment generating unit generates the comment according to a genre to which the content belongs. The information processing device according to any one of (1) to (11) above. (13) The comment generating unit modifies the generated comment based on the information about the relationship. The information processing device according to any one of (1) to (12) above. (14) The comment generating unit modifies the ending of the generated comment based on the hierarchical relationship between the poster and the viewer. The information processing device according to (13) above. (15) further comprising an extraction unit that extracts an event of the content; The comment generating unit generates the comment for the event. The information processing device according to any one of (1) to (14) above. (16) The comment generating unit skips generating a comment for the next event when a time difference between an occurrence time of a previous event and an occurrence time of a next event in the content is less than a predetermined time. The information processing device according to (15) above. (17) When a time difference between an occurrence time of a previous event and an occurrence time of a next event in the content is less than a predetermined time and the priority of the next event is higher than the priority of the previous event, the comment generating unit generates a comment for the next event and requests the user to stop uttering the comment generated for the previous event. The information processing device according to (15) above. (18) The comment generation unit generates comments to be spoken by each of the two or more virtual commentators. The information processing device according to any one of (1) to (17) above. (19) The comment generating unit generates the comment so that a second virtual commentator among the two or more virtual commentators will make a statement after a first virtual commentator among the two or more virtual commentators has completed making a statement. The information processing device according to (18) above. (20) The comment generating unit generates the comment that causes one of the two or more virtual commentators to make a statement directed to another of the two or more virtual commentators. The information processing device according to (18) or (19). (twenty one) the acquisition unit acquires the number of viewers currently viewing the content; The comment generation unit generates the comments to be spoken by each of the virtual commentators, the number of which corresponds to the number of the viewers, and increases or decreases the number of the virtual commentators according to an increase or decrease in the number of the viewers. The information processing device according to any one of (1) to (20) above. (twenty two) The comment generation unit acquires feedback from the viewer and generates the comment in response to the feedback. The information processing device according to any one of (1) to (21) above. (twenty three) The content editing and rendering unit further includes: an editing and rendering unit that incorporates at least one of character data and audio data corresponding to the comment into the content; The information processing device according to any one of (1) to (22) above. (twenty four) further comprising an animation generation unit that generates an animation of the virtual commentator; The editing and rendering unit superimposes animation of the virtual commentator on the content. The information processing device according to (23). (twenty five) The editing and rendering unit adjusts the position of the animation in the content according to the position of the virtual commentator. The information processing device according to (24). (26) a speech control unit for converting the comment into the voice data; The information processing device according to any one of (23) to (25). (27) a speech control unit that converts the comment into voice data; an editing and rendering unit that incorporates the audio data into the content; an animation generation unit that generates an animation of the virtual commentator; Furthermore, The acquisition unit sets a target of a comment to be made by the virtual commentator, the animation generation unit generates the animation in which the orientation of the virtual commentator is adjusted according to a target of the comment; The editing and rendering unit superimposes animation of the virtual commentator on the content. The information processing device according to any one of (9) to (11) above. (28) The acquisition unit acquires line-of-sight information of the contributor or the viewer, The comment generating unit generates the comment based on the line-of-sight information. The information processing device according to any one of (1) to (27) above. (29) The comment generating unit adjusts a display position of the comment in the content based on the line-of-sight information. The information processing device according to (28). (30) Obtaining information regarding a relationship between a contributor of content and a viewer of said content; A comment to be made by the virtual commentator is generated based on the information about the relationship. An information processing method including: (31) An information processing system in which a first user terminal, an information processing device, and a second user terminal are connected via a predetermined network, The information processing device includes: an acquisition unit that acquires information regarding a relationship between a contributor who posts content from the first user terminal to the information processing device and a viewer who views the content via the second user terminal; a comment generating unit that generates a comment to be uttered by a virtual commentator based on the information regarding the relationship; Equipped with an information processing system. [Explanation of symbols]

[0229] 1. Information Processing Systems 10, 110, M1 user terminal 20 User information storage unit 30 Event Extraction Unit 31 Image analysis unit 32 Audio analysis section 40 Comment Generation 41 Position and target control section 50 Speech control unit 60 Avatar Generation Unit 70 Editing and Rendering Department 80 Distribution Department 90 Network 100 Cloud 120 Setting section 1001 CPU 1002 ROM 1003 RAM 1010 Internal Bus 1100 Input / Output Interface 1101 Sensor input section 1102 Operation unit 1103 Display section 1104 Audio output unit 1105 Storage section 1106 Communications Department A10, A11, A12, A13 viewers A100 Viewer side U1 user U100 User side

Claims

1. An acquisition unit that acquires information regarding a relationship between a contributor of content and a viewer of the content; an extraction unit that extracts an event of the content; a comment generating unit that generates a comment for the event, the comment being made by a virtual commentator based on information about the relationship; The comment generating unit skips generating a comment for the next event when a time difference between an occurrence time of a previous event and an occurrence time of a next event in the content is less than a predetermined time. Information processing device.

2. An acquisition unit that acquires information regarding a relationship between a contributor of content and a viewer of the content; a comment generating unit that generates a comment to be uttered by a virtual commentator based on the information about the relationship; the acquisition unit acquires the number of viewers currently viewing the content; The comment generation unit generates the comments to be spoken by each of the virtual commentators, the number of which corresponds to the number of viewers, and increases or decreases the number of the virtual commentators according to an increase or decrease in the number of viewers. Information processing device.

3. An acquisition unit that acquires information regarding a relationship between a contributor of content and a viewer of the content; a comment generating unit that generates a comment to be uttered by a virtual commentator based on the information regarding the relationship; an animation generation unit that generates an animation of the virtual commentator; an editing and rendering unit that incorporates at least one of character data and audio data corresponding to the comment into the content and superimposes an animation of the virtual commentator on the content; An information processing device comprising:

4. An acquisition unit that acquires information regarding a relationship between a contributor of content and a viewer of the content; a comment generating unit that generates a comment to be uttered by a virtual commentator based on the information about the relationship; The acquisition unit acquires line-of-sight information of the contributor or the viewer, The comment generating unit generates the comment based on the line-of-sight information. Information processing device.

5. further comprising an extraction unit that extracts an event of the content, the comment generation unit generates the comment for the event; The comment generating unit skips generating a comment for the next event when a time difference between an occurrence time of a previous event and an occurrence time of a next event in the content is less than a predetermined time. The information processing device according to any one of claims 2 to 4.

6. the acquisition unit acquires the number of viewers currently viewing the content; The comment generation unit generates the comments to be spoken by each of the virtual commentators, the number of which corresponds to the number of viewers, and increases or decreases the number of the virtual commentators according to an increase or decrease in the number of viewers.

5. The information processing device according to claim 1, 3 or 4.

7. an animation generation unit that generates an animation of the virtual commentator; an editing and rendering unit that incorporates at least one of character data and audio data corresponding to the comment into the content and superimposes an animation of the virtual commentator on the content; The information processing device according to claim 1 , further comprising:

8. The editing and rendering unit adjusts the position of the animation in the content according to the position of the virtual commentator. The information processing device according to claim 3 or 7.

9. The acquisition unit acquires line-of-sight information of the contributor or the viewer, The comment generating unit generates the comment based on the line-of-sight information. The information processing device according to any one of claims 1 to 3.

10. The comment generating unit adjusts a display position of the comment in the content based on the line-of-sight information. The information processing device according to claim 4 or 9.

11. The comment generating unit modifies the generated comment based on the information about the relationship. The information processing device according to any one of claims 1 to 10.

12. The comment generation unit acquires feedback from the viewer and generates the comment in response to the feedback. The information processing device according to any one of claims 1 to 11.

13. An information processing method executed by an information processing device, comprising: an acquisition step of acquiring information about a relationship between a contributor of content and a viewer of the content; an extraction step of extracting an event from the content; a comment generating step of generating a comment for the event, the comment being made to be uttered by a virtual commentator based on information about the relationship; In the comment generation step, if the time difference between the occurrence time of the previous event and the occurrence time of the next event in the content is less than a predetermined time, generation of a comment for the next event is skipped. Information processing methods.

14. An information processing method executed by an information processing device, comprising: an acquisition step of acquiring information about a relationship between a contributor of content and a viewer of the content; a comment generating step of generating a comment to be uttered by a virtual commentator based on the information about the relationship; In the acquiring step, the number of viewers currently viewing the content is acquired; In the comment generating step, the comments to be uttered by each of the virtual commentators, the number of which corresponds to the number of the viewers, are generated, and the number of the virtual commentators is increased or decreased according to an increase or decrease in the number of the viewers. Information processing methods.

15. An information processing method executed by an information processing device, comprising: Obtaining information regarding a relationship between a contributor of content and a viewer of said content; generating a comment to be uttered by a virtual commentator based on the information about the relationship; generating an animation of the virtual commentator; At least one of character data and audio data corresponding to the comment is incorporated into the content, and an animation of the virtual commentator is superimposed on the content. Information processing methods.

16. An information processing method executed by an information processing device, comprising: an acquisition step of acquiring information about a relationship between a contributor of content and a viewer of the content; a comment generating step of generating a comment to be uttered by a virtual commentator based on the information about the relationship; In the acquiring step, gaze information of the contributor or the viewer is acquired, In the comment generating step, the comment is generated based on the line-of-sight information. Information processing methods.

17. An information processing system in which a first user terminal, an information processing device, and a second user terminal are connected via a predetermined network, The information processing device includes: an acquisition unit that acquires information regarding a relationship between a contributor who posts content from the first user terminal to the information processing device and a viewer who views the content via the second user terminal; an extraction unit that extracts an event of the content; a comment generating unit that generates a comment for the event, the comment being made by a virtual commentator based on information about the relationship; The comment generating unit skips generating a comment for the next event when a time difference between an occurrence time of a previous event and an occurrence time of a next event in the content is less than a predetermined time. Information processing system.

18. An information processing system in which a first user terminal, an information processing device, and a second user terminal are connected via a predetermined network, The information processing device includes: an acquisition unit that acquires information regarding a relationship between a contributor who posts content from the first user terminal to the information processing device and a viewer who views the content via the second user terminal; a comment generating unit that generates a comment to be uttered by a virtual commentator based on the information about the relationship; the acquisition unit acquires the number of viewers currently viewing the content; The comment generation unit generates the comments to be spoken by each of the virtual commentators, the number of which corresponds to the number of viewers, and increases or decreases the number of the virtual commentators according to an increase or decrease in the number of viewers. Information processing system.

19. An information processing system in which a first user terminal, an information processing device, and a second user terminal are connected via a predetermined network, The information processing device includes: an acquisition unit that acquires information regarding a relationship between a contributor who posts content from the first user terminal to the information processing device and a viewer who views the content via the second user terminal; a comment generating unit that generates a comment to be uttered by a virtual commentator based on the information regarding the relationship; an animation generation unit that generates an animation of the virtual commentator; an editing and rendering unit that incorporates at least one of character data and audio data corresponding to the comment into the content and superimposes an animation of the virtual commentator on the content; An information processing system comprising:

20. An information processing system in which a first user terminal, an information processing device, and a second user terminal are connected via a predetermined network, The information processing device includes: an acquisition unit that acquires information regarding a relationship between a contributor who posts content from the first user terminal to the information processing device and a viewer who views the content via the second user terminal; a comment generating unit that generates a comment to be uttered by a virtual commentator based on the information about the relationship; The acquisition unit acquires line-of-sight information of the contributor or the viewer, The comment generating unit generates the comment based on the line-of-sight information. Information processing system.

Citation Information

Patent Citations

  • Information management system, server device and program

    JP2011210157A

  • Game management device, game system, game management method, and program

    JP2013236929A

  • Virtual processing server, control method for virtual processing server, content distribution system, and application program for terminal

    JP2018174456A

  • Object control system and object control method

    JP2018187712A

  • System for providing automatic comment

    KR1020150145280A