Cross-platform co-viewing device
The cross-platform co-viewing device facilitates co-viewing across different platforms by converting SNS posts into metaverse expressions, improving the social presence and enjoyment of content viewing.
Patent Information
- Application Number
- JP2024110447
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-09
- Publication Date
- 2026-01-22
AI Technical Summary
Co-viewing experiences are limited to users within the same platform or service, preventing users in a metaverse from enjoying content together with users on different platforms.
A cross-platform co-viewing device that collects content-related posts from SNS users using a content-related post collection unit and converts these posts into expressions in a metaverse co-viewing space through an expression conversion unit, allowing metaverse users to experience social presence with SNS users.
Enables metaverse users to feel as if they are co-viewing content with SNS users, enhancing the social presence and enjoyment of the viewing experience.
Smart Images

Figure 2026010523000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a cross-platform co-viewing device, and more particularly to a cross-platform co-viewing device for providing services to users on a metaverse who view content. [Background technology]
[0002] Broadcasting has played a social role in creating communication between people through programs. One typical example of this is the co-viewing of broadcast programs (broadcast content). Co-viewing is generally defined as the act of watching the same content while communicating with others. Typical examples of co-viewing include watching a drama program with family on the living room TV, or watching a live sports event on a streaming service and posting support comments on social networking services (SNS) to share with other supporters.
[0003] Co-viewing is said to have the effect of enriching the viewing experience more than viewing alone. For example, research has shown that it increases the subjective enjoyment of the content and reduces feelings of loneliness. Furthermore, it is said that the degree to which the viewing experience is enriched by co-viewing increases with the strength of social presence, which is the sense of social connection felt with co-viewers. Co-viewing therefore plays an important role in human communication.
[0004] In recent years, co-viewing has also expanded into virtual spaces. For example, users of a microblog (e.g., X (Non-Patent Document 1)), a tweeting service on a social networking site (hereinafter referred to as "SNS users"), can communicate by tweeting (text) their impressions while watching a television program.
[0005] Furthermore, users (hereinafter referred to as metaverse users) on the metaverse (for example, VRChat (Non-Patent Document 2)) operate their own avatars to gather in front of a virtual content playback device such as a television set installed on the metaverse, and communicate via the avatars while watching videos (content). At this time, communication is carried out using voice chat and the behavior of the avatars. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] “X.com”, [online], [Retrieved June 7, 2024], Internet<URL:https: / / x.com / > [Non-patent document 2] “VRChat”, [online], [Retrieved June 7, 2024], Internet<URL:https: / / hello.vrchat.com / > Summary of the Invention [Problem to be solved by the invention]
[0007] However, currently, co-viewing is limited to communication between users who participate in the same platform or service. For example, a user watching a program on a metaverse can communicate with other users in the same co-viewing space, but cannot enjoy the program together with other users watching the same program on a platform outside the metaverse.
[0008] The present invention has been made in consideration of the above-mentioned problems, and its objective is to provide a cross-platform co-viewing device that allows metaverse users viewing content to feel as if they are co-viewing the same content with SNS users who are viewing the same content.
[0009] In order to solve the above problems, the cross-platform co-viewing device of the present invention is a cross-platform co-viewing device used for a service in which users of different platforms co-view the same content, and is equipped with a content-related post collection unit that collects text posting data from an SNS server that acquires, stores, and publishes text posting data from SNS users who use microblogging services, and extracts content-related posts, which are posting data related to the content of the same content, which are provided to content playback devices located in a metaverse co-viewing space, which is a virtual space, and in real space, and an expression conversion unit that brings the extracted content-related posts onto the metaverse service and, based on these content-related posts, converts the expression of the metaverse co-viewing space presented to metaverse users who view the content in the metaverse co-viewing space in a predetermined manner.
[0010] According to this configuration, the cross-platform co-viewing device uses a content-related post collection unit to collect content-related posts related to the same content as the content being viewed by a metaverse user from among the publicly posted data of SNS users. As a result, the content-related post collection unit can acquire data on reactions to the content posted by SNS users who view the same content on platforms outside the metaverse. The cross-platform co-viewing device then uses a representation conversion unit to convert the representation of the metaverse co-viewing space presented to the metaverse user in a predetermined manner based on the content-related posts brought onto the metaverse service. Here, the representation conversion method for the metaverse co-viewing space may be a method of converting the posted content into avatar behavior or a method of converting the number of posts into a presentation that allows the metaverse user to experience it. As a result, the representation conversion unit can represent the reactions of SNS users outside the metaverse to the content on the metaverse in a manner that makes it easy for the metaverse user to feel the social presence of the SNS user. [Effects of the Invention]
[0011] The present invention provides the following excellent effects. The cross-platform co-viewing device according to the present invention can give a metaverse user who is viewing content the feeling that he or she is viewing the content together with other SNS users who are viewing the same content. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a schematic configuration diagram of a co-viewing service providing system including a cross-platform co-viewing device according to a first embodiment of the present invention. [Figure 2] 1 is a block diagram showing functions of a cross-platform co-viewing device according to a first embodiment. FIG. [Figure 3] 1A is a conceptual diagram showing content viewing in a real space, and FIG. 1B is a conceptual diagram showing content viewing in a virtual space. [Figure 4] FIG. 10 is a sequence diagram showing the flow of content distribution in the shared viewing service providing system. [Figure 5] FIG. 10 is a sequence diagram showing a processing flow (part 1) of the cross-platform shared viewing device. [Figure 6] FIG. 10 is a sequence diagram showing a processing flow (part 2) of the cross-platform shared viewing device. [Figure 7] FIG. 10 is a block diagram showing functions of the cross-platform co-viewing device according to the second embodiment. [Figure 8] FIG. 11 is a block diagram showing functions of the cross-platform co-viewing device according to the third embodiment. [Figure 9] (a) is a conceptual diagram showing posts in real space, and (b) is a conceptual diagram showing examples of presentations according to the number of posts. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, an embodiment of the cross-platform co-viewing device according to the present invention will be described. (First embodiment) [Overview of the cross-platform shared viewing device] An overview of the cross-platform shared viewing device according to the first embodiment will be described with reference to FIGS. 1, 2, and 3. FIG. The cross-platform co-viewing device 100 is a device used for a service in which users of different platforms co-view the same content. As shown in Fig. 2, the cross-platform co-viewing device 100 includes a content-related post collection unit 110 and an expression conversion unit 120. The content-related post collection unit 110 collects post data from the SNS servers 18A-18N, which acquire, store, and publish text post data from the SNS users 2A-2N who use the microblog service. The content-related post collection unit 110 extracts content-related posts, which are post data related to the content of the same content provided to the content playback devices 14A-14N, 210 (see FIG. 3) located in the metaverse shared viewing space 200 (see FIG. 3), which is a virtual space, and the real spaces 30A-30N (see FIG. 3), respectively. The expression conversion unit 120 brings the extracted content-related posts onto the metaverse service, and based on these content-related posts, converts the expression of the metaverse shared viewing space 200 to be presented to the metaverse user 2X who watches the content in the metaverse shared viewing space 200 in a predetermined manner.
[0014] [Shared viewing service provision system] Next, a configuration example of a system for providing a shared viewing service including the cross-platform shared viewing device 100 will be described with reference to FIG. 1 (and also with reference to FIG. 3 as needed). The co-viewing service providing system 1 provides a service (co-viewing service) in which users of different platforms co-view the same content. As shown in Fig. 1, the co-viewing service providing system 1 includes a content server 12, content playback devices 14A to 14N, SNS servers 18A to 18N, user terminals 15A to 15N, a client terminal 21, a virtual space providing device 25, and a cross-platform co-viewing device 100.
[0015] Here, the shared-viewing service providing system 1 includes N SNS servers 18A to 18N corresponding to multiple types of SNSs A to N (hereinafter referred to as SNS_A to SNS_N). For the sake of simplicity, it is assumed that, for each SNS, one SNS user watches content on a content playback device and uses a user terminal. The number of each device in the shared-viewing service providing system 1 is arbitrary.
[0016] The content server 12 simultaneously distributes content to the real spaces 30A to 30N (see FIG. 3) and the metaverse shared viewing space 200 (see FIG. 3). Here, the content is, for example, broadcast content. The broadcast may be, for example, satellite broadcast, terrestrial broadcast, cable broadcast, or IP (Internet Protocol) multicast broadcast. It may also be online distribution (simultaneous online distribution, etc.) in which a digital broadcast television program is distributed via communication (Internet). The broadcast may be in the form of a live broadcast, such as a live sports program, or it may not be in the form of a live broadcast, but it is preferable that the content be distributed in real time with a known content start time. Note that the content is not limited to broadcast content.
[0017] Metaverse user 2X views content in metaverse shared-viewing space 200. Client terminal 21 is a terminal used by metaverse user 2X. For example, it is a PC (personal computer) or a dedicated terminal for using metaverse services. Note that a head-mounted display is not essential and may or may not be used. Metaverse user 2X uses client terminal 21 to view content using a virtual content playback device 210 in metaverse shared-viewing space 200, as shown in FIG. 3(b). Metaverse user 2X has a self-avatar 220 in metaverse shared-viewing space 200. The same content provided to content playback device 210 is provided to content playback devices 14A-14N located in real spaces 30A-30N.
[0018] The virtual space providing device 25 provides virtual space information to the client terminal 21 that is connected for communication via the network 3. As the virtual space information, various types of information for using general metaverse services can be used. The virtual space information includes, for example, information on virtual objects in the virtual space, location information, avatar information, background information, etc. Note that the operator that operates the virtual space providing device 25 and the operator that operates the cross-platform co-viewing device 100 may be different operators or may be the same operator.
[0019] SNS user 2A is a user who uses SNS_A and posts content. Content playback device 14A is a device that plays content distributed from content server 12. Content playback device 14A is, for example, a television, a mobile terminal, a personal computer, etc. SNS user 2A uses content playback device 14A to view content. User terminal 15A is a terminal that SNS user 2A operates to use SNS_A. User terminal 15A is, for example, a mobile terminal such as a smartphone, a personal computer, etc.
[0020] An application 16A, which is software for using SNS_A, has been downloaded to user terminal 15A. Application 16A provides an interface for SNS user 2A to post. Application 16A also transmits posted data to SNS server 18A. Note that the application is referred to as "app" in the figure. SNS server 18A provides the SNS_A service. SNS server 18A acquires, stores, and publishes text posted data from SNS user 2A. Note that an SNS may be any system in which a user posts comments (text) using a user terminal and the comments are published so that they can be viewed by other users.
[0021] The SNS user 2N is a user who uses the SNS_N and posts. The content playback device 14N plays back content distributed from the content server 12. The SNS user 2N uses the content playback device 14N to view the content. The user terminal 15N is a terminal that the SNS user 2N operates to use the SNS_N. An application 16N, which is software for using the SNS_N, is downloaded to the user terminal 15N. The application 16N provides an interface for the SNS user 2N to post. The application 16N also transmits the posted data to the SNS server 18N. The SNS server 18N provides the services of the SNS_N. The SNS server 18N acquires the text posted data from the SNS user 2N, stores it, and makes it public.
[0022] [Components of the cross-platform shared viewing device 100] Next, each unit of the cross-platform shared viewing device 100 will be described with reference to FIG. 2 (and also with reference to FIG. 1 and FIG. 3 as appropriate). In the following, it is assumed that the microblog service is a distributed SNS. That is, the SNS servers 18A to 18N are distributed SNS servers. The content-related post collection unit 110 collects post data from the SNS servers 18A to 18N via an API (Application Programming Interface).
[0023] A decentralized SNS is an SNS that has the following two characteristics: (1) It is operated by multiple independent servers. (2) It uses an open protocol. Decentralized SNSs based on the same protocol have a high degree of interoperability. For example, if SNS_A and SNS_N use the same protocol, data can be transferred between them, such as posting a comment on SNS_N using an SNS_A account. This is because the data vocabulary is shared. Due to this high degree of interoperability, a group of decentralized SNSs that use the same protocol can behave as if they were a single giant SNS, and are therefore also known as the Fediverse (see, for example, Reference 1 below). Reference 1: “Fediverse.Party - explore federated networks”, [online], [Retrieved June 7, 2024], Internet<URL:https: / / fediverse.party / >
[0024] Typical SNSs provide APIs that allow external services to obtain user posted data, but in the case of decentralized SNSs in particular, by supporting a single protocol, data can be obtained from other SNSs that also use the same protocol, and the content can be interpreted and used.
[0025] An example of a standard protocol used in decentralized SNS is ActivityPub (see, for example, Reference 2 below), which is standardized by the W3C (World Wide Web Consortium). Reference 2: “ActivityPub - W3C”, [online], [Retrieved June 7, 2024], Internet<URL:https: / / www.w3.org / TR / activitypub / >
[0026] Examples of decentralized SNSs that use ActivityPub include the microblogging services Mastodon (registered trademark, the same applies hereinafter; see Reference 3 below) and Misskey (registered trademark). Reference 3: “Mastodon - Decentralized social media”, [online], [Retrieved June 7, 2024], Internet<URL:https: / / joinmastodon.org / >
[0027] The content-related post collection unit 110 acquires, from a distributed SNS, reactions to content from users of platforms outside the metaverse. In the cross-platform shared viewing device 100, the content-related post collection unit 110, which is a means for acquiring data, is designed based on standard specifications and standards, and is therefore able to acquire data in a general manner without depending on a specific platform. As a result, the content-related post collection unit 110 acquires data on the reactions of SNS users outside the metaverse to content.
[0028] The content-related post collection unit 110 selects, from the collected post data, post data (content-related posts) related to the content currently being played in the metaverse co-viewing space 200 (see FIG. 3 ). The content-related post collection unit 110 sends the selected content-related posts to the emotion estimation unit 132, which will be described later. Here, the method by which the content-related post collection unit 110 selects content-related posts is arbitrary, and a desired method may be determined in advance by the provider, etc., of the cross-platform co-viewing device 100. As an example, the content-related post collection unit 110 may select content-related posts based on hashtags related to the content currently being played. Also, for example, the content-related post collection unit 110 may select replies to a specific post as content-related posts. Alternatively, the content-related post collection unit 110 may select posts sent to a specific server as content-related posts. Furthermore, an estimator may be created that estimates the degree of relevance to content from the posting time and post content, and the estimator may determine the content-related posts to be selected. In this case, the estimator may be located either inside or outside the content-related post collector 110 .
[0029] It is believed that communication in the metaverse makes it easier for users to feel a sense of social presence toward others who operate avatars. This is believed to be because both linguistic means, such as text, and non-verbal means, such as avatar gestures, are used. However, posts to SNSs are primarily text and do not include non-verbal communication elements such as avatar gestures. The expression conversion unit 120 of this embodiment estimates appropriate gestures to express the poster's emotions based on the content of the post to the decentralized SNS.
[0030] For this purpose, the expression conversion unit 120 includes a text-motion conversion unit 130, as shown in FIG. 2. The text-motion conversion unit 130 converts the text data of a content-related post into avatar motion. The text-motion conversion unit 130 includes an emotion estimation unit 132 and a motion generation unit 134. The emotion estimation unit 132 estimates the emotion of the SNS user 2A-2N who posted the post based on the content of the post data. The motion generation unit 134 generates avatar motion data corresponding to the estimated emotion and sends the motion data to the co-viewing avatars 230, 240 (see FIG. 3).
[0031] In this embodiment, the motion generation unit 134 generates co-viewing avatars 230, 240 (see FIG. 3) that correspond to the SNS users 2A, 2N who posted on the metaverse co-viewing space 200 (see FIG. 3). The co-viewing avatars are dummy avatars that represent other people. Here, the co-viewing avatar 230 is a co-viewing avatar that reflects the emotion estimated from the content posted by the SNS user 2A. The co-viewing avatar 240 is a co-viewing avatar that reflects the emotion estimated from the content posted by the SNS user 2N.
[0032] The motion generation unit 134 also sends the posted content (text data) to the co-viewing avatars 230 and 240 (see FIG. 3). As a result, the motion generation unit 134 makes the co-viewing avatar behave as if the poster were directly controlling the co-viewing avatar. 3(b), the co-viewing avatars 230 and 240 perform gestures according to the emotions estimated from the posted content. Note that the posted content may be displayed near the co-viewing avatars 230 and 240. At this time, as shown in FIG. 3(a), SNS users 2A and 2N viewing content in real spaces 30A and 30N post content related to their ongoing viewing of the content. As shown in FIG. 3(b), in metaverse co-viewing space 200, co-viewing avatars 230 and 240 perform gestures based on the posted content (text) corresponding to the content at the time of viewing, in accordance with the emotions of joy and sadness that can be discerned from the posted content (text). In other words, the actions of co-viewing avatars 230 and 240 reflect the reactions of SNS users 2A and 2N viewing the same content outside the metaverse. At the same time, metaverse user 2X is viewing the content at the time SNS users 2A and 2N viewed it. Therefore, metaverse user 2X also shares the emotions of joy and sadness that SNS users 2A and 2N felt when viewing the content, and experiences a sense of co-viewing. In this way, metaverse user 2X can feel a sense of social presence toward SNS users 2A and 2N outside the metaverse who posted the content.
[0033] In this embodiment, the emotion estimation unit 132 classifies the content of post data into multiple predetermined emotions using a learning model that has been trained to estimate the emotions contained in input text. Emotions may be classified as, for example, fear, relaxation, joy, surprise, happiness, sadness, boredom, neutrality, embarrassment, and anger. The emotion estimation unit 132 extracts the emotion of the SNS user who posted the post from the post data (text) based on a predetermined emotion classification. The emotion estimation unit 132 sends the estimation result of the emotion of the SNS user who posted the post to the motion generation unit 134.
[0034] In this embodiment, the motion generation unit 134 selects motion data corresponding to the emotion estimated by the emotion estimation unit 132 from among a plurality of pieces of motion data created in advance corresponding to a plurality of types of emotions. For example, the emotion of happiness (joy) can be associated with the gesture of raising one's arms up. For example, the emotion of sadness can be expressed by the gesture of holding one's head with both hands. Such motion programs for the jaesuture can be prepared in advance so that they can be selected. The motion generation unit 134 sends motion data to the corresponding co-viewing avatars 230 and 240 in the metaverse co-viewing space 200.
[0035] [Operation in the shared viewing service provision system] Next, the operation of the shared-viewing service providing system 1 will be described with reference to FIGS. 4 to 6 under the following two preconditions. <Prerequisite 1> N users (SNS users 2A to 2N) are viewing, in their respective viewing environments, videos distributed from the same content server 12. The SNS users 2A to 2N are each using a different distributed SNS (SNS_A to SNS_N). <Prerequisite 2> Metaverse user 2X is viewing the same content as SNS users 2A to 2N while operating his / her own avatar 220 in metaverse co-viewing space 200. Also present in metaverse co-viewing space 200 are co-viewing avatars 230 and 240 that act as agents for SNS users 2A to 2N. In the drawings and the following description, the operation of two representative SNSs (SNS_A, SNS_N) will be described.
[0036] As shown in FIG. 4, first, content server 12 simultaneously distributes content to the real space and the metaverse shared viewing space (step S11). Virtual space providing device 25 also provides virtual space information (step S12). Client terminal 21 used by metaverse user 2X accepts an operation of the metaverse user's avatar (step S13). Metaverse user 2X uses client terminal 21 to view the content (step S14). At the same time, SNS user 2A views the content played on content playback device 14A (step S14), and SNS user 2N views the content played on content playback device 14N (step S14).
[0037] 5, SNS user 2A operates application 16A of user terminal 15A to post content-related information (step S21). At this time, user terminal 15A connects to SNS server 18A that provides the SNS_A service via application 16A. Similarly, SNS user 2N operates application 16N of user terminal 15N to post content-related information (step S21). At this time, user terminal 15N connects to SNS server 18N that provides the SNS_N service via application 16N.
[0038] As shown in Figure 3(a), when the content is a real-time soccer broadcast program, SNS user 2A, who is cheering for the team attacking the goal, posts "Go!". On the other hand, SNS user 2N, who is cheering for the team that lost the goal, posts "We're done for!".
[0039] Then, following step S21, SNS server 18A aggregates posts using the SNS_A service and makes the posted data public (step S22). Similarly, SNS server 18N aggregates posts using the SNS_N service and makes the posted data public (step S22). Then, in the cross-platform co-viewing device 100, content-related post collection unit 110 collects post data using the SNS_A service via the API of SNS server 18A (step S23). Similarly, content-related post collection unit 110 collects post data using the SNS_N service via the API of SNS server 18N (step S23). Note that content-related post collection unit 110 finds time periods when content is being simultaneously distributed to the real space and the metaverse co-viewing space and collects data in real time.
[0040] As shown in FIG. 6, in the cross-platform co-viewing device 100, the content-related post collection unit 110 selects a content-related post from the collected posting data (step S31). The selected content-related post is sent to the emotion estimation unit 132. The emotion estimation unit 132 then estimates the emotion of the user who posted the content-related post based on the content of the content-related post (step S32). The estimation result is sent to the motion generation unit 134. The motion generation unit 134 then generates motion data for an avatar corresponding to the estimated emotion (step S33). The motion generation unit 134 then sends the motion data to the corresponding co-viewing avatars 230 and 240 in the metaverse co-viewing space 200 (step S34). As a result, the corresponding co-viewing avatars 230 and 240 in the metaverse co-viewing space 200 perform an action corresponding to the content of the posting data (step S35).
[0041] As shown in FIG. 3(b), the co-viewing avatar 230 assumes a pose with its arms raised, which is a gesture selected in response to happiness (joy), which is the emotion inferred in response to the post content "Go!" from SNS user 2A. Furthermore, co-viewing avatar 240 assumes a pose of holding his head with both hands, which is a gesture selected in response to sadness (sorrow), which is an emotion estimated in response to the content of the post "I'm hurt" by SNS user 2N. This allows the metaverse user 2X to enjoy the program while feeling as if he or she is watching the program together with the SNS users 2A and 2N outside the metaverse who posted the program. At this time, the content posted by SNS user 2A, "Go!", may be displayed near the top of co-viewing avatar 230. Similarly, the content posted by SNS user 2N, "I got you!", may be displayed near the top of co-viewing avatar 240.
[0042] (Second embodiment) Next, a cross-platform shared viewing device according to the second embodiment will be described with reference to Fig. 7. Note that the same components as those in the first embodiment are denoted by the same reference numerals, and the description thereof will be omitted. 7, the cross-platform shared viewing device 100B according to the second embodiment includes a content-related post collection unit 110 and an expression conversion unit 120B. The expression conversion unit 120B includes a text-to-motion conversion unit 130 and a speech generation unit 140.
[0043] The speech generation unit 140 generates speech data by converting the text data of the content-related post into speech, and sends the speech data to the co-viewing avatars 230 and 240 . In the example shown in FIG. 3(b), co-viewing avatar 230 poses with its arms raised while uttering the words "Go!". Co-viewing avatar 240 poses with its head held in its hands while uttering the words "I got you." At this time, the content posted by SNS user 2A, "Go!", or the content posted by SNS user 2N, "I got you," may be displayed.
[0044] In the metaverse, in addition to the behavior of avatars, voice chat using a microphone by users can also be used as a communication tool. By having the co-viewing avatars 230, 240 speak as in this embodiment, the metaverse user 2X can feel as if the poster is speaking directly through the co-viewing avatars 230, 240. This makes it easier for the metaverse user 2X to get the feeling that they are co-viewing content with SNS users 2A, 2N outside the metaverse, allowing them to enjoy the program even more.
[0045] (Third embodiment) Next, a cross-platform co-viewing device according to the third embodiment will be described with reference to Fig. 8. Note that the same components as those in the second embodiment are denoted by the same reference numerals, and the description thereof will be omitted. As shown in FIG. 8, a cross-platform co-viewing device 100C according to the third embodiment includes a content-related post collection unit 110 and an expression conversion unit 120C. The expression conversion unit 120C includes a text-motion conversion unit 130, a speech sound generation unit 140, a video data generation unit 150, and a sound data generation unit 160.
[0046] The video data generation unit 150 generates video data to be displayed in the metaverse shared viewing space 200 based on the content-related posts extracted by the content-related post collection unit 110 . The sound data generator 160 generates sound data to be output to the metaverse shared viewing space 200 based on the content-related posts extracted by the content-related post collector 110 .
[0047] In this embodiment, the content-related post collection unit 110 is capable of acquiring content-related posts and poster IDs from multiple distributed SNSs. The content-related post collection unit 110 appropriately manages the information of the post data so that the poster's personal information is not leaked to the outside. The conversion of images (visual expression) and sounds (auditory expression) performed by the expression conversion unit 120C can be performed in a variety of ways. Representative patterns are listed below. For example, the operator of the cross-platform shared viewing device 100 can select and set a pattern as appropriate in advance.
[0048] [Pattern 1-1: Expressing the number of posts related to the content] In pattern 1-1, the expression conversion unit 120C produces a metaverse space according to the number of posts related to the content. The video data generation unit 150 generates video data to be displayed in the metaverse shared viewing space 200 based on the number of content-related posts extracted by the content-related post collection unit 110. Examples of types of video include balloons, fireworks, speech bubbles, etc. The video data generation unit 150 may draw balloons, fireworks, and speech bubbles for each post. Taking into account the total number of content-related posts extracted within a predetermined time while the content is being provided, the larger the total number, the larger the balloons and fireworks may be.
[0049] 9(b) is a diagram showing the metaverse co-viewing space 200 when the post shown in FIG. 9(a) is made. In this example, fireworks 250 are launched above the head of the co-viewing avatar 230, and the co-viewing avatar 230 is posing with its arms raised. Note that balloons 260 may be displayed instead of or together with the fireworks 250. The fireworks 250 or balloons 260 may be displayed near the content playback device 210, or may be superimposed on the content playback device 210. This allows the metaverse user 2X to enjoy the program with the feeling that they are watching together with the SNS user 2A outside the metaverse who posted the program. At this time, the post content of the SNS user 2A, "Go!", may be displayed.
[0050] The sound data generator 160 generates sound data to be output to the metaverse shared viewing space 200 based on the number of content-related posts extracted by the content-related post collector 110. Examples of types of sound include cheers, applause, background music, and various sound effects. For cheers, sound data of cheers may be generated by converting the posted data (text) into audio. The sound data generator 160 may play audio for each post. Furthermore, taking into consideration the total number of content-related posts extracted within a predetermined time while the content is being provided, the greater the total number, the louder the volume may be.
[0051] According to pattern 1-1, the cross-platform co-viewing device 100C creates effects such as balloons, fireworks, speech bubbles, and cheers in the metaverse co-viewing space 200 depending on the number of posts, allowing the metaverse user 2X to feel the social presence of others (SNS users 2A to 2N) who are viewing the same content, thereby improving the experience through co-viewing.
[0052] [Pattern 1-2: Expressing the number of contributors related to the content] In pattern 1-2, the representation conversion unit 120C produces a metaverse space according to the ID of the poster who posted a content-related message. The video data generation unit 150 generates video data to be displayed in the metaverse shared viewing space 200 based on the IDs of the posters (contributors) of the content-related posts extracted by the content-related post collection unit 110. The video data generation unit 150 may, for example, draw balloons with different colors for each poster's ID.
[0053] The sound data generator 160 generates sound data to be output to the metaverse shared viewing space 200 based on the ID of the poster (contributor) of the content-related post extracted by the content-related post collector 110. The sound data generator 160 may change the tone of the cheers, for example, for each poster ID. According to pattern 1-2, as with pattern 1-1, metaverse user 2X can sense the social presence of others (SNS users 2A to 2N) who are viewing the same content, and the experience can be improved by co-viewing.
[0054] [Pattern 2: Expressions based on the poster's emotions estimated from content-related posts] In pattern 2, the expression conversion unit 120C produces the metaverse space according to the poster's emotion estimated by the emotion estimation unit 132. Pattern 2 will be explained in more detail below.
[0055] [Pattern 2-1: Expressions corresponding to the most common emotion among all posters] In pattern 2-1, video data generation unit 150 generates video data to be displayed in metaverse shared-viewing space 200 based on, for example, the most prevalent emotion among all the emotions of all contributors estimated by emotion estimation unit 132. At this time, the video data to be displayed in metaverse shared-viewing space 200 may be an object showing the weather as a background image of a virtual object such as content playback device 210. In this case, video data generation unit 150 may draw an image of rain if the most prevalent emotion is negative.
[0056] In the example shown in FIG. 9(b), rain 270 is falling above the head of co-viewing avatar 240, and co-viewing avatar 240 is holding his / her head with both hands. This allows metaverse user 2X to enjoy the program while feeling like he / she is co-viewing with SNS user 2N, who is the poster and is outside the metaverse. At this time, the content posted by SNS user 2N, "I got you," may be displayed.
[0057] The sound data generation unit 160 generates sound data to be output to the metaverse shared viewing space 200 based on, for example, the most prevalent emotion among all the emotions of all contributors estimated by the emotion estimation unit 132. If the most prevalent emotion is negative, the sound data generation unit 160 may play back the sound of booing.
[0058] [Pattern 2-2: Expressions according to the distribution of emotions of all contributors] In pattern 2-2, video data generation unit 150 may generate video data to be displayed in metaverse shared-viewing space 200, for example, based on the distribution of emotions of all contributors estimated by emotion estimation unit 132. In this case, the video data to be displayed in metaverse shared-viewing space 200 may be stamps that are superimposed on objects or background images in the virtual space. In this case, video data generation unit 150 may draw stamps corresponding to positive emotions and stamps corresponding to negative emotions.
[0059] The sound data generation unit 160 generates sound data to be output to the metaverse shared viewing space 200, for example, based on the distribution of emotions of all contributors estimated by the emotion estimation unit 132. The sound data generation unit 160 may also play cheers corresponding to positive emotions and angry shouts corresponding to negative emotions.
[0060] [Pattern 2-3: Expressions that correspond to the individual poster's emotions] In pattern 2-3, video data generation unit 150 may generate video data to be displayed in metaverse shared-viewing space 200 based on the emotion of each individual poster estimated by emotion estimation unit 132. At this time, the video data to be displayed in metaverse shared-viewing space 200 may be an avatar for each speaker's ID. In this case, video data generation unit 150 may draw the avatar for each speaker's ID so as to perform a motion corresponding to the emotion of the speaker (contributor).
[0061] The sound data generation unit 160 generates sound data to be output to the metaverse shared viewing space 200 based on, for example, the emotion of each individual poster estimated by the emotion estimation unit 132. The sound data generation unit 160 may also generate sound data for an avatar for each speaker's ID to speak in accordance with the emotion of the speaker (contributor).
[0062] According to the patterns 2-1, 2-2, and 2-3, similar to the patterns 1-1 and 1-2, the metaverse user 2X can sense the social presence of others (SNS users 2A to 2N) who are viewing the same content, and the shared viewing experience can be improved. In any of the patterns 1-1, 1-2, 2-1, 2-2, and 2-3, the metaverse can be produced according to the popularity of microblogging. Furthermore, it is possible to produce a production that will liven up the metaverse space 200 based on data from the microblogging service.
[0063] 8 includes a video data generation unit 150 and a sound data generation unit 160, it may also be configured to include only one of the video data generation unit 150 and the sound data generation unit 160. In this case, in the expression conversions listed in the above-mentioned patterns 1-1, 1-2, 2-1, 2-2, and 2-3, only the expression conversion to video (visual expression) may be performed, or only the expression conversion to sound (auditory expression) may be performed.
[0064] The expression conversion unit 120C may be modified to have a configuration that excludes the text-to-motion conversion unit 130 and the speech sound generation unit 140 and includes at least one of the video data generation unit 150 and the sound data generation unit 160. Furthermore, the expression conversion unit 120C may be modified to have a configuration that includes at least one of the video data generation unit 150 and the sound data generation unit 160 and the emotion estimation unit 132.
[0065] [Example] In order to confirm the effect of the cross-platform co-viewing device, a co-viewing service providing system including the cross-platform co-viewing device 100 according to the first embodiment was implemented in the following manner. The content server 12, the content playback devices 14A and 14N, and the content playback device 210 were implemented based on MPEG-DASH (see Reference 4 below). Reference 4: “ISO / IEC 230091:2022”, [online], [Retrieved June 7, 2024], Internet<URL:https: / / www.iso.org / standard / 83314.html> The SNS applications (applications 16A to 16N) and the SNS servers 18A to 18N were implemented using Mastodon (registered trademark: see Reference 3 mentioned above).
[0066] The content-related post collection unit 110 of the cross-platform co-viewing device 100 acquires posted data from the SNS servers 18A to 18N using the Mastodon streaming API. The content-related post collection unit 110 acquires the poster's Mastodon account ID along with the posted data. After collecting the posted data, the content-related post collection unit 110 checks whether a hashtag has been added to each piece of posted data, and subjects only posted data that has a hashtag associated with the content being played in the metaverse co-viewing space 200 to subsequent processing.
[0067] The text-to-motion conversion unit 130 associates the SNS account with the co-viewing avatar in the metaverse co-viewing space 200, and determines the motion of the co-viewing avatar according to the content of the posted data of the SNS account. In the text-to-motion conversion unit 130, the emotion estimation unit 132 first estimates the emotion of the poster based on the posted data. Then, the motion generation unit 134 selects a motion corresponding to the estimated emotion from the motion data prepared in advance, and causes the co-viewing avatar to execute it.
[0068] The emotion estimation unit 132 was implemented using BERT (see Reference 5). Reference 5: Devlin, J., Chang, M.-W., Lee, K. and Toutanova, K., “BERT: Pre training of Deep Bidirectional Transformers for Language Understanding,” Proceedings of NAACL-HLT 2019, P. 4171-4186, 2019 The emotion estimation unit 132 classified the emotion implied in the text into one of the 10 emotions mentioned above.
[0069] The experimental method was to conduct verification under the limited condition that only one co-viewing avatar was controlled based on the posting data obtained from one distributed SNS server. We also measured the time it took for the avatar in the metaverse space to start moving after posting a comment. To do this, we used a video camera to record the smartphone posting the comment and the PC running the avatar. We measured the time it took for the avatar to start moving after the user pressed the post button five times.
[0070] Experimental results confirmed that when a comment was posted from a decentralized SNS application, the co-viewing avatar in the metaverse performed a gesture corresponding to the posted content. Furthermore, the maximum time from posting a comment until the avatar in the metaverse space began to move was 0.3 seconds. This confirmed the conceptual feasibility of the system.
[0071] Although the cross-platform co-viewing device according to each embodiment of the present invention has been described above, the scope of the present invention is not limited to these descriptions and should be broadly interpreted based on the claims. Furthermore, it goes without saying that various changes and modifications based on these descriptions are also included in the scope of the present invention.
[0072] For example, in the above-described embodiment, the cross-platform shared viewing device is described as independent hardware, but the present invention is not limited to this. For example, the present invention can also be realized by a program that causes hardware resources such as a CPU, memory, and hard disk of a computer to operate cooperatively as the above-described cross-platform shared viewing device. This program may be distributed via a communication line or written to a recording medium such as a CD-ROM or flash memory and distributed.
[0073] Furthermore, the relationship between SNS commenters and co-viewing avatars is not limited to one-to-one. Multiple (e.g., 100) co-viewing avatars may be generated based on the posting data of one commenter. Alternatively, based on the posting data of multiple (e.g., 100) commenters, one co-viewing avatar may be made to behave, display, speak, etc. individually in accordance with multiple (e.g., 100) comments.
[0074] Furthermore, although the microblogging service is assumed to be a decentralized SNS, even in the case of a general SNS, if the user agrees to make their posts public outside the SNS to which they belong, the content-related post collection unit 110 can collect posts from users of the general SNS.
[0075] Furthermore, the expression conversion patterns into video (visual expression) and sound (auditory expression) performed by expression conversion unit 120C are set by the business operator, but may be selectable by the service user. Furthermore, the roles of SNS users and metaverse users are not fixed, and an SNS user may become a metaverse user, or vice versa. [Explanation of symbols]
[0076] 1. Shared viewing service provision system 2A~2N SNS users 2X Metaverse Users 3 Network 12 Content Server 14A~14N Content playback device 15A~15N User terminal 16A~16N Applications 18A~18N SNS Server 21 Client terminals 25 Virtual space providing device 30A~30N Real Space 100, 100B, 100C Cross-platform shared viewing device 110 Content-related Post Collection Department 120, 120B, 120C Expression conversion section 130 Text-Motion Conversion Unit 132 Emotion estimation part 134 Motion Generation Unit 140 Speech generation unit 150 Video data generation unit 160 Sound data generation unit 200 Metaverse Shared Viewing Space 210 Content playback device 220 Self Avatar 230,240 co-viewing avatars 250 Fireworks 260 Balloons 270 Rain
Claims
1. A cross-platform co-viewing device used in a service in which users of different platforms co-view the same content, a content-related post collection unit that collects text posting data from an SNS server that acquires, accumulates, and publishes the text posting data from SNS users who use a microblog service, and extracts content-related posts that are posting data related to the content of the same content that is provided to content playback devices that are located in a metaverse shared viewing space, which is a virtual space, and in the real space; A cross-platform co-viewing device characterized by comprising: an expression conversion unit that brings the extracted content-related posts onto a metaverse service and, based on these content-related posts, converts the expression of the metaverse co-viewing space presented to metaverse users who watch the content in the metaverse co-viewing space in a predetermined manner.
2. The microblog service is a distributed SNS, The cross-platform co-viewing device according to claim 1 , wherein the content-related post collection unit collects the post data from a distributed SNS server through an API.
3. The representation conversion unit a text-motion conversion unit that generates conversion data for converting the text data of the content-related post into an avatar's motion; The text-to-motion conversion unit an emotion estimation unit that estimates the emotion of the SNS user who posted the post based on the content of the post data; The cross-platform co-viewing device of claim 2, further comprising a motion generation unit that generates a co-viewing avatar in the metaverse co-viewing space that corresponds to the SNS user who posted, generates motion data for the avatar that corresponds to the estimated emotion, and sends the motion data to the co-viewing avatar.
4. the emotion estimation unit classifies the content of the posted data into a plurality of predetermined emotion types using a learning model that has been trained to estimate the emotion contained in the input text; The cross-platform co-viewing device according to claim 3, wherein the motion generation unit selects motion data corresponding to the emotion estimated by the emotion estimation unit from a plurality of motion data previously created corresponding to the plurality of types of emotions.
5. The representation conversion unit The cross-platform co-viewing device of claim 4, further comprising a speech generation unit that generates audio data by converting text data of the content-related post into audio and sends the audio data to the co-viewing avatar.
6. The representation conversion unit A cross-platform co-viewing device as described in any one of claims 2 to 5, characterized in that it is further characterized by comprising a video data generation unit that generates video data to be displayed in the metaverse co-viewing space based on the content-related posts extracted by the content-related post collection unit.
7. The representation conversion unit A cross-platform co-viewing device as described in any one of claims 2 to 5, characterized in that it is equipped with a sound data generation unit that generates sound data to be output to the metaverse co-viewing space based on the content-related posts extracted by the content-related post collection unit.