Content management system, method, device, and program
The content management system addresses the inefficiency of thumbnail images in locating playback positions by using voice-generated comments at specific time points, enhancing playback adjustment efficiency and reducing network resource consumption.
Patent Information
- Application Number
- JP2025201752
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-04
- Estimated Expiration
- 2045-11-21
AI Technical Summary
Thumbnail images are not effective for locating playback positions in content with little or no visual change, increasing the effort required to adjust playback positions and network resource usage.
A content management system that includes a voice generation unit to associate voice with comments at specific time points in content, allowing audio playback at reference points indicated by an interface, reducing the need to reposition content playback and minimizing network resource consumption.
Reduces the effort required to adjust playback positions and suppresses the increase in network resources used by enabling audio-based navigation through content with minimal visual changes.
Smart Images

Figure 0007811303000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a content management system, method, apparatus, and program. [Background technology]
[0002] 2. Description of the Related Art A technique is known for displaying thumbnail images of video content as an aid when searching for a playback position of video content using a seek bar.
[0003] For example, Patent Document 1 discloses a technology in which a scale is displayed overlaid on a video image, a cursor indicating a desired playback position is displayed on the scale, and an image index closest to the desired playback position calculated from the cursor position is displayed. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 6-153130 Summary of the Invention [Problem to be solved by the invention]
[0005] In content with little or no change in visual information, there was a problem in that thumbnail images were not useful for locating the playback position when using the seek bar to locate the playback position in the content.
[0006] In one aspect, the present disclosure has been made in consideration of the above circumstances, and the purpose of the present disclosure is to reduce the effort required to adjust the playback position of content and to suppress an increase in network resource usage. [Means for solving the problem]
[0007] A content management system according to one aspect of the present disclosure includes: a content distribution unit that distributes image content that is a moving image; a voice generating unit that associates a voice read out of text based on comments posted corresponding to time points included between two time points in the image content with the two time points; a content display control unit that displays image content and an interface that specifies a reference time point of the image content; and an audio playback unit that plays audio associated with two points in time that include the reference point in time indicated by the interface.
[0008] A content management system according to one aspect of the present disclosure includes: a content distribution unit that distributes audio content associated with image content; a voice generating unit that associates a voice read out of text based on comments posted corresponding to time points included between two time points in the image content with the two time points; a content display control unit that displays image content and an interface that specifies a reference time point of the image content; and an audio playback unit that plays audio associated with two points in time that include the reference point in time indicated by the interface.
[0009] A content management system according to one aspect of the present disclosure includes: a content distribution unit that distributes image content including a plurality of unit contents whose display order is determined; a voice generating unit that associates a voice reading out a text based on a comment posted in correspondence with two or more consecutive unit contents in the image content with the two or more consecutive unit contents; a content display control unit that displays image content and an interface that designates reference unit content of the image content; and an audio playback unit that plays back audio associated with two or more consecutive unit contents that include the reference time point indicated by the interface therebetween. [Effects of the Invention]
[0010] According to one aspect, the present disclosure can reduce the effort required to adjust the playback position of content, thereby suppressing an increase in the amount of network resources used. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram illustrating the system configuration of a content management system according to the present disclosure. [Figure 2] FIG. 2 is a diagram illustrating a hardware configuration related to the content management system of the present disclosure. [Figure 3] FIG. 3 is a diagram illustrating a functional configuration related to the content management system of the present disclosure. [Figure 4] FIG. 4 is a diagram illustrating an example of a content display screen of the content management system of the present disclosure. [Figure 5] FIG. 5 is a diagram illustrating another example of a content display screen of the content management system of the present disclosure. [Figure 6] FIG. 6 is a diagram illustrating another example of a content display screen of the content management system of the present disclosure. [Figure 7] FIG. 7 is a sequence diagram illustrating the processing flow of the content management system of the present disclosure. [Figure 8] FIG. 8 is a diagram illustrating a functional configuration related to a content management system according to a first modified example. [Figure 9] FIG. 9 is a diagram illustrating an example of a content display screen of a content management system according to the first modified example. [Figure 10] FIG. 10 is a diagram illustrating a functional configuration related to a content management system according to the second modification. [Figure 11] FIG. 11 is a sequence diagram illustrating the processing flow of the content management system according to the second modified example. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In the drawings, identical parts will be designated by the same reference numerals and their description will be omitted. The configurations of the embodiments described below and the actions and effects brought about by such configurations are merely examples and are not limited to the following description. In addition, ordinal numbers such as "first" and "second" will be used as necessary below, but these ordinal numbers are used for convenience of identification and do not indicate any particular meaning such as a particular priority unless otherwise specified. When comparing the magnitude relationship of two numerical values, either of the two criteria of "greater than or equal to" and "greater than" may be used, unless otherwise specified, or either of the two criteria of "less than or equal to" and "less than" may be used.
[0013] [System Overview] The content management system according to the present disclosure is a computer system that manages content played on a user terminal. The content refers to information provided by a computer or computer system that can be recognized by humans, such as audio content, video content expressed by video that may include audio, and e-books that include at least one of text and images. The video content includes video generated in advance by an imaging device or the like, and computer graphics (CG) generated in real time.
[0014] In one example, the content management system provides content to users by distributing content data to user terminals. Distributing content refers to a process performed to provide content to users, such as transmitting information to user terminals via a communication network N.
[0015] The user terminal reproduces the content based on the received data. In one example, the content is provided by a distributor. A distributor is a person who intends to provide information to viewers through a content management system and is a content provider. A viewer is a person who intends to obtain information through a content management system and is a content user.
[0016] The content management system transmits content data provided by a distributor terminal to a viewer terminal. The viewer terminal processes the content data and displays the content on a screen. In this disclosure, distributors and viewers are sometimes collectively referred to as users, and distributor terminals and viewer terminals are sometimes collectively referred to as user terminals.
[0017] In one example, the content management system may distribute video content in real time. In this case, for example, a distributor terminal processes captured video to generate content data and transmits the content data to the content management system in real time. "Real time" means that data is transmitted and received over the same connection after communication between the server (content management system) and the client terminal (user terminal) is established, and there may be delays in the transmission of the content data due to data communication, data processing, etc.
[0018] The content management system receives the content data and transmits the content data in real time to the viewer terminal, thereby delivering live streaming content.
[0019] The content data may be generated in a content management system. That is, the content management system may encode real-time video provided from a distributor terminal to generate content data, and transmit the content data to a viewer terminal in real time.
[0020] The content management system may also be used to archive content so that it can be viewed for a given period of time after real-time distribution.
[0021] In one example, the content management system may perform on-demand distribution, allowing viewers to view content at any time. In this case, the content management system stores content data generated by processing previously captured video, content data generated by processing previously recorded audio, or content data containing at least one of previously created text and images in at least one of storage devices such as an origin server and a cache server.
[0022] The content management system transmits the stored content data to the viewer terminal in response to a request from the viewer.
[0023] The content management system displays content generated based on the content data on the viewer terminal. If there are multiple viewer terminals, the content may be displayed on at least one of the multiple viewer terminals.
[0024] [System Configuration] FIG. 1 is an overall configuration diagram of a content management system 1 according to an embodiment of the present disclosure. As shown in FIG. 1, the content management system 1 includes a server 10, a distributor terminal 20, and a viewer terminal 30. The server 10, the distributor terminal 20, and the viewer terminal 30 are connected to a network N such as the Internet and are capable of communicating with each other. The content management system 1 of this embodiment will be described assuming a known client-server type content management system, but is not limited to this. The number of distributor terminals 20, servers 10, and viewer terminals 30 illustrated in FIG. 1 is not limited to the number shown. The communication network N may include the Internet or an intranet.
[0025] The server 10 may be configured by one or more computers. When multiple computers are used, these computers may be connected to each other via a communication network N to logically configure a single server 10.
[0026] The distributor terminal 20 is a computer used by a distributor. In one example, the distributor terminal 20 has a function of accessing the content management system 1 and transmitting content data indicating content. The distributor terminal 20 may be a smartphone, a tablet terminal, a head-mounted display (HMD), smart glasses, a smart watch, a mobile terminal such as a laptop personal computer or a mobile phone, a stationary terminal such as a desktop personal computer, or an imaging system having a function of capturing, recording, and transmitting video.
[0027] The viewer terminal 30 is a computer used by a viewer. In one example, the viewer terminal 30 has a function of accessing the content management system 1 to receive and display content data indicating the content. The viewer terminal 30 may be a smartphone, a tablet terminal, a head-mounted display (HMD), smart glasses, a smart watch, a mobile terminal such as a laptop personal computer or a mobile phone, or a stationary terminal such as a desktop personal computer.
[0028] Distributors operate distributor terminal 20 to log in to content management system 1, thereby providing content to viewers. Viewers operate viewer terminal 30 to log in to content management system 1, thereby allowing viewers to view content. This disclosure assumes that viewers and distributors of content management system 1 are already logged in. Viewers may be omitted from logging in to content management system 1. In other words, content management system 1 may transmit content data to viewer terminal 30 of a general viewer who is not logged in. In this case, general viewers who are not logged in can also view the content.
[0029] FIG. 2 is a block diagram showing a hardware configuration related to a content management system 1 according to an embodiment of the present disclosure. As an example, a server computer 100 includes a processor 101, a main memory 102, an auxiliary memory 103, and a communication unit 104. The processor 101 is a computing device that executes an operating system and application programs, and is, for example, a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or a field programmable gate array (FPGA). The main memory 102 stores programs to be executed, computation results, etc., and is generally capable of reading and writing data faster than the auxiliary memory 103 (described later), and is configured, for example, by a volatile storage medium such as a dynamic random access memory (DRAM). The auxiliary memory 103 is generally capable of storing larger amounts of data than the main memory 102, and is configured, for example, by a nonvolatile storage medium such as a hard disk or flash memory. The auxiliary storage unit 103 stores a server program P1 and various data for causing the server computer 100 to function as the server 10. The communication unit 104 is a device that executes data communication with other computers via the communication network N.
[0030] In this embodiment, the content management program is implemented as a server program P1. Each functional element of the server 10 is realized by loading the server program P1 onto the processor 101 or the main memory unit 102 and having the processor 101 execute the program. The server program P1 includes code for realizing each functional element of the server 10. The processor 101 operates the communication unit 104 in accordance with the server program P1, and executes reading and writing of data from and to the main memory unit 102 or the auxiliary memory unit 103.
[0031] As an example, the terminal computer 200 includes, as hardware components, a processor 201, a main memory unit 202, an auxiliary memory unit 203, a communication unit 204, an input interface 205, and an output interface 206.
[0032] The processor 201 is a computing device that executes an operating system and application programs, and is, for example, a CPU, a GPU, an ASIC, or an FPGA.
[0033] The main memory unit 202 is a device that stores programs to be executed, calculation results, etc., and is generally a device that can read and write data faster than the auxiliary memory unit 203 described below, and is composed of a volatile storage medium such as a DRAM (Dynamic Random Access Memory).
[0034] The auxiliary storage unit 203 is generally a device capable of storing larger amounts of data than the main storage unit 202, and is configured by a non-volatile storage medium such as a hard disk or flash memory. The auxiliary storage unit 203 stores the client program P2 for causing the terminal computer 200 to function as the distributor terminal 20 or the viewer terminal 30, and various data.
[0035] The communication unit 204 is a device that executes data communication with other computers via the communication network N.
[0036] The input interface 205 is a device that accepts data based on user operations or actions, and is configured with at least one of a keyboard, operation buttons, a pointing device, a touch panel, a microphone, a sensor, and a camera, for example.
[0037] The output interface 206 is a device that outputs data processed by the terminal computer 200, and is configured to include, for example, a display device such as a display.
[0038] Each functional element of the distributor terminal 20 or the viewer terminal 30 is realized by loading the corresponding client program P2 into the processor 201 or the main memory unit 202 and having the processor 201 execute the program. The client program P2 includes code for realizing each functional element of the distributor terminal 20 or the viewer terminal 30. The processor 201 operates the communication unit 204, the input interface 205, the output interface 206, or the imaging unit 207 in accordance with the client program P2, and reads and writes data from and to the main memory unit 202 or the auxiliary memory unit 203.
[0039] At least one of the server program P1 and the client program P2 may be provided by being non-temporarily recorded on a tangible recording medium such as a magnetic tape, a magneto-optical disk, an optical disk, a magnetic disk, or a semiconductor memory. Alternatively, at least one of these programs may be provided as a data signal superimposed on a carrier wave via the communications network N. These programs may be provided separately or together.
[0040] 3 is a diagram showing an example of a functional configuration related to the content management system 1. The server 10 includes a voice generation unit 11, a content distribution unit 12, and a comment reception unit 13 as functional elements.
[0041] The audio generation unit 11 is a functional element that generates a reading audio associated with a reference position of the content. The reference position of the content is information for searching a specific position of the content, such as the playback time of the video content and the audio content, and the page number indicating the display position of the e-book content.
[0042] If the content is video content, the reference positions are any two points in the video. For example, the content positions are specified as "2 minutes 30 seconds" and "2 minutes 45 seconds" based on the start of the video (this can also be said to specify the time from 2 minutes 30 seconds to 2 minutes 45 seconds in the video).
[0043] If the content is audio content, image content (hereinafter referred to as cover art as an example) associated with the audio content is distributed. This cover art may be a video or a still image. The reference positions are any two points in time in the cover art. If the cover art is a still image, the thumbnails associated with either reference position may be reduced versions of the still image.
[0044] If the content is an e-book, it is composed of multiple pages (hereinafter, each of which may be referred to as a unit content) with a set display order. Each page contains at least one of an image and text. Even if each page contains only text, it will be called image content because it will be displayed as an image when displayed. The reference position is any two or more consecutive pages, for example, specified as "pages 5 to 10."
[0045] The content distribution unit 12 is a functional element that distributes the reading audio, the reference position, and the content data to the viewer terminal 30. As described above, if the content is video content, the content data is video data. If the content is audio, the content data is cover art. If the content is an electronic book, the content data is image data of multiple pages (unit content).
[0046] The content distribution unit 12 may distribute data in the same file format as the content data received from the distributor terminal 20 as content data, or may distribute data that has been transcoded from the content data received from the distributor terminal 20 as content data.
[0047] The reading audio may be composed of a single file that combines all of the reading audio to be read at each reference position in the content, or may be composed of multiple files that combine all of the reading audio to be read at each reference position in the content.
[0048] The comment receiving unit 13 is a functional element that receives comments from the viewer terminal 30.
[0049] Although not shown in FIG. 3, the server 10 may be configured to further include a thumbnail generation unit as a functional element. The thumbnail generation unit is a functional element that generates thumbnail images associated with reference positions of content. In one example, the thumbnail images are distributed as sprite images. When the content is video content, the sprite images are composed of image files in which multiple frames included in the video are arranged in a tiled pattern. The images that make up the sprite images may be reduced images of the original content. The sprite images may be a single image file or multiple image files. The thumbnail generation unit may generate a thumbnail image for each reference position associated with the reading audio. In this case, there is a 1:1 correspondence between the reading audio and the thumbnail.
[0050] If the server 10 includes a thumbnail generating unit, the content distribution unit 12 may distribute thumbnail images.
[0051] The distributor terminal 20 includes a functional element, a content generation unit 21. The content generation unit 21 is a functional element that generates content data and transmits it to the server 10.
[0052] The viewer terminal 30 includes functional elements including a content display unit 31 and a comment posting unit 32. The content display unit 31 is a functional element that receives content, a read-aloud voice, and a reference position associated with the read-aloud voice from the server 10, displays the content and a seek bar that can specify a playback position of the content via the output interface 206, and plays the read-aloud voice according to the reference position indicated by the cursor position of the seek bar. The comment posting unit 32 is a functional element that posts a comment on the displayed content in association with a reference position. If the content is a video, the reference position may be the time when the comment was displayed at the time of sending (i.e., the time in the video when the comment was posted). If the content is audio, the reference position may be the time when the comment was played at the time of sending (i.e., the time in the audio when the comment was posted). If the content is an e-book, the reference position may be the page when the comment was displayed at the time of sending (i.e., the page in the e-book when the comment was posted).
[0053] The seek bar is an example of an interface. The reference point may be specified by other interfaces (such as by directly inputting a specific reference point).
[0054] If the audio is composed of a single file, the content display unit 31 determines the position at which to start playing the audio file based on the reference position associated with the audio. If the audio is composed of multiple files, the content display unit 31 determines which audio file to play based on the reference position associated with the audio.
[0055] The content display unit 31 may display a thumbnail image according to the reference position indicated by the seek bar. When the reference positions associated with a plurality of thumbnail images correspond one-to-one to the reference positions associated with a plurality of read-aloud voices, the content display unit 31 may display a thumbnail image associated with the same reference position when the reference position indicated by the cursor position of the seek bar is the reference position associated with the read-aloud voice.
[0056] [Content display screen] 4 is a diagram showing an example of a content display screen when the content is video content. In this example, a video player, a comment posting field for posting comments on the video content, and an information display field for displaying information about the video content being played are displayed on the display screen 300 of the viewer terminal 30.
[0057] The video player displays content 301, a seek bar 302, and comments 303. In Fig. 4, content 301 shows the point 48 minutes and 15 seconds after the start of distribution of the live video (LIVE).
[0058] The viewer terminal 30 receives comments posted on the displayed content from the server 10 and displays them on the video player.
[0059] In this example, comment 303 is displayed superimposed on content 301, moving from right to left according to the timestamp associated with the comment. The direction of movement of comment 303 is not limited to right to left, and for example, comment 303 may be displayed superimposed on content 301, moving from bottom to top. Furthermore, comment 303 may be displayed as a pop-up superimposed on content 301, and then disappear after a certain period of time has passed.
[0060] In this example, the user moves cursor C of seek bar 302 to a point 48 minutes and 15 seconds after the start of distribution, so that a reading audio associated with two points in time, including the point 48 minutes and 15 seconds, is played.
[0061] The comment posting field displays a text box 304 for entering a comment and a post button 305 for posting the entered text. A viewer or distributor playing video content can post a comment on the video content being played by entering text in text box 304 and pressing post button 305.
[0062] The information display field displays a number of views 306, a number of comments 307, and a comment list 308. The number of views 306 indicates the number of times the video content being played has been played. The number of comments 307 indicates the total number of comments posted for the video content being played. The comment list 308 displays a list of comments posted for the video content being played. The comment list 308 does not need to display all comments at once, as long as all comments can be viewed by scrolling. The comment list 308 may not display all posted comments, but may display only comments from all posted comments, excluding inappropriate comments, etc.
[0063] 5 is a diagram showing an example of a content display screen when the content is audio content. In this example, an audio player, a comment posting field for posting comments on the audio content, and an information display field for displaying information about the audio content being played are displayed on the display screen 300a of the viewer terminal 30a.
[0064] The audio player displays content 301a, a seek bar 302a, and comments 303a. In this example, content 301a is cover art for the audio content. Content 301a is displayed while the audio player is playing the audio for the audio content. In FIG. 5, the audio is being played at 1 minute 8 seconds into a 3 minute 36 second piece of music, and the cover art at that time is displayed as content 301a. Note that the cover art may be a still image that is not dependent on the playback time.
[0065] The viewer terminal 30a receives comments posted on the displayed content from the server 10 and displays them on the audio player.
[0066] The viewer terminal 30 receives comments posted on the displayed content from the server 10 and displays them on the video player.
[0067] In this example, comment 303a is displayed superimposed on content 301a, moving from right to left according to the timestamp associated with the comment. The direction of movement of comment 303a is not limited to right to left, and for example, comment 303a may be displayed superimposed on content 301a, moving from bottom to top. Furthermore, comment 303a may be displayed as a pop-up superimposed on content 301a, and then disappear after a certain time has passed.
[0068] In this example, the user moves cursor C of seek bar 302 to the point 1 minute 8 seconds after the start of distribution, so that the audio reading associated with the two points in time including the 1 minute 8 second point is played.
[0069] The comment posting field displays a text box 304a for entering a comment and a post button 305a for posting the entered text. A viewer or a distributor playing audio content can post a comment on the audio content being played by entering text in the text box 304a and pressing the post button 305a.
[0070] The information display field displays a playback count 306a, a comment count 307a, and a comment list 308a. The playback count 306a indicates the number of times the audio content being played has been played. The comment count 307a indicates the total number of comments posted for the audio content being played. The comment list 308a displays a list of comments posted for the audio content being played, which were received from the server 10. The comment list 3098 does not need to display all comments at once, as long as all comments can be viewed by scrolling. The comment list 308a may not display all posted comments, but may display only comments from all posted comments, excluding inappropriate comments, etc.
[0071] 6 is a diagram showing an example of a content display screen when the content is an electronic book content. In this example, an electronic book viewer, a comment posting field for posting comments on the electronic book content, and an information display field for displaying information about the electronic book content being played are displayed on the display screen 300b of the viewer terminal 30b.
[0072] The electronic book viewer displays content 301b, a seek bar 302b, and comments 303b. In Fig. 6, page 206 of the 328-page electronic book is displayed as content 301b.
[0073] The viewer terminal 30b receives comments posted on the displayed content from the server 10 and displays them on the electronic book viewer.
[0074] The viewer terminal 30 receives comments posted on the displayed content from the server 10 and displays them on the video player.
[0075] In this example, comment 303b is displayed superimposed on content 301b, moving from right to left depending on the display position associated with the comment (for example, the page number). The movement direction of comment 303b is not limited to right to left, and for example, comment 303b may be displayed superimposed on content 301b, moving from bottom to top. Furthermore, comment 303b may be displayed as a pop-up superimposed on content 301b, and then disappear after a certain period of time has passed.
[0076] In this example, the user has moved cursor C of the seek bar 302 to the position of page 206, so that a reading voice associated with two or more consecutive unit contents including page 206 therebetween is played back.
[0077] The comment posting field displays a text box 304b for entering a comment and a post button 305b for posting the entered text. A viewer or a distributor playing back electronic book content can post a comment on the electronic book content being played back by entering text in the text box 304b and pressing the post button 305b.
[0078] The information display field displays a number of views 306b, a number of comments 307b, and a comment list 308b. The number of views 306b indicates the number of times the displayed e-book content has been viewed. The number of comments 307b indicates the total number of comments posted for the displayed e-book content. The comment list 308b displays a list of comments posted for the displayed e-book content, which have been received from the server 10. The comment list 308b does not need to display all comments at once, as long as all comments can be viewed by scrolling. The comment list 308b may not display all posted comments, but may display only comments from all posted comments, excluding inappropriate comments and the like.
[0079] [System Operation] The processing flow of the content management system 1 according to an embodiment of the present disclosure will be described with reference to FIG.
[0080] In step S11, the content generating unit 21 generates content data indicating the content.
[0081] In step S12, the content generation unit 21 transmits content data to the server 10 with which a connection with the distributor terminal 20 has been established. The content generation unit 21 may sequentially generate content data indicating parts of the content and sequentially transmit the generated content data to the server 10. In this case, the content may be provided to the viewer terminal 30 as live distribution content. The content generation unit 21 may also transmit content data including the entire content to the server 10. In this case, the content may be provided to the viewer terminal 30 as on-demand distribution content.
[0082] In step S13, the content distribution unit 12 distributes content data to the viewer terminal 30 in response to a request from the viewer terminal 30 for which a connection with the server 10 has been established.
[0083] In step S14, the content display unit 31 plays back the content based on the received content data.
[0084] In step S15, the comment posting unit 32 posts a comment in association with the reference position of the content being played back.
[0085] In step S16, the voice generation unit 11 generates a reading voice based on the comment received by the comment receiving unit 13. The reading voice is associated with a reference position at which the reading voice is played. In this embodiment, the reading voice is generated by generating a voice that reads out text based on a comment associated with a reference position included in the reference position associated with the reading voice. The reference position associated with the reading voice may be set by the start position and end position at which the reading voice is played, or may be set by the length of the position at which the reading voice starts to be played and the order in which the reading voice is played.
[0086] The comment-based text may be some or all of the comments associated with the reference location included in the reference location associated with the spoken audio.
[0087] The comment-based text may be a portion of comments selected based on a predetermined criterion from all comments associated with reference positions included in the reference positions associated with the speech reading. In this case, at least one of the following criteria may be used: the number of favorable reactions is greater than a predetermined number; the comment was posted more recently than a predetermined date and time; and the comment does not contain inappropriate language. A known method may be used for the selection, and for example, artificial intelligence technology may be used.
[0088] The text based on the comment may be a text summarizing all or some of the comments associated with the reference position included in the reference position associated with the audio reading. The text based on the comment may also be a text summarizing the comment and at least one of the content situation of the reference position associated with the audio reading and the audio played at the reference position associated with the audio reading. Any known method may be applied for summarization or transcription, and artificial intelligence technology, for example, may be used. When a summary is used in the audio reading, the playback time of the audio reading can be shortened compared to when all the comments are read aloud, and the content of the comment can be explained, thereby suppressing an increase in network resource consumption.
[0089] When the content display unit 31 displays thumbnail images and further displays comments superimposed on the content, the text based on the comments may be part or all of the comments superimposed on the content when the scene corresponding to the thumbnail image is displayed. For example, if the thumbnail image is a reduced image of the video content at 2 minutes 30 seconds, part or all of the comments displayed during the display of the scene at 2 minutes 30 seconds in the video content may be the text based on the comments.
[0090] As an example, the voice generation unit 11 can generate a reading voice in the following procedure. First, the voice generation unit 11 determines a predetermined section indicating the start position and end position of a reference position to be associated with the reading voice. Next, the voice generation unit 11 acquires a comment associated with the reference position included in the predetermined section. Finally, the voice generation unit 11 inputs the acquired comment into voice synthesis software to generate a voice that reads the comment aloud.
[0091] When the content is video content, the audio generation unit 11 can generate a reading audio from text based on comments posted corresponding to a point in time between two points in time in the video content, and associate the two points in time with the generated reading audio.
[0092] As a specific example, the voice generation unit 11 generates a reading voice from text based on a comment posted when the video content is displayed at any time between two points in time, "2 minutes 30 seconds" and "2 minutes 45 seconds," in the video content. The voice generation unit 11 also associates the generated reading voice with the two points in time, "2 minutes 30 seconds" and "2 minutes 45 seconds."
[0093] If no comment has been posted at the relevant time, the voice generating unit 11 does not need to generate a voice (the same applies below).
[0094] More specifically, when the comment-based text is a summary of the comment, the speech generation unit 11 can generate the reading speech by, for example, the following procedure. First, the speech generation unit 11 determines a predetermined section indicating the start and end positions of a reference position to be associated with the reading speech. Next, the speech generation unit 11 acquires the comment associated with the reference position included in the predetermined section. Thereafter, the speech generation unit 11 generates a summary of the comment using a machine learning model such as a large-scale language model. Finally, the speech generation unit 11 inputs the acquired summary into speech synthesis software to generate a speech that reads out the summary.
[0095] In step S17, the content distribution unit 12 transmits the content data, the read-out voice, and the reference position to the viewer terminal 30 in response to a request from the viewer terminal 30 that has established a connection with the server 10.
[0096] In step S18, the content display unit 31 plays back the content based on the received content data. The content display unit 31 plays back the content (displaying video in video content, playing audio and displaying cover art in audio content, displaying each page in an e-book, etc.) and displays a seek bar that allows the user to specify the playback position of the content.
[0097] In step S19, when the user of the viewer terminal 30 changes the cursor position of the slider of the seek bar, the content display unit 31 plays back the read-aloud voice associated with the reference position corresponding to the cursor position.
[0098] For example, if the content is video or audio content and the seek bar indicates 2 minutes 40 seconds as the reference time, the content display unit 31 plays back the audio readout associated with two time points that include 2 minutes 40 seconds (in the above example, 2 minutes 30 seconds to 2 minutes 45 seconds).
[0099] Suppose the content image is an electronic book, and the seek bar indicates page 7 as the reference page. In this case, the content display unit 31 plays back a reading voice associated with two or more consecutive pages (pages 5 to 10 in the above example) including page 7.
[0100] By playing back the audio reading, it becomes possible to adjust the playback position of the content by listening to the audio reading, even for content with little or no change in visual information, so there is no need to acquire and play back the content at the corresponding position each time to adjust the playback position, thereby suppressing an increase in the amount of network resources used.
[0101] If the content includes audio, the content display unit 31 may lower the volume of the audio playback of the content or stop the playback of the audio of the content while the reading audio is being played back.
[0102] For example, if the content is video content, the content display unit 31 reduces or mutes the volume of the video while the reading audio is being played. If the content is video content, the content display unit 31 may stop the playback of the video while the reading audio is being played, thereby stopping the playback of the audio.
[0103] For example, if the content is audio content, the content display unit 31 reduces the volume of the audio content or mutes the audio content while the reading audio is being played back. If the content is audio content, the content display unit 31 may stop the playback of the audio by stopping the playback of the audio content while the reading audio is being played back.
[0104] By lowering or muting the volume of the video while the reading voice is being played, it is possible to prevent the reading voice from becoming difficult to hear due to the content audio and the reading voice overlapping.
[0105] The content display unit 31 may start playing the reading audio when the reference position indicated by the cursor of the seek bar remains at a position included in a specified section for a specified period of time or more, or when it indicates the same position for a specified period of time or more.
[0106] For example, suppose the content is video or audio content, and the reading audio is associated with 2 minutes 30 seconds to 2 minutes 45 seconds. In this case, if the reference time indicated by the cursor of the seek bar stays between 2 minutes 30 seconds and 2 minutes 45 seconds for a predetermined time (e.g., 1 second) or more, the content display unit 31 plays the reading audio associated with 2 minutes 30 seconds to 2 minutes 45 seconds. Alternatively, if the reference time indicated by the cursor of the seek bar stays between 2 minutes 30 seconds and 2 minutes 45 seconds (e.g., 2 minutes 40 seconds) for a predetermined time (e.g., 1 second) or more, the content display unit 31 plays the reading audio associated with 2 minutes 30 seconds to 2 minutes 45 seconds.
[0107] Assume that the content image is an electronic book and pages 5 to 10 are associated with the reading audio. In this case, if the reference page indicated by the cursor of the seek bar stays between pages 5 and 10 for a predetermined time (e.g., one second) or more, the content display unit 31 plays the reading audio associated with pages 5 to 10. Alternatively, if the reference page indicated by the cursor of the seek bar stays on a page between pages 5 and 10 (e.g., page 7) for a predetermined time (e.g., one second) or more, the content display unit 31 plays the reading audio associated with pages 5 to 10.
[0108] By controlling the timing at which the playback of the read-out voice starts as described above, it is possible to prevent an increase in consumption of calculation resources when the cursor of the seek bar is moved quickly.
[0109] [Content Management System According to First Modification] The functional configuration and operation of a content management system 1A according to a first modified example will be described with reference to Figures 8 and 9. The content management system 1A includes a server 10A, a distributor terminal 20A, and a viewer terminal 30A. In the following, explanations of points that are the same between the content management system 1 and the content management system 1A may be omitted.
[0110] 8 is a diagram showing an example of a functional configuration related to a content management system 1A. The server 10A includes, as functional elements, a text generation unit 11A, a content distribution unit 12A, and a comment receiving unit 13A. In the content management system 1A according to this modification, the server 10A does not need to include a voice generation unit.
[0111] The text generation unit 11A is a functional element that generates a reading text associated with a reference position of the content. The content distribution unit 12A is a functional element that distributes the reading text, the reference position, and the content data to the viewer terminal 30A.
[0112] The distributor terminal 20A includes a content generation unit 21A as a functional element. The content generation unit 21A is a functional element that generates content data and transmits it to the server 10A.
[0113] The viewer terminal 30A includes, as functional elements, a content display unit 31A, a comment posting unit 32A, and a voice generation unit 33A. The content display unit 31A is a functional element that receives content, a voice-reading text, and a reference position associated with the voice-reading text from the server 10A, displays the content and a seek bar that can specify a playback position of the content via the output interface 206, and plays a voice-reading sound according to the reference position indicated by the cursor position of the seek bar. The voice generation unit 33A is a functional element that generates a voice that reads the voice-reading text using speech synthesis software or the like.
[0114] [System Operation] The processing flow of the content management system 1A according to an embodiment of the present disclosure will be described with reference to FIG.
[0115] In step S11A, the content generating unit 21A generates content data indicating the content.
[0116] In step S12A, the content generation unit 21A transmits the content data to the server 10A with which a connection with the distributor terminal 20A has been established.
[0117] In step S13A, the content distribution unit 12A distributes content data to the viewer terminal 30A in response to a request from the viewer terminal 30A that has established a connection with the server 10A.
[0118] In step S14A, the content display unit 31A plays back the content based on the received content data.
[0119] In step S15A, the comment posting unit 32A posts a comment in association with the reference position of the content being played back.
[0120] In step S16A, the text generation unit 11A generates a reading text based on the comment received by the comment receiving unit 13. The reading text is associated with a reference position at which a reading voice that reads the reading text aloud is played. The reading text is text based on a comment associated with a reference position included in the reference position associated with the reading text.
[0121] The comment-based text may be some or all of the comments associated with the reference location included in the reference location associated with the spoken text.
[0122] The comment-based text may be a portion of comments selected based on a predetermined criterion from all comments associated with reference positions included in the reference positions associated with the spoken text. In this case, at least one of the following criteria may be used: the number of favorable reactions is greater than a predetermined number; the comment was posted more recently than a predetermined date and time; and the comment does not contain inappropriate language. A known method may be used for the selection, and for example, artificial intelligence technology may be used.
[0123] The comment-based text may be a text summarizing all or some of the comments associated with the reference positions included in the reference positions associated with the text-to-speech. Alternatively, the comment-based text may be a text summarizing at least one of the content status of the reference positions associated with the text-to-speech and the audio played at the reference positions associated with the text-to-speech, and the comment. When a summary is used in the text-to-speech, the playback time of the audio can be shortened compared to when all the comments are read aloud, and the content of the comment can be explained, thereby suppressing the increase in consumption of network resources. A known method may be applied for summarization and transcription, and artificial intelligence technology, for example, may be used.
[0124] When the content display unit 31A displays thumbnail images and further displays comments superimposed on the content, the text based on the comments may be part or all of the comments superimposed on the content when the scene corresponding to the thumbnail image is displayed. For example, if the thumbnail image is a reduced image of the video content at 2 minutes 30 seconds, part or all of the comments displayed during the display of the scene at 2 minutes 30 seconds in the video content may be the text based on the comments.
[0125] In step S17A, the content distribution unit 12A transmits content data, text to be read, and a reference position to the viewer terminal 30A in response to a request from the viewer terminal 30A that has established a connection with the server 10. The transmitted comments include not only comments posted by the viewer terminal 30A itself, but also comments posted by other viewer terminals. The comments are also associated with the reference position of the content.
[0126] In step S18A, the content display unit 31A plays the content based on the received content data (displaying video in video content, playing audio and displaying cover art in audio content, displaying each page in an e-book, etc.), and displays a seek bar that allows the playback position of the content to be specified.
[0127] In step S19A, the voice generation unit 33A inputs the received text to speech synthesis software to generate a speech. The voice generation unit 33A associates a reference position corresponding to the text to be spoken with the speech. The speech synthesis software may be software stored in the viewer terminal 30A, or may be a web application that can acquire the speech via a network.
[0128] In step S20A, when the user of the viewer terminal 30A changes the cursor position of the slider of the seek bar, the content display unit 31A plays back the read-aloud voice associated with the reference position corresponding to the cursor position.
[0129] For example, if the content is video or audio content and the seek bar indicates 2 minutes 40 seconds as the reference time, the content display unit 31 plays back the audio readout associated with two time points that include 2 minutes 40 seconds (in the above example, 2 minutes 30 seconds to 2 minutes 45 seconds).
[0130] Suppose the content image is an electronic book, and the seek bar indicates the reference page as page 7. In this case, the content display unit 31 plays back a reading voice associated with two or more consecutive pages (in the above example, pages 5 to 10) including page 7.
[0131] By playing back the audio reading, it becomes possible to adjust the playback position of the content by listening to the audio reading, even for content with little or no change in visual information, so there is no need to acquire and play back the content at the corresponding position each time to adjust the playback position, thereby suppressing an increase in the amount of network resources used.
[0132] If the content includes audio, the content display unit 31A may lower the volume of the audio playback of the content or stop the playback of the audio of the content while the reading audio is being played back.
[0133] For example, if the content is video content, the content display unit 31A reduces the volume of the video or mutes the video while reproducing the reading audio. If the content is video content, the content display unit 31A may stop the playback of the video while reproducing the reading audio, thereby stopping the playback of the audio.
[0134] For example, if the content is audio content, the content display unit 31A may reduce the volume of the audio content or mute the audio content while reproducing the reading audio. If the content is audio content, the content display unit 31A may stop the audio reproduction by stopping the reproduction of the audio content while reproducing the reading audio.
[0135] By lowering or muting the volume of the video while the reading voice is being played, it is possible to prevent the reading voice from becoming difficult to hear due to the content audio and the reading voice overlapping.
[0136] The content display unit 31A may start playing the reading audio when the reference position indicated by the cursor of the seek bar remains at a position included in a specified section for a specified period of time or more, or when it indicates the same position for a specified period of time or more.
[0137] For example, suppose the content is video or audio content, and the reading audio is associated with 2 minutes 30 seconds to 2 minutes 45 seconds. In this case, if the reference time indicated by the cursor of the seek bar stays between 2 minutes 30 seconds and 2 minutes 45 seconds for a predetermined time (e.g., 1 second) or more, the content display unit 31 plays the reading audio associated with 2 minutes 30 seconds to 2 minutes 45 seconds. Alternatively, if the reference time indicated by the cursor of the seek bar stays between 2 minutes 30 seconds and 2 minutes 45 seconds (e.g., 2 minutes 40 seconds) for a predetermined time (e.g., 1 second) or more, the content display unit 31 plays the reading audio associated with 2 minutes 30 seconds to 2 minutes 45 seconds.
[0138] Assume that the content image is an electronic book and pages 5 to 10 are associated with the reading audio. In this case, if the reference page indicated by the cursor of the seek bar stays between pages 5 and 10 for a predetermined time (e.g., one second) or more, the content display unit 31 plays the reading audio associated with pages 5 to 10. Alternatively, if the reference page indicated by the cursor of the seek bar stays on a page between pages 5 and 10 (e.g., page 7) for a predetermined time (e.g., one second) or more, the content display unit 31 plays the reading audio associated with pages 5 to 10.
[0139] By controlling the timing at which the playback of the read-out voice starts as described above, it is possible to prevent an increase in consumption of calculation resources when the cursor of the seek bar is moved quickly.
[0140] [Content Management System According to the Second Modification] The functional configuration and operation of a content management system 1B according to a second modified example will be described with reference to Figures 10 and 11. The content management system 1B includes a server 10B, a distributor terminal 20B, and a viewer terminal 30B. In the following, explanations of points that are the same between the content management system 1 and the content management system 1B may be omitted.
[0141] 10 is a diagram showing an example of a functional configuration related to a content management system 1B. The server 10B includes, as functional elements, a content distribution unit 12B and a comment receiving unit 13B. In the content management system 1B according to this modification, the server 10B does not need to include a voice generation unit.
[0142] The distributor terminal 20B includes a content generation unit 21B as a functional element. The content generation unit 21B is a functional element that generates content data and transmits it to the server 10B.
[0143] The viewer terminal 30B includes functional elements including a content display unit 31B, a comment posting unit 32B, and a voice generation unit 33B. The content display unit 31B is a functional element that receives content from the server 10B, displays the content via the output interface 206 and a seek bar that can specify a playback position in the content, and plays back a voice read aloud according to a reference position indicated by the cursor position of the seek bar. The voice generation unit 33B is a functional element that generates a voice read aloud text associated with a reference position in the content, and generates a voice read out the voice read aloud text using voice synthesis software or the like.
[0144] [System Operation] A processing flow of the content management system 1B according to an embodiment of the present disclosure will be described with reference to FIG.
[0145] In step S11B, the content generation unit 21B generates content data indicating the content.
[0146] In step S12B, the content generation unit 21B transmits the content data to the server 10B with which a connection with the distributor terminal 20B has been established.
[0147] In step S13B, the content distribution unit 12B distributes content data to the viewer terminal 30B in response to a request from the viewer terminal 30B that has established a connection with the server 10B.
[0148] In step S14B, the content display unit 31 plays back the content based on the received content data.
[0149] In step S15B, the comment posting unit 32B posts a comment in association with the reference position of the content being played back.
[0150] In step S16B, the content distribution unit 12B transmits the content data and the comments received by the comment receiving unit 13B to the viewer terminal 30B in response to a request from the viewer terminal 30B that has established a connection with the server 10B.
[0151] In step S17B, the content display unit 31B plays the content based on the received content data (displaying video in video content, playing audio and displaying cover art in audio content, displaying each page in an e-book, etc.), and displays a seek bar that allows the playback position of the content to be specified.
[0152] In step S18B, the voice generating unit 33B generates a reading text based on the comment, and inputs the reading text into voice synthesis software to generate a reading voice. The voice generating unit 33B associates the reading voice with a reference position.
[0153] The comment-based text may be all comments associated with reference locations included in the reference location associated with the spoken audio.
[0154] The comment-based text may be a portion of comments selected based on a predetermined criterion from all comments associated with reference positions included in the reference positions associated with the speech reading. In this case, at least one of the following may be used as the criterion: the number of favorable responses is greater than a predetermined number; the comment was posted more recently than a predetermined date and time; and the comment does not contain inappropriate language.
[0155] The comment-based text may be a text summarizing all or some of the comments associated with the reference positions included in the reference positions associated with the speech-to-speech. Alternatively, the comment-based text may be a text summarizing at least one of the content status of the reference positions associated with the speech-to-speech and the speech played at the reference positions associated with the speech-to-speech, and the comment. When a summary sentence is used in the speech-to-speech, the playback time of the speech-to-speech can be shortened compared to when all the comments are read aloud, and the content of the comment can be explained, thereby suppressing an increase in consumption of network resources.
[0156] When the content display unit displays thumbnail images and further displays comments superimposed on the content, the text based on the comments may be all comments displayed superimposed on the content when the scene corresponding to the thumbnail image is displayed.
[0157] When the content display unit displays a thumbnail image, the sound generation unit 33B may set the reference position associated with the thumbnail image as the reference position to be associated with the reading sound.
[0158] In step S19B, when the user of the viewer terminal 30B changes the cursor position of the slider of the seek bar, the content display unit 31B plays back the read-aloud voice associated with the reference position corresponding to the cursor position.
[0159] For example, if the content is video or audio content and the seek bar indicates 2 minutes 40 seconds as the reference time, the content display unit 31 plays back the audio readout associated with two time points that include 2 minutes 40 seconds (in the above example, 2 minutes 30 seconds to 2 minutes 45 seconds).
[0160] Suppose the content image is an electronic book, and the seek bar indicates page 7 as the reference page. In this case, the content display unit 31 plays back a reading voice associated with two or more consecutive pages (pages 5 to 10 in the above example) including page 7.
[0161] By playing back the audio reading, it becomes possible to adjust the playback position of the content by listening to the audio reading, even for content with little or no change in visual information, so there is no need to acquire and play back the content at the corresponding position each time to adjust the playback position, thereby suppressing an increase in the amount of network resources used.
[0162] If the content includes audio, the content display unit 31B may lower the volume of the audio playback of the content while playing back the read-out audio, or may stop the playback of the audio of the content.
[0163] For example, if the content is video content, the content display unit 31 reduces or mutes the volume of the video while the reading audio is being played. If the content is video content, the content display unit 31 may stop the playback of the video while the reading audio is being played, thereby stopping the playback of the audio.
[0164] For example, if the content is audio content, the content display unit 31 reduces the volume of the audio content or mutes the audio content while the reading audio is being played back. If the content is audio content, the content display unit 31 may stop the playback of the audio by stopping the playback of the audio content while the reading audio is being played back.
[0165] By lowering or muting the volume of the video while the reading voice is being played, it is possible to prevent the reading voice from becoming difficult to hear due to the content audio and the reading voice overlapping.
[0166] The content display unit 31B may start playing the reading audio when the reference position indicated by the cursor of the seek bar remains at a position included in a specified section for a specified period of time or more, or when it indicates the same position for a specified period of time or more.
[0167] For example, suppose the content is video or audio content, and the reading audio is associated with 2 minutes 30 seconds to 2 minutes 45 seconds. In this case, if the reference time indicated by the cursor of the seek bar stays between 2 minutes 30 seconds and 2 minutes 45 seconds for a predetermined time (e.g., 1 second) or more, the content display unit 31 plays the reading audio associated with 2 minutes 30 seconds to 2 minutes 45 seconds. Alternatively, if the reference time indicated by the cursor of the seek bar stays between 2 minutes 30 seconds and 2 minutes 45 seconds (e.g., 2 minutes 40 seconds) for a predetermined time (e.g., 1 second) or more, the content display unit 31 plays the reading audio associated with 2 minutes 30 seconds to 2 minutes 45 seconds.
[0168] Assume that the content image is an electronic book and pages 5 to 10 are associated with the reading audio. In this case, if the reference page indicated by the cursor of the seek bar stays between pages 5 and 10 for a predetermined time (e.g., one second) or more, the content display unit 31 plays the reading audio associated with pages 5 to 10. Alternatively, if the reference page indicated by the cursor of the seek bar stays on a page between pages 5 and 10 (e.g., page 7) for a predetermined time (e.g., one second) or more, the content display unit 31 plays the reading audio associated with pages 5 to 10.
[0169] By controlling the timing at which the playback of the read-out voice starts as described above, it is possible to prevent an increase in consumption of calculation resources when the cursor of the seek bar is moved quickly.
[0170] Any part or all of the functional units described herein may be realized by a program. The programs described herein may be distributed by being non-temporarily recorded on a computer-readable recording medium, distributed via a communication line (including wireless communication) such as the Internet, or distributed in a state where they are installed on any terminal. Alternatively, the programs may be programs that run on a web browser (so-called web apps). In this case, a computer may receive a program written in a markup language file (e.g., an HTML file) from a server and execute the program using the web browser.
[0171] Based on the above description, a person skilled in the art may be able to conceive additional effects and various modifications of the present invention, but the aspects of the present invention are not limited to the individual embodiments described above. Various additions, modifications, and partial deletions are possible within the scope of the conceptual idea and spirit of the present invention, which is derived from the content defined in the claims and their equivalents.
[0172] For example, what is described in this specification as a single device (or component, the same applies hereinafter) (including what is depicted as a single device in the drawings) may be realized by multiple devices. Conversely, what is described in this specification as multiple devices (including what is depicted as multiple devices in the drawings) may be realized by a single device. Alternatively, some or all of the means or functions included in a certain device (e.g., a server) may be included in another device (e.g., a user terminal). Furthermore, a "system" may consist of a single device, or may consist of two or more devices (e.g., a server and a user terminal, or multiple user terminals).
[0173] Furthermore, not all of the features described in this specification are essential requirements. In particular, features described in this specification but not included in the claims can be considered optional additional features.
[0174] It should be noted that the applicant is merely aware of the inventions disclosed in the documents listed in the "Prior Art Documents" section of this specification, and the present invention does not necessarily aim to solve the problems of the disclosed inventions. The problem that the present invention aims to solve should be determined by taking into consideration the entire specification. For example, if this specification states that a specific configuration achieves a certain effect, it can also be said that the present invention solves a problem that is the reverse of that effect. However, it is not necessarily intended that such a specific configuration be an essential requirement.
[0175] Any part or all of the functional units described herein may be realized by a program. The programs described herein may be distributed by being non-temporarily recorded on a computer-readable recording medium, distributed via a communication line (including wireless communication) such as the Internet, or distributed in a state where they are installed on any terminal. Alternatively, the programs may be programs that run on a web browser (so-called web apps). In this case, a computer may receive a program written in a markup language file (e.g., an HTML file) from a server and execute the program using the web browser.
[0176] Based on the above description, a person skilled in the art may be able to conceive additional effects and various modifications of the present invention, but the aspects of the present invention are not limited to the individual embodiments described above. Various additions, modifications, and partial deletions are possible within the scope of the conceptual idea and spirit of the present invention, which is derived from the content defined in the claims and their equivalents.
[0177] For example, what is described in this specification as a single device (or component, the same applies hereinafter) (including what is depicted as a single device in the drawings) may be realized by multiple devices. Conversely, what is described in this specification as multiple devices (including what is depicted as multiple devices in the drawings) may be realized by a single device. Alternatively, some or all of the means or functions included in a certain device (e.g., a server) may be included in another device (e.g., a user terminal). Furthermore, a "system" may consist of a single device, or may consist of two or more devices (e.g., a server and a user terminal, or multiple user terminals).
[0178] Furthermore, not all of the matters described in this specification are essential requirements. In particular, matters described in this specification but not described in the claims can be considered optional additional matters. Furthermore, if a certain generic concept is described in this specification but a specific subordinate concept is not mentioned, that specific subordinate concept may or may not be applied (i.e., this specification discloses the generic concept excluding the specific subordinate concept). Furthermore, if there is no mention of whether or not a certain device (method) includes a specific configuration (step), that specific configuration (step) may or may not be included.
[0179] It should be noted that the applicant is merely aware of the inventions disclosed in the documents listed in the "Prior Art Documents" section of this specification, and the present invention does not necessarily aim to solve the problems of the disclosed inventions. The problem that the present invention aims to solve should be determined by taking into consideration the entire specification. For example, if this specification states that a specific configuration achieves a certain effect, it can also be said that the present invention solves a problem that is the reverse of that effect. However, it is not necessarily intended that such a specific configuration be an essential requirement.
[0180] [Note] The present disclosure discloses the following configurations. [Section 1] a content distribution unit that distributes image content that is a moving image; a voice generating unit that reads out a text based on a comment posted corresponding to a time point included between two time points in the image content, and associates the text with the two time points; a content display control unit that displays the image content and an interface that specifies a reference time point of the image content; and an audio playback unit that plays back the audio associated with two points in time that include the reference point indicated by the interface. [Section 2] a content distribution unit that distributes audio content associated with image content; a voice generating unit that reads out a text based on a comment posted corresponding to a time point included between two time points in the image content, and associates the text with the two time points; a content display control unit that displays the image content and an interface that specifies a reference time point of the image content; and an audio playback unit that plays back the audio associated with two points in time that include the reference point indicated by the interface. [Section 3] a content distribution unit that distributes image content including a plurality of unit contents whose display order is determined; a voice generating unit that reads out a text based on a comment posted in correspondence with two or more consecutive content units in the image content, and associates the text with the two or more consecutive content units; a content display control unit that displays the image content and an interface that designates a reference unit content of the image content; an audio playback unit that plays back the audio associated with the two or more consecutive unit contents that include the reference time point indicated by the interface. [Section 4] 3. The content management system according to any one of items 1 to 2, wherein the audio playback unit reduces the playback volume of the audio of the image content while playing back the audio. [Section 5] 3. The content management system according to any one of items 1 to 2, wherein the audio playback unit stops playback of the audio of the image content while playing back the audio. [Section 6] A content management system described in any one of items 1 to 5, wherein the audio playback unit plays the audio when the reference point indicated by the interface indicates a point that is included in the two points in time for a predetermined period of time or more. [Section 7] 6. The content management system according to any one of claims 1 to 5, wherein the audio playback unit plays the audio when the reference point indicated by the interface indicates the same point for a predetermined period of time or more. [Section 8] 8. The content management system according to any one of claims 1 to 7, wherein the voice generation unit uses as the text a sentence summarizing comments posted corresponding to a time point included between the two time points. [Section 9] The thumbnail generator is generating thumbnail images corresponding to two time points in the image content; Associating the two time points with the generated thumbnail image; The content display control unit 9. The content management system according to any one of claims 1 to 8, wherein thumbnail images associated with two time points between which the reference time point indicated by the interface is interposed are displayed. [Section 10] One or more hardware processors A first step of delivering image content that is a moving image; a second step of associating a voice reading a text based on a comment posted corresponding to a time point included between the two time points in the image content with the two time points; a third step of displaying the image content and an interface for specifying a reference time point of the image content; A content management method that executes a fourth step of playing the audio associated with two points in time that include the reference point indicated by the interface. [Section 11] One or more hardware processors a first step of delivering audio content having associated image content; a second step of associating a voice reading a text based on a comment posted corresponding to a time point included between the two time points in the image content with the two time points; a third step of displaying the image content and an interface for specifying a reference time point of the image content; A content management method that executes a fourth step of playing the audio associated with two points in time that include the reference point indicated by the interface. [Section 12] One or more hardware processors A first step of distributing image content including unit content included in a plurality of unit content whose display order is determined; a second step of associating a voice reading out text based on comments posted corresponding to two or more consecutive unit contents in the image content with the two or more consecutive unit contents; a third step of displaying the image content and an interface for specifying a reference unit content of the image content; a fourth step of playing the audio associated with the two or more consecutive content units that include the reference time point indicated by the interface therebetween. [Section 13] a content receiving unit that receives image content, which is a moving image, and audio; a content display control unit that displays the image content and an interface that specifies a reference time point of the image content; an audio playback unit that plays back the audio associated with two time points that include a reference time point indicated by the interface; A content display device, wherein the audio is a reading of text based on comments posted corresponding to points in time between two points in time in the image content. [Section 14] a content receiving unit that receives audio content associated with image content and audio; a content display control unit that displays the audio content and an interface that specifies a reference time point of the audio content; an audio playback unit that plays back the audio associated with two time points that include a reference time point indicated by the interface; A content display device, wherein the audio is a reading of text based on comments posted corresponding to points in time between two points in time in the audio content. [Section 15] a content receiving unit for receiving image content including a plurality of unit contents whose display order is determined, and audio; a content display control unit that displays the image content and an interface that specifies a reference time point of the image content; an audio playback unit that plays back the audio associated with two or more consecutive content units that include the reference time point indicated by the interface; A content display device, wherein the audio is a reading of text based on comments posted in association with unit content included in two or more consecutive unit content in the image content. [Section 16] One or more hardware processors receiving image content, which is a moving image, and audio, the audio being a reading of text based on comments posted corresponding to time points in the image content that are included between two time points; displaying the image content and an interface for specifying a reference time point of the image content; and playing the audio associated with two points in time that include the reference point in time indicated by the interface. [Section 17] One or more hardware processors receiving audio content associated with image content and audio, the audio being a reading of text based on comments posted corresponding to time points in the audio content that fall between two time points; displaying the audio content and an interface for specifying a reference time point of the audio content; and playing the audio associated with two points in time that include the reference point in time indicated by the interface. [Section 18] One or more hardware processors receiving image content including a plurality of unit contents whose display order is determined, and audio, wherein the audio is a reading of text based on comments posted corresponding to the unit contents included in two or more unit contents in the image content; displaying the image content and an interface for specifying a reference time point of the image content; and playing back the audio associated with two or more unit content pieces that include the reference time point indicated by the interface therebetween. [Explanation of symbols]
[0181] 1...content management system, 10...server, 20...distributor terminal, 30...viewer terminal, 11...audio generation unit, 12...content distribution unit, 13...comment receiving unit, 21...content generation unit, 31...content display unit, 32...comment posting unit
Claims
1. a content distribution unit that distributes image content that is a moving image; a voice generating unit that reads out text based on comments posted corresponding to time points included between two time points in the image content, and associates the text with the two time points; a content display control unit that displays the image content and an interface that specifies a reference time point of the image content; and an audio playback unit that plays back the audio associated with two points in time that include the reference point indicated by the interface.
2. a content distribution unit that distributes audio content associated with image content; a voice generating unit that reads out text based on comments posted corresponding to time points included between two time points in the image content, and associates the text with the two time points; a content display control unit that displays the image content and an interface that specifies a reference time point of the image content; and an audio playback unit that plays back the audio associated with two points in time that include the reference point indicated by the interface.
3. a content distribution unit that distributes image content including a plurality of unit contents whose display order is determined; a voice generating unit that reads out a text based on a comment posted in correspondence with two or more consecutive content units in the image content, and associates the text with the two or more consecutive content units; a content display control unit that displays the image content and an interface that designates a reference unit content of the image content; an audio playback unit that plays back the audio associated with the two or more consecutive unit contents that include the reference time point indicated by the interface therebetween.
4. The content management system according to claim 1 , wherein the audio playback unit reduces the playback volume of the audio of the image content while playing back the audio.
5. The content management system according to claim 1 , wherein the audio playback unit stops playback of the audio of the image content while the audio is being played back.
6. The content management system according to claim 1 , wherein the audio playback unit plays back the audio when the reference time point indicated by the interface indicates a time point that is included in the two time points for a predetermined time period or more.
7. The content management system according to claim 1 , wherein the audio playback unit plays back the audio when the reference time indicated by the interface remains the same for a predetermined period of time or more.
8. The content management system according to claim 1 , wherein the voice generating unit uses, as the text, a sentence summarizing comments posted corresponding to points in time between the two points in time.
9. a thumbnail generation unit that generates thumbnail images corresponding to two time points in the image content and associates the generated thumbnail images with the two time points; The content display control unit The content management system according to claim 1 , wherein the interface displays thumbnail images associated with two time points between which the reference time point indicated by the interface is located.
10. one or more hardware processors, A first step of delivering image content that is a moving image; a second step of associating a voice reading a text based on a comment posted corresponding to a time point included between two time points in the image content with the two time points; a third step of displaying the image content and an interface for specifying a reference time point of the image content; a fourth step of playing the audio associated with two points in time that include the reference point indicated by the interface.
11. one or more hardware processors, a first step of delivering audio content associated with image content; a second step of associating a voice reading a text based on a comment posted corresponding to a time point included between two time points in the image content with the two time points; a third step of displaying the image content and an interface for specifying a reference time point of the image content; a fourth step of playing the audio associated with two points in time that include the reference point indicated by the interface.
12. one or more hardware processors, a first step of distributing image content including a plurality of unit contents whose display order is determined; a second step of associating a voice reading out text based on comments posted corresponding to two or more consecutive content units included in the image content with the two or more consecutive content units; a third step of displaying the image content and an interface for specifying a reference unit content of the image content; a fourth step of playing back the audio associated with the two or more consecutive content units that include the reference time point indicated by the interface therebetween.
13. a content receiving unit that receives image content, which is a moving image, and audio; a content display control unit that displays the image content and an interface that specifies a reference time point of the image content; an audio playback unit that plays back the audio associated with two time points that include a reference time point indicated by the interface; A content display device, wherein the audio is a reading of text based on comments posted corresponding to points in time between two points in time in the image content.
14. a content receiving unit that receives audio content associated with image content and audio; a content display control unit that displays the audio content and an interface that specifies a reference time point of the audio content; an audio playback unit that plays back the audio associated with two time points that include a reference time point indicated by the interface; A content display device, wherein the audio is a reading of text based on comments posted corresponding to points in time between two points in time in the audio content.
15. a content receiving unit for receiving image content including a plurality of unit contents whose display order is determined, and audio; a content display control unit that displays the image content and an interface that specifies a reference time point of the image content; an audio playback unit that plays back the audio associated with two or more consecutive content units that include the reference time point indicated by the interface, A content display device, wherein the audio is a reading of text based on comments posted in association with unit content included in two or more consecutive unit content in the image content.
16. one or more hardware processors, receiving image content, which is a moving image, and audio, the audio being a reading of text based on comments posted corresponding to time points in the image content that are included between two time points; displaying the image content and an interface for specifying a reference time point of the image content; and playing back the audio associated with two points in time that include the reference point in time indicated by the interface.
17. one or more hardware processors, receiving audio content associated with image content and audio, the audio being a reading of text based on comments posted corresponding to time points in the audio content that are included between two time points; displaying the audio content and an interface for specifying a reference time point of the audio content; and playing back the audio associated with two points in time that include the reference point in time indicated by the interface.
18. one or more hardware processors, receiving image content including a plurality of unit contents whose display order is determined, and audio, wherein the audio is a reading of text based on comments posted corresponding to the unit contents included in two or more unit contents in the image content; displaying the image content and an interface for specifying a reference time point of the image content; and playing back the audio associated with two or more unit contents that include the reference time point indicated by the interface therebetween.
Citation Information
Patent Citations
Comment distribution system, terminal, comment output method, and program
JP2010114571A
Video content receiver, comment exchange system, method for generating comment data, and program
JP2010232716A
Comment delivery system, terminal device, comment delivery method, and program
JP2011166833A
Voice output control device, voice output control method, program, and recording medium
JP2014011509A
Server system
JP2020089716A