System and data manufacturing method
The system identifies and extracts characteristic video portions by analyzing viewer and performer actions, providing accurate and high-accuracy edited video data.
Patent Information
- Application Number
- JP2024047389
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-23
- Publication Date
- 2025-10-06
AI Technical Summary
Existing systems fail to effectively identify characteristic portions in video data, such as exciting, topic, or standard content, based on viewer and performer actions.
A system and method that utilizes computers to acquire action information from viewers and performers during video distribution, determine start and end points of video portions based on this information, and generate edited video data highlighting these portions.
Enables accurate identification of characteristic video portions by analyzing viewer and performer actions, allowing for high-accuracy extraction and presentation of exciting or significant content.
Smart Images

Figure 2025147164000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a system and a method for producing data. [Background technology]
[0002] As an invention relating to a conventional system, for example, a program described in Patent Document 1 is known. In this program, a communication unit receives a video distributed from a server. A display unit displays the distributed video. A control unit acquires first information (gift information) based on a first input by a user of the terminal to the display unit on which the video is displayed. The control unit acquires second information (timestamp) related to the video based on the first information (gift information). The display unit displays the second information. Based on an input by the user of the terminal to the second information (timestamp), the control unit plays back on the display unit a portion of the video corresponding to the first input. This allows the user to easily review the scene in which the gift was sent. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 7128338 Summary of the Invention [Problem to be solved by the invention]
[0004] Incidentally, there is a demand for identifying characteristic parts of video data.
[0005] SUMMARY OF THE INVENTION It is therefore an object of the present invention to provide a system and a data production method that can identify characteristic portions in video data. [Means for solving the problem]
[0006] The first form is A system comprising one or more computers, The one or more computers Obtain video data distributed by a distributor to one or more viewers; acquire action information indicating performer actions by performers when the video data is distributed and / or viewer actions by the one or more viewers with respect to the video data when the video data is distributed; determining a start point and an end point of a first portion of the video data based on the action information; It is a system.
[0007] The second form is the one or more computers are one or more servers; The system according to the first aspect.
[0008] The third form is The first portion is an exciting portion in the video in which the performer or the one or more viewers are excited, a topic portion in the video containing content related to a specific topic, or a standard portion in the video containing standard content. The system according to either the first aspect or the second aspect.
[0009] The fourth form is the action information is viewer action information indicating the viewer actions taken by the one or more viewers, the viewer action is an action taken by each of the one or more viewers to operate one or more viewer terminals, a video of the video data including a viewer action indication indicating the viewer actions by the one or more viewers; the one or more computers determine a start point and an end point of the first portion of the video data based on the viewer action information; The system according to the first or second aspect.
[0010] The fifth form is The viewer action information is information showing the viewer's goodwill toward the performer, information showing a comment from the viewer to the performer, or information showing a gift given by the viewer to the performer. The system according to the fourth aspect.
[0011] The sixth form is In the process of determining the start point and the end point of the first portion of the video data, the one or more computers generate a count result of counting the frequency of occurrence of the viewer action based on the viewer action information, and determine the start point and the end point of the first portion of the video data based on the count result. The system according to the fourth or fifth aspect.
[0012] The seventh form is the action information is performer action information indicating the performer action by the performer, In the process of acquiring the action information, the one or more computers generate the performer action information based on sounds included in the video of the video data and / or actions of the performers included in the video of the video data; the one or more computers determine a start point and an end point of the first portion of the video data based on the performer action information; The system is according to any one of the first to sixth aspects.
[0013] The eighth form is In the process of acquiring the action information, the one or more computers generate the performer action information by extracting keywords from sounds included in the video of the video data. The system according to the seventh aspect.
[0014] The ninth form is The action information is associated with action occurrence time information indicating the time at which the performer action and / or the viewer action occurred, In the process of determining the start point and the end point of the first portion of the video data, the one or more computers determine the start point and the end point of the first portion of the video data based on the action information and the action occurrence time information. The system according to any one of the first to eighth aspects.
[0015] The tenth form is In the process of determining the start point and the end point of the first portion of the video data, the one or more computers generate time distribution information indicating a relationship between time and a parameter related to the number of occurrences of the performer actions or the number of occurrences of the viewer actions, based on the action information and the action occurrence time information, and determine the start point and the end point of the first portion of the video data based on the time distribution information. The system according to the ninth aspect.
[0016] The 11th form is In the process of determining the start point and the end point of the first portion of the video data, the one or more computers determine the start point and the end point of a period in which the parameter of the time distribution information is greater than a predetermined value as the start point and the end point of the first portion of the video data, respectively. The system according to the tenth aspect.
[0017] The 12th form is M is a natural number, The viewer actions include a first viewer action through an M-th viewer action, In the process of determining the start point and the end point of the first portion of the video data, the one or more computers calculate the parameters by weighting the first viewer action through the Mth viewer action, respectively, according to the contents of the first viewer action through the Mth viewer action. The system according to either the tenth or eleventh aspect.
[0018] The 13th form is the one or more computers extract first edited video data corresponding to the first portion from the video data; The system according to any one of the first to twelfth aspects.
[0019] The 14th form is N is an integer equal to or greater than 2, In the process of determining the start point and the end point of the N-th portion of the video data, the one or more computers determine the start point and the end point of the N-th portion of the video data based on the action information; The one or more computers extracting N-th edited video data corresponding to the N-th portion from the video data; generating edited video data by combining the first edited video data through the Nth edited video data; The system according to the thirteenth aspect.
[0020] The 15th form is 1. A data production method, comprising: The one or more computers Acquire video data distributed by a distributor to viewers, acquire action information indicating performer actions by performers when the video data is distributed and / or viewer actions by the one or more viewers with respect to the video data when the video data is distributed; extracting a portion of the video data based on the action information to generate first edited video data having a structure that allows a computer to reproduce the portion of the video data; It is a data production method. [Effects of the Invention]
[0021] According to the present disclosure, it is possible to identify characteristic parts of video data. [Brief explanation of the drawings]
[0022] [Figure 1] FIG. 1 is a block diagram of systems 1, 1a to 1d. [Figure 2] FIG. 2 shows the images displayed on the viewer terminals 210-1 to 210-n. [Figure 3] FIG. 3 shows a viewer action table stored in the distributor terminal 10. As shown in FIG. [Figure 4] FIG. 4 shows a weighting table stored in the distributor terminal 10. [Figure 5] FIG. 5 is an explanatory diagram of the operation of the server 110. [Figure 6] FIG. 6 is a block diagram of the distributor terminal 10. [Figure 7] FIG. 7 is a block diagram of the server 110. [Figure 8] FIG. 8 is a block diagram of the viewer terminal 210-1. [Figure 9] FIG. 9 is a flowchart showing the operation of the control unit 112 of the server 110. [Figure 10] FIG. 10 is a flowchart showing the operation of the control unit 112 of the server 110. [Figure 11] FIG. 11 is a flowchart showing the operation of the control unit 112 of the server 110. [Figure 12] FIG. 12 shows the images displayed on the viewer terminals 210-1 to 210-n. [Figure 13] FIG. 13 is an explanatory diagram of the operation of the server 110. [Figure 14] FIG. 14 is a flowchart showing the operation of the control unit 112 of the server 110. [Figure 15] FIG. 15 is a flowchart showing the operation of the control unit 112 of the server 110. [Figure 16] FIG. 16 is a flowchart showing the operation of the control unit 112 of the server 110. [Figure 17] FIG. 17 shows the images displayed on the viewer terminals 210-1 to 210-n. [Figure 18] FIG. 18 is a flowchart showing the operation of the control unit 112 of the server 110. [Figure 19] FIG. 19 is an explanatory diagram of the operation of the server 110. [Figure 20] FIG. 20 is a flowchart showing the operation of the control unit 112 of the server 110. [Figure 21] FIG. 21 is a flowchart showing the operation of the control unit 112 of the server 110. DETAILED DESCRIPTION OF THE INVENTION
[0023] (Embodiment) A system 1 according to an embodiment of the present disclosure will be described with reference to the drawings.
[0024] [System 1 Overview] First, the overall configuration of the system 1 will be described with reference to the drawings. Figure 1 is a block diagram of the systems 1, 1a to 1d.
[0025] The system 1 shown in Fig. 1 includes a distributor terminal 10, a server 110, and viewer terminals 210-1 to 210-n, where n is a natural number. The distributor terminal 10, the server 110, and the viewer terminals 210-1 to 210-n can communicate with each other via a communication network. The network may be the Internet, an intranet, or the like.
[0026] The distributor terminal 10 is an information processing device used by a distributor of video content. The distributor terminal 10 is, for example, a smartphone, a tablet terminal, a home game console, a portable game console, or a personal computer.
[0027] The viewer terminals 210-1 to 210-n are information processing devices used by multiple viewers. The viewer terminals 210-1 to 210-n are, for example, smartphones, tablet terminals, home game consoles, portable game consoles, or personal computers.
[0028] Server 110 is an information processing device used by the operator of a video distribution service. Server 110 acquires video data from distributor terminal 10 and distributes the video data to viewer terminals 210-1 to 210-n. Server 110 is a computer.
[0029] [System 1 Operation Overview] Next, an overview of the operation of the system 1 will be described. Fig. 2 shows images displayed on the viewer terminals 210-1 to 210-n. Fig. 3 shows a viewer action table stored in the distributor terminal 10. Fig. 4 shows a weighting table stored in the distributor terminal 10. Fig. 5 is an explanatory diagram of the operation of the server 110.
[0030] In System 1, the following two actions are performed: (1) Watching videos (2) Generation of edited video data
[0031] (1) Watching videos When watching a video, a viewer watches a video in which performers appear. More specifically, distributor terminal 10 photographs the performers shown in FIG. 2 and superimposes images of avatars onto the performers. Note that, for ease of explanation, the avatars will be referred to as performers hereinafter. In this way, distributor terminal 10 generates video data D0. Then, distributor terminal 10 transmits video data D0 to server 110. Server 110 distributes video data D0 to viewer terminals 210-1 to 210-n. In this way, viewer terminals 210-1 to 210-n receive video data D0. Then, multiple viewers can watch a video including images of the performers shown in FIG. 2 using viewer terminals 210-1 to 210-n.
[0032] Multiple viewers can perform viewer actions while watching a video. A viewer action is an action taken by multiple viewers on the video data D0 when the video data D0 is distributed. Specific examples of viewer actions will be described below with reference to FIG. 2.
[0033] The image shown in FIG. 2 includes a heart button, a gift button, and a text input box. The heart button is called the "Like" button. The gift button is called the "Gift" button. The text input box is called the "Input" box. Viewer actions include a viewer pressing the "Like" button, a viewer pressing the "Gift" button, and a viewer sending a comment. When a viewer presses the "Like" button, a heart button is displayed on viewer terminals 210-1 to 210-n, as shown in the image in FIG. 2. When a viewer presses the "Gift" button, a gift button is displayed on viewer terminals 210-1 to 210-n, as shown in the image in FIG. 2, and the viewer gives a gift to the performer. The gift may be, for example, data for an item to liven up the video, data for currency that can be used in a video distribution service, or data for a coupon that can be exchanged for an item in a video distribution service. The viewer may also select a gift from multiple options. When a viewer enters a comment in the input box and presses the "Send" button, the comment is displayed on viewer terminals 210-1 to 210-n, as shown in the image in FIG. 2.
[0034] When a viewer action occurs, viewer terminals 210-1 to 210-n transmit viewer action information indicating the viewer action to server 110. The viewer action information includes the viewer's account and the type of viewer action. In this way, server 110 acquires the viewer action information.
[0035] Server 110 stores the viewer action table shown in Fig. 3. The viewer action table records the viewer's account name, the type of viewer action, and the time when the action occurred. The time indicates the elapsed time from the start of video data D0. The action occurrence time indicates the time when the viewer action occurred. When video data D0 is distributed, server 110 records the viewer's account name, the type of viewer action, and the time when the action occurred in the viewer action table based on the viewer action information acquired from viewer terminals 210-1 to 210-n.
[0036] (2) Generation of edited video data In generating edited video data, server 110 extracts a portion of video data D0 to generate edited video data DX. Edited video data DX has a structure that allows a computer such as distributor terminal 10 or viewer terminals 210-1 to 210-N to play back that portion of video data D0. More specifically, server 110 stores the weighting table shown in FIG. 4. The weighting table records types of viewer actions and points. Points are parameters assigned to viewer actions. Points indicate the level of excitement of a viewer action. Viewer actions that are highly exciting to the viewer are assigned high points. Viewer actions that are less exciting to the viewer are assigned low points.
[0037] After the video data D0 is distributed, the server 110 generates the relationship between the excitement parameter P and time shown in FIG. 5 based on the weighting table shown in FIG. 4 and the viewer action table shown in FIG. 3. The excitement parameter P indicates the level of excitement at each time point in the video. The excitement level depends on the activity of the actions taken by the viewer and / or broadcaster, the number of viewers, etc. The more active the actions taken by the viewer and / or broadcaster, the higher the level of excitement. The larger the number of viewers, the higher the level of excitement. The activity of the actions taken by the viewer and / or broadcaster depends on the number of actions taken by the viewer and / or broadcaster per unit time and the magnitude of the actions taken by the viewer and / or broadcaster. The magnitude of the actions taken by the viewer and / or broadcaster refers to the volume of the viewer's and / or broadcaster's voice, the frequency of uttering keywords, and the magnitude of the viewer's and / or broadcaster's movements. The server 110 then identifies the first portion A1 and the second portion A2 having an excitement parameter P greater than a predetermined value p0. The first portion A1 and the second portion A2 are portions that multiple viewers enjoyed. The server 110 extracts the first portion A1 to generate first edited video data DX1, and extracts the second portion A2 to generate second edited video data DX2. The server 110 generates edited video data DX by combining the first edited video data DX1 and the second edited video data DX2.
[0038] [Structure of the distributor terminal 10] The structure of the distributor terminal 10 will be described with reference to the drawings. Figure 6 is a block diagram of the distributor terminal 10.
[0039] As shown in FIG. 6, the distributor terminal 10 includes a control unit 12, a memory unit 14, a network interface 16, a camera 17, a graphics processing unit 18, a display 20, an audio processing unit 22, a speaker 24, and an operation unit 26.
[0040] The storage unit 14 stores programs and data and is, for example, a combination of a read-only memory (ROM), a random access memory (RAM), and a storage (for example, a flash memory or a hard disk).
[0041] The programs include, for example, the following programs: OS (Operating System) programs - Programs for applications that process information (e.g., web browsers or target apps described below)
[0042] The data includes, for example, the following data: Databases referenced in information processing Data obtained by performing information processing (i.e., the results of performing information processing)
[0043] The control unit 12 executes the programs stored in the storage unit 14 to realize the functions of the distributor terminal 10. The control unit 12 is, for example, at least one of the following: ·CPU(Central Processing Unit) ·GPU(Graphic Processing Unit) ·ASIC(Application Specific Integrated Circuit) ·FPGA(Field Programmable Array)
[0044] The network interface 16 controls communication between the distributor terminal 10 and an external device. The external device is a server 110.
[0045] The camera 17 captures the surroundings of the distributor terminal 10 and generates video data. In this embodiment, the camera 17 captures the performers.
[0046] The graphics processing unit 18 displays an image on the display 20 based on the image data generated by the control unit 12. The display 20 is a liquid crystal display or an organic EL (Electro Luminescence) display.
[0047] The audio processing unit 22 causes the speaker 24 to output sound based on the audio data generated by the control unit 12 .
[0048] The operation unit 26 generates an operation signal based on an operation by the distributor, and outputs the operation signal to the control unit 12.
[0049] [Server 110 Structure] Next, the structure of the server 110 will be described with reference to the drawings.
[0050] The server 110 is capable of communicating with the distributor terminal 10 and the viewer terminals 210-1 to 210-n. As shown in FIG.
[0051] The storage unit 114 stores programs and data and is, for example, a combination of a read-only memory (ROM), a random access memory (RAM), and a storage (for example, a flash memory or a hard disk). The programs include, for example, the following programs: OS (Operating System) programs - Programs for applications that process information (e.g., web browsers or target apps described below)
[0052] The data includes, for example, the following data: Databases referenced in information processing Data obtained by performing information processing (i.e., the results of performing information processing)
[0053] The control unit 112 executes the programs stored in the storage unit 114 to realize the functions of the server 110. The control unit 112 is, for example, at least one of the following: ·CPU(Central Processing Unit) ·GPU(Graphic Processing Unit) ·ASIC(Application Specific Integrated Circuit) ·FPGA(Field Programmable Array)
[0054] The control unit 112 includes a data acquisition unit 120, an information acquisition unit 122, a start point / end point determination unit 124, an extraction unit 126, and a combination unit 128 as functional blocks.
[0055] The network interface 116 controls communication between the server 110 and external devices. The external devices are the distributor terminal 10 and the viewer terminals 210-1 to 210-n.
[0056] [Structure of viewer terminals 210-1 to 210-n] Next, the structure of the viewer terminals 210-1 to 210-n will be described with reference to the drawings. Figure 8 is a block diagram of the viewer terminal 210-1.
[0057] As shown in FIG. 8, the viewer terminal 210-1 includes a control unit 212, a memory unit 214, a network interface 216, a camera 217, a graphics processing unit 218, a display 220, an audio processing unit 222, a speaker 224, and an operation unit 226.
[0058] The storage unit 214 stores programs and data and is, for example, a combination of a read-only memory (ROM), a random access memory (RAM), and storage (for example, a flash memory or a hard disk).
[0059] The programs include, for example, the following programs: OS (Operating System) programs - Programs for applications that process information (e.g., web browsers or target apps described below)
[0060] The data includes, for example, the following data: Databases referenced in information processing Data obtained by performing information processing (i.e., the results of performing information processing)
[0061] The control unit 212 realizes the functions of the viewer terminal 210-1 by executing the program stored in the storage unit 214. The control unit 212 is, for example, at least one of the following: ·CPU(Central Processing Unit) ·GPU(Graphic Processing Unit) ·ASIC(Application Specific Integrated Circuit) ·FPGA(Field Programmable Array)
[0062] The network interface 216 controls communication between the viewer terminal 210-1 and an external device, which is the server 110.
[0063] The camera 217 captures the surroundings of the viewer terminal 210-1 and generates video data.
[0064] The graphics processing unit 218 displays an image on the display 220 based on the image data generated by the control unit 212. The display 220 is a liquid crystal display or an organic EL (Electro Luminescence) display.
[0065] The audio processing unit 222 causes the speaker 224 to output sound based on the audio data generated by the control unit 212 .
[0066] The operation unit 226 generates an operation signal based on the viewer's operation, and outputs the operation signal to the control unit 212 .
[0067] The structure of the viewer terminals 210-2 to 210-n is the same as that of the viewer terminal 210-1, so a description thereof will be omitted.
[0068] [System 1 operation] Next, the operation of the system 1 will be described with reference to the drawings. Figures 9 to 11 are flowcharts showing the operation of the control unit 112 of the server 110.
[0069] When the control unit 112 of the server 110 reads out the programs stored in the memory unit 114, these programs cause the server 110 to execute the operations described below. The programs then cause the server 110 to function as a data acquisition unit 120, an information acquisition unit 122, a start point / end point determination unit 124, an extraction unit 126, and a combination unit 128.
[0070] (1) Watching videos This process is started when the distributor operates the operation unit 26 of the distributor terminal 10. The control unit 12 of the distributor terminal 10 photographs the performers using the camera 17 and superimposes avatar images onto the performers. To superimpose the avatar images, the control unit 12 of the distributor terminal 10 may use motion capture. As a result, the control unit 12 of the distributor terminal 10 generates video data D0. Furthermore, the control unit 12 of the distributor terminal 10 transmits the video data D0 to the server 110 via the network interface 16. In response, the network interface 116 of the server 110 receives the video data D0 and outputs the video data D0 to the control unit 112. As a result, the control unit 112 of the server 110 acquires the video data D0.
[0071] Next, control unit 112 of server 110 starts distribution of video data D0 to viewer terminals 210-1 to 210-n (step S1). Until the distribution of video data D0 is completed, control unit 112 of server 110 continues to acquire video data D0 from distributor terminal 10 and continues to distribute video data D0 to viewer terminals 210-1 to 210-n.
[0072] When distribution of video data D0 starts, control unit 112 of server 110 determines whether viewer action information indicating a viewer action has been acquired (step S2). A viewer action is an action in which multiple viewers operate viewer terminals 210-1 to 210-n. Viewer actions include, for example, an action in which a viewer presses the like button, an action in which a viewer presses the gift button, and an action in which a viewer sends a comment. Thus, viewer action information is information indicating a viewer's goodwill toward a performer, information indicating a comment from a viewer to a performer, or information indicating a gift given by a viewer to a performer. A viewer's goodwill toward a performer is an action in which a viewer presses the like button.
[0073] Viewer terminals 210-1 to 210-n generate viewer action information in response to the viewer action. Viewer terminals 210-1 to 210-n transmit the viewer action information to server 110 via network interface 216. In response, network interface 116 of server 110 receives the viewer action information and outputs the viewer action information to control unit 112. Control unit 112 of server 110 then acquires the viewer action information, and the process proceeds to step S3. On the other hand, if the viewer action information has not been generated, control unit 112 of server 110 does not acquire the viewer action information, and the process proceeds to step S6.
[0074] If control unit 112 of server 110 does not acquire viewer action information, control unit 112 of server 110 transmits video data D0 to viewer terminals 210-1 to 210-n via network interface 116 (step S6). As a result, viewer terminals 210-1 to 210-n acquire video data D0. As shown in FIG. 2, viewer terminals 210-1 to 210-n display video of video data D0 that does not include viewer action indicators showing viewer actions by multiple viewers. The viewer action indicators are heart marks, gift marks, and comments. After this, the process returns to step S2.
[0075] When the control unit 112 of the server 110 acquires the viewer action information, the control unit 112 of the server 110 generates composite video data D0 by combining the video of the video data D0 with viewer action indicators showing viewer actions. The control unit 112 of the server 110 then transmits the composite image video data D0 to the viewer terminals 210-1 to 210-n via the network interface 116 (step S3). As a result, the viewer terminals 210-1 to 210-n acquire the composite video data D0. The viewer terminals 210-1 to 210-n display the video of the video data D0, which includes viewer action indicators showing viewer actions taken by multiple viewers, as shown in FIG. 2.
[0076] Next, the control unit 112 of the server 110 records the viewer action information in the viewer action table shown in Fig. 4 (step S4). More specifically, the control unit 112 of the server 110 records the viewer account and the type of viewer action included in the viewer action information in the viewer action table. Furthermore, the control unit 112 of the server 110 records the time at which the viewer action information is acquired as the action occurrence time in the viewer action table. In this way, the action information is associated with action occurrence time information that indicates the action occurrence time at which the viewer action occurred.
[0077] Next, the control unit 112 of the server 110 determines whether or not to complete the distribution of the video data D0 (step S5). Specifically, the control unit 112 of the server 110 determines whether or not the transmission of the video data D0 from the distributor terminal 10 has stopped. If the transmission of the video data D0 has stopped, the control unit 112 of the server 110 completes the distribution of the video data D0. In this case, the process proceeds to step S5. On the other hand, if the transmission of the video data D0 has not stopped, the control unit 112 of the server 110 does not complete the distribution of the video data D0. In this case, the process returns to step S2. Then, the process repeats steps S2 to S6 until the distribution of the video data D0 is completed. By repeating steps S2 to S6, the control unit 112 (information acquisition unit 122) of the server 110 acquires action information indicating viewer actions taken by multiple viewers on the video data D0 during the distribution of the video data D0.
[0078] When the distribution of the video data D0 is completed, the control unit 112 of the server 110 stores the video data D0 in the storage unit 114 (step S7). After the above processing, this processing ends.
[0079] (2) Generation of edited video data This process is started when the distributor operates the operation unit 26 of the distributor terminal 10. The control unit 12 of the distributor terminal 10 transmits a data generation instruction indicating that edited video data DX is to be generated to the server 110 via the network interface 16. In response, the network interface 116 of the server 110 receives the data generation instruction and outputs the data generation instruction to the control unit 112. As a result, the control unit 112 of the server 110 acquires the data generation instruction (step S11).
[0080] Next, the control unit 112 of the server 110 reads out the video data D0 stored in the storage unit 114. As a result, the control unit 112 of the server 110 acquires the video data D0 distributed by the distributor to multiple viewers (step S12).
[0081] Next, the control unit 112 of the server 110 sets N to 1 (step S13), where N is a natural number. Then, the control unit 112 of the server 110 determines the start point SPN and end point EPN of the Nth portion AN of the video data D0 based on the viewer action information (step S14). Step S14 will be described in more detail below with reference to FIG. 11.
[0082] First, the control unit 112 (start and end point determination unit 124) of the server 110 generates a counting result of counting the frequency of viewer actions based on the viewer action information, as shown in FIG. 5 (step S101). In this embodiment, the control unit 112 (start and end point determination unit 124) of the server 110 generates time distribution information indicating the relationship between excitement parameter P and time based on the viewer action information and action occurrence time information. The excitement parameter P is an example of a parameter related to the number of viewer actions that occur. The time distribution information corresponds to the counting result. The excitement parameter P and the time distribution information will be described below.
[0083] The viewer actions include an action of a viewer pressing the like button (first viewer action), an action of a viewer sending a comment (second viewer action), and an action of a viewer pressing a gift button (third viewer action). The control unit 112 (start and end point determination unit 124) of the server 110 calculates the excitement parameter P by weighting the action of a viewer pressing the like button (first viewer action), the action of a viewer sending a comment (second viewer action), and the action of a viewer pressing a gift button (third viewer action) according to their contents. Specifically, the weighting table shown in FIG. 3 records points according to the contents of the viewer action. The control unit 112 (start and end point determination unit 124) of the server 110 weights each of the multiple viewer actions shown in FIG. 3 using the weighting table shown in FIG. 4. Then, the control unit 112 (start and end point determination unit 124) of the server 110 arranges the weighted multiple viewer actions in order of the action occurrence time. The control unit 112 (start and end point determination unit 124) of the server 110 generates time distribution information by calculating a moving average of the weighted multiple viewer actions. The period of the moving average is, for example, 10 seconds.
[0084] Next, the control unit 112 (start and end point determination unit 124) of the server 110 determines the start point SPN and end point EPN of the Nth portion AN of the video data D0 based on the counting result generated in step S101 (step S102). Specifically, as shown in Fig. 5, the control unit 112 (start and end point determination unit 124) of the server 110 determines the start point and end point, respectively, of a period in which the excitement parameter P of the time distribution information is greater than a predetermined value p0 as the start point SPN and end point EPN of the Nth portion of the video data D0. In this way, the control unit 112 (start and end point determination unit 124) of the server 110 determines the start point SPN and end point EPN of the Nth portion AN of the video data D0 based on the action information and action occurrence time information.
[0085] Next, the control unit 112 (start and end point determination unit 124) of the server 110 corrects the start point SPN and end point EPN of the Nth portion AN of the video data D0 (step S103). More specifically, the start point SPN of the Nth portion AN may not coincide with the start of the conversation between the performers. Similarly, the end point EPN of the Nth portion AN may not coincide with the end of the conversation between the performers. Therefore, in step S103, the control unit 112 (start and end point determination unit 124) of the server 110 corrects the start point SPN of the Nth portion AN of the video data D0 so that the start point SPN coincides with the start of the conversation between the performers. Similarly, the control unit 112 (start and end point determination unit 124) of the server 110 corrects the end point EPN of the Nth portion AN of the video data D0 so that the start point SPN coincides with the end of the conversation between the performers.
[0086] Specifically, as shown in FIG. 5, the control unit 112 (start and end point determination unit 124) of the server 110 analyzes the voices of the performers in the video data D0 to determine the time when the performers start talking near the start point SPN as the corrected start point SPaN. The time when the performers start talking is, for example, the end point of a predetermined time when the volume of the performers' voices is zero for that predetermined time. Furthermore, the control unit 112 (start and end point determination unit 124) of the server 110 analyzes the voices of the performers in the video data D0 to determine the time when the performers end talking near the end point EPN as the corrected end point EPaN. The time when the performers end talking is, for example, the start point of a predetermined time when the volume of the performers' voices is zero for that predetermined time. Furthermore, the control unit 112 (start and end point determination unit 124) of the server 110 may determine the corrected start point SPaN and the corrected end point EPaN by using AI (Artificial Intelligence) such as LLM (Large Language Model) to determine the start and end of a conversation. In this case, the control unit 112 (start and end point determination unit 124) of the server 110 generates text data based on the voices of the performers. Then, the control unit 112 (start and end point determination unit 124) of the server 110 identifies the start and end of the conversation based on the text data. With the above processing, step S14 ends, and the process proceeds to step S15.
[0087] Next, the control unit 112 (extraction unit 126) of the server 110 extracts the Nth edited video data DXN corresponding to the Nth portion AN between the corrected start point SPaN and the corrected end point EPaN from the video data D0 (step S15). The control unit 112 (extraction unit 126) of the server 110 then determines whether or not a period in which the excitement parameter P is greater than a predetermined value p0 exists at a time after the Nth portion AN (step S16). If a period in which the excitement parameter P is greater than the predetermined value p0 exists, the process proceeds to step S17. If a period in which the excitement parameter P is greater than the predetermined value p0 does not exist, the process proceeds to step S18.
[0088] If there is a period during which excitement parameter P is greater than predetermined value p0, control unit 112 of server 110 increments N by 1 (step S17). After this, the process returns to step S14. Then, by repeating steps S14 to S17, control unit 112 of server 110 extracts first edited video data DX1 through Nth edited video data DXN corresponding to first portion A1 through Nth portion AN from video data D0.
[0089] If there is no period in which the excitement parameter P is greater than the predetermined value p0, the control unit 112 of the server 110 generates edited video data DX by combining the first edited video data DX1 through the Nth edited video data DXN (step S18). In this way, the control unit 112 of the server 110 extracts the first edited video data DX1 through the Nth edited video data DXN (portions) of the video data D0 based on the action information, thereby generating edited video data DX having a structure that allows a computer to play back the first edited video data DX1 through the Nth edited video data DXN (portions) of the video data D0. This completes the process.
[0090] The edited video data DX as described above is stored in the storage unit 114 of the server 110. The distributor terminal 10 and the viewer terminals 210-1 to 210-n can then obtain the edited video data DX from the server 110 by accessing a web page managed by the server 110. This allows the distributor and multiple viewers to view the edited video data DX.
[0091] Furthermore, the distributor can advertise the video data D0 by pasting the URL of the link destination of the edited video data DX on the web page into a page of a social networking service (SNS).
[0092] [effect] According to the system 1, it is possible to identify characteristic portions of the video data D0. More specifically, the control unit 112 of the server 110 acquires viewer action information indicating viewer actions taken by multiple viewers on the video data D0 when the video data D0 is distributed. Then, the control unit 112 of the server 110 determines the start point SPN and end point EPN of the Nth portion of the video data D0 based on the viewer action information. Such an Nth portion AN is a portion of the video data D0 where many viewer actions occurred. Therefore, the Nth portion AN is a characteristic portion of the video data D0. For the above reasons, according to the system 1, it is possible to identify characteristic portions of the video data D0.
[0093] According to the system 1, characteristic portions of the video data D0 can be identified with high accuracy. The control unit 112 of the server 110 generates an aggregation result that aggregates the frequency at which viewer actions occur based on the viewer action information. Furthermore, the control unit 112 of the server 110 determines the start SPN and end EPN of the Nth portion AN of the video data D0 based on the aggregation result. In this way, the aggregation result that aggregates the frequency at which viewer actions occur is used to determine the start SPN and end EPN of the Nth portion AN. By using the quantified aggregation result, characteristic portions of the video data D0 can be identified with high accuracy.
[0094] According to the system 1, characteristic portions of the video data D0 can be identified with high accuracy. The control unit 112 of the server 110 determines the start point SPN and end point EPN of the Nth portion AN of the video data D0 based on the viewer action information and the action occurrence time information. This allows the control unit 112 of the server 110 to accurately identify the Nth portion AN corresponding to the period in which many viewer actions occurred, thereby enabling ... characteristic portions of the video data D0.
[0095] According to the system 1, characteristic portions of the video data D0 can be identified with high accuracy. The control unit 112 of the server 110 generates time distribution information indicating the relationship between the excitement parameter P (a parameter related to the number of viewer actions) and time based on the viewer action information and the action occurrence time information, and determines the start point SPN and end point EPN of the Nth portion AN of the video data D0 based on the time distribution information. In this way, by using the quantified time distribution information, characteristic portions of the video data D0 can be identified with high accuracy.
[0096] According to the system 1, characteristic portions of the video data D0 can be identified with high accuracy. The control unit 112 of the server 110 calculates the excitement parameter P by weighting the actions of the viewer pressing the like button, the viewer sending a comment, and the viewer pressing the gift button according to the content of the action. This allows the content of the viewer action to be reflected in the excitement parameter P. Therefore, the excitement parameter P can more appropriately indicate the time of an important portion in the video data D0. By using this excitement parameter P to determine the start point SPN and end point EPN of the Nth portion AN, characteristic portions of the video data D0 can be identified with high accuracy.
[0097] (First Modification) The system 1a according to the first modified example will be described below with reference to the drawings. Fig. 12 shows images displayed on the viewer terminals 210-1 to 210-n. Fig. 13 is an explanatory diagram of the operation of the server 110. Figs. 14 and 15 are flowcharts showing the operation of the control unit 112 of the server 110.
[0098] System 1a differs from System 1 in that the control unit 112 of the server 110 determines the start point SPN and end point EPN of the Nth portion AN of the video data D0 based on performer action information. Performer action information is information that indicates performer actions by performers. The following explanation will focus on these differences.
[0099] [System 1a operation overview] In the system 1a, the following two operations are performed. (1) Watching videos (2) Generation of edited video data
[0100] (1) Watching videos Viewing of videos in the system 1a is the same as viewing of videos in the system 1, so a description thereof will be omitted.
[0101] (2) Generation of edited video data In generating edited video data, the server 110 extracts a portion of the video data D0 to generate edited video data DX, and transmits the edited video data DX to the distributor terminal 10. More specifically, the performers perform performer actions in the video. Examples of performer actions include a performer screaming action, a performer laughing loudly, and a performer shouting loudly, as shown in FIG.
[0102] Therefore, as shown in FIG. 13, the server 110 analyzes the voices of the performers in the video data D0 to generate a graph showing the relationship between the performer's vocal volume and time. The server 110 then determines that a performer action occurred during a period in which the performer's volume was greater than a predetermined value p1. The server 110 then identifies a first portion A1 and a second portion A2 in which the performer's volume was greater than the predetermined value p1. The first portion A1 and the second portion A2 are exciting portions in the video in which the performers were excited. The server 110 extracts the first portion A1 to generate first edited video data DX1, and extracts the second portion A2 to generate second edited video data DX2. The server 110 generates edited video data DX by combining the first edited video data DX1 and the second edited video data DX2.
[0103] [System 1a operation] Next, the operation of the system 1a will be described with reference to the drawings. Figures 14 and 15 are flowcharts showing the operation of the control unit 112 of the server 110.
[0104] This process is started when a performer operates the operation unit 26 of the distributor terminal 10. The control unit 12 of the distributor terminal 10 transmits a data generation instruction to the server 110 via the network interface 16, indicating that edited video data DX is to be generated. In response, the network interface 116 of the server 110 receives the data generation instruction and outputs the data generation instruction to the control unit 112. As a result, the control unit 112 of the server 110 acquires the data generation instruction (step S11).
[0105] Next, the control unit 112 of the server 110 reads out the video data D0 stored in the storage unit 114. As a result, the control unit 112 of the server 110 acquires the video data D0 distributed by the distributor to multiple viewers (step S12).
[0106] Next, the control unit 112 of the server 110 analyzes the voices of the performers in the video of the video data D0 (step S31). More specifically, the control unit 112 (information acquisition unit 122) of the server 110 generates a graph showing the relationship between time and the voice volume of the performers, as shown in FIG.
[0107] Next, the control unit 112 (information acquisition unit 122) of the server 110 identifies a period during which the voice volume of the performer is greater than a predetermined value p1 based on the graph of FIG. 13. Furthermore, the control unit 112 (information acquisition unit 122) of the server 110 identifies a performer action by analyzing the voice of the performer during the period during which the voice volume of the performer is greater than the predetermined value p1. Specifically, the control unit 112 (information acquisition unit 122) of the server 110 identifies whether the performer action is a scream, a loud laugh, or a loud shout. This allows the control unit 112 (information acquisition unit 122) of the server 110 to acquire action information indicating the performer action taken by the performer when the video data D0 is distributed (step S32). The performer action information includes the type of the performer action. In this way, the control unit 112 (information acquisition unit 122) of the server 110 generates performer action information based on the sounds included in the video of the video data D0. Furthermore, the control unit 112 (information acquisition unit 122) of the server 110 identifies the action occurrence time when the performer action occurred. As a result, the control unit 112 (information acquisition unit 122) of the server 110 acquires action occurrence time information indicating the action occurrence time (step S32). In this way, the performer action information is associated with action occurrence time information indicating the action occurrence time when the performer action occurred.
[0108] Next, the control unit 112 of the server 110 sets N to 1 (step S13), where N is a natural number. Then, the control unit 112 of the server 110 determines the start point SPN and end point EPN of the Nth portion AN of the video data D0 based on the performer action information (step S33). Step S33 will be described in more detail below with reference to FIG. 15.
[0109] First, the control unit 112 (start and end point determination unit 124) of the server 110 determines the start point SPN and end point EPN of the Nth part AN of the video data D0 based on the performer action information and the action occurrence time information (step S201). In this embodiment, the control unit 112 (start and end point determination unit 124) of the server 110 determines the start point of the action occurrence time indicated by the action occurrence time information as the start point SPN, and determines the end point of the action occurrence time indicated by the action occurrence time information as the end point EPN, as shown in Fig. 13 .
[0110] Next, the control unit 112 (start and end point determination unit 124) of the server 110 corrects the start SPN and end EPN of the Nth portion of the video data D0 (step S103). More specifically, the start SPN of the Nth portion may not coincide with the start of the conversation between the performers. Similarly, the end EPN of the Nth portion may not coincide with the end of the conversation between the performers. Therefore, in step S103, the control unit 112 (start and end point determination unit 124) of the server 110 corrects the start SPN of the Nth portion AN of the video data D0 so that the start SPN coincides with the start of the conversation between the performers. Similarly, the control unit 112 (start and end point determination unit 124) of the server 110 corrects the end EPN of the Nth portion AN of the video data D0 so that the start SPN coincides with the end of the conversation between the performers.
[0111] Specifically, the control unit 112 (start and end point determination unit 124) of the server 110 analyzes the voices of the performers in the video of the video data D0 to determine the corrected start point SPaN as the time when the performers start talking near the start point SPN. The time when the performers start talking is, for example, the end point of a predetermined time when the volume of the performers' voices is zero for that predetermined time. Furthermore, the control unit 112 (start and end point determination unit 124) of the server 110 analyzes the voices of the performers in the video of the video data D0 to determine the corrected end point EPaN as the time when the performers end talking near the end point EPN. The time when the performers end talking is, for example, the start point of a predetermined time when the volume of the performers' voices is zero for that predetermined time. With the above processing, step S33 ends, and the process proceeds to step S15.
[0112] Next, the control unit 112 (extraction unit 126) of the server 110 extracts the Nth edited video data DXN corresponding to the Nth portion AN between the corrected start point SPaN and the corrected end point EPaN from the video data D0 (step S15). The control unit 112 (extraction unit 126) of the server 110 determines whether or not there is any further performer action at a time later than the time when the performer action information processed in step S33 occurred (step S34). If there is performer action information, the process proceeds to step S17. If there is no performer action information, the process proceeds to step S18.
[0113] If performer action information exists, control unit 112 of server 110 increments N by 1 (step S17). After this, the process returns to step S33. Then, by repeating steps S33, S15, S34, and S17, control unit 112 of server 110 extracts first edited video data DX1 through Nth edited video data DXN, which correspond to first portion A1 through Nth portion AN, from video data D0.
[0114] If no performer action information exists, the control unit 112 of the server 110 generates edited video data DX by combining the first edited video data DX1 through the Nth edited video data DXN (step S18). This ends the process.
[0115] The edited video data DX as described above is stored in the storage unit 114 of the server 110. The distributor terminal 10 and the viewer terminals 210-1 to 210-n can then obtain the edited video data DX from the server 110 by accessing a web page managed by the server 110. This allows the distributor and multiple viewers to view the edited video data DX.
[0116] [effect] According to the system 1a, characteristic portions of the video data D0 can be identified. More specifically, the control unit 112 of the server 110 acquires performer action information indicating performer actions taken by performers when the video data D0 is distributed. Then, the control unit 112 of the server 110 determines the start point SPN and end point EPN of the Nth portion of the video data D0 based on the performer action information. Such an Nth portion AN is the portion of the video data D0 where a performer action occurred. Therefore, the Nth portion AN is a characteristic portion of the video data D0. For the above reasons, according to the system 1a, characteristic portions of the video data D0 can be identified.
[0117] According to the system 1a, characteristic portions of the video data D0 can be identified with high accuracy. The control unit 112 of the server 110 determines the start point SPN and end point EPN of the Nth portion AN of the video data D0 based on the performer action information and the action occurrence time information. This allows the control unit 112 of the server 110 to accurately identify the Nth portion AN corresponding to the period in which the performer action occurred, thereby accurately identifying characteristic portions of the video data D0.
[0118] (Second Modification) The following describes a system 1b according to a second modified example with reference to the drawings. System 1b differs from system 1a in that performer actions are movements of performers included in the video of video data D0. The following description will focus on this difference.
[0119] [System 1b Operation Overview] In system 1b, the following two operations are performed. (1) Watching videos (2) Generation of edited video data
[0120] (1) Watching videos Viewing of moving images in the system 1b is the same as viewing of moving images in the system 1a, so a description thereof will be omitted.
[0121] (2) Generation of edited video data In generating edited video data, the server 110 extracts a portion of the video data D0 to generate edited video data DX, and transmits the edited video data DX to the distributor terminal 10. More specifically, the performers perform performer actions in the video. The performer actions are, for example, a performer's surprised action, a performer's happy action, and a performer's angry action, as shown in FIG.
[0122] Therefore, the server 110 analyzes the video of the performers in the video data D0 to generate performer information related to the performers' facial expressions and body movements. The server 110 then determines whether a performer action has occurred based on the performer information. Specifically, the server 110 determines that a performer action has occurred if the performer is surprised, happy, or angry. The server 110 then identifies the portions of the video data D0 where the performer action occurred as a first portion A1 and a second portion A2. The first portion A1 and the second portion A2 are exciting portions of the video where the performers were excited. The server 110 extracts the first portion A1 to generate first edited video data DX1 and extracts the second portion A2 to generate second edited video data DX2. The server 110 generates edited video data DX by combining the first edited video data DX1 and the second edited video data DX2.
[0123] [System 1b operation] Next, the operation of the system 1b will be described with reference to the drawing. Fig. 16 is a flowchart showing the operation of the control unit 112 of the server 110.
[0124] Steps S11 and S12 of the system 1b are the same as steps S11 and S12 of the system 1a, and therefore a description thereof will be omitted.
[0125] Next, the control unit 112 of the server 110 generates performer information related to the facial expressions and body movements of the performers by analyzing the images of the performers in the video data D0 (step S41). The performer information is information in which the movements of the performers are quantified, for example, by motion capture, facial emotion recognition AI, or the like.
[0126] Next, the control unit 112 (information acquisition unit 122) of the server 110 acquires performer action information indicating performer actions based on the performer information (step S42). Specifically, the control unit 112 (information acquisition unit 122) of the server 110 identifies performer actions that indicate surprise, joy, and anger in the video of the video data D0. For this determination, motion capture, facial emotion recognition AI, etc. are used. In this manner, the control unit 112 (information acquisition unit 122) of the server 110 generates performer action information based on the actions of the performers included in the video of the video data D0. Furthermore, the control unit 112 (information acquisition unit 122) of the server 110 identifies the action occurrence time at which the performer action occurred. As a result, the control unit 112 (information acquisition unit 122) of the server 110 acquires action occurrence time information indicating the action occurrence time (step S42). In this manner, the performer action information is associated with action occurrence time information indicating the action occurrence time at which the performer action occurred.
[0127] Steps S13, S33, S15, S34, S17, and S18 of the system 1b are the same as steps S13, S33, S15, S34, S17, and S18 of the system 1a, and therefore will not be described here.
[0128] The system 1b as described above can achieve the same effects as the system 1a.
[0129] (Third Modification) Below, a system 1c according to the third modified example will be described with reference to the drawings. Fig. 17 shows images displayed on viewer terminals 210-1 to 210-n. System 1c differs from system 1 in that server 110 extracts the opening part, topic part, and closing part of video data D0 and combines the opening part, topic part, and closing part. The following description will focus on these differences.
[0130] [System 1c Operation Overview] In system 1c, the following two operations are performed. (1) Watching videos (2) Generation of edited video data
[0131] (1) Watching videos Viewing of moving images in the system 1c is the same as viewing of moving images in the system 1a, so a description thereof will be omitted.
[0132] (2) Generation of edited video data In generating edited video data, server 110 extracts a portion of video data D0 to generate edited video data DX and transmits the edited video data DX to distributor terminal 10. More specifically, performers perform performer actions in the video. Performer actions include, for example, the performer giving an opening greeting as shown in the opening video of FIG. 17, the performer talking about a topic as shown in the topic video of FIG. 17, and the performer giving a closing greeting as shown in the closing video of FIG. 17.
[0133] Therefore, the server 110 analyzes the voices of the performers in the video data D0 to generate performer information that is a text version of the performers' conversations. The server 110 then determines whether a performer action has occurred based on the performer information. Specifically, the server 110 determines that a performer action has occurred when a performer gives an opening greeting, when a performer talks about a topic, or when a performer gives a closing greeting. Furthermore, the server 110 identifies the portions of the video data D0 where a performer action has occurred as an opening portion A11, a topic portion A12, and a closing portion A13. The opening portion A11 (first portion) and the closing portion A13 (first portion) are standard portions in a video that contain standard content. The topic portion (first portion) is a portion in a video that contains content related to a specific topic.
[0134] Server 110 extracts opening portion A11 from video data D0 to generate opening-edited video data DX11, extracts topic portion A12 from video data D0 to generate topic-edited video data DX12, and extracts closing portion A13 from video data D0 to generate closing-edited video data DX 13. Furthermore, server 110 generates edited video data DX by combining opening-edited video data DX11, topic-edited video data DX12, and closing-edited video data DX13.
[0135] [System 1c behavior] Next, the operation of the system 1c will be described with reference to the drawing. Fig. 18 is a flowchart showing the operation of the control unit 112 of the server 110.
[0136] This process is started when the distributor operates the operation unit 26 of the distributor terminal 10. The control unit 12 of the distributor terminal 10 sends a data generation instruction to the server 110 via the network interface 16, indicating that edited video data DX is to be generated. The data generation instruction includes an instruction to extract an opening portion A11, a topic portion A12, and a closing portion A13 from the video data D0. In response, the network interface 116 of the server 110 receives the data generation instruction and outputs the data generation instruction to the control unit 112. As a result, the control unit 112 of the server 110 acquires the data generation instruction (step S11).
[0137] Next, the control unit 112 of the server 110 reads out the video data D0 stored in the storage unit 114. As a result, the control unit 112 of the server 110 acquires the video data D0 distributed by the performers to a plurality of viewers (step S12).
[0138] Next, the control unit 112 of the server 110 analyzes the voices of the performers in the video of the video data D0 (step S51). More specifically, the control unit 112 (information acquisition unit 122) of the server 110 analyzes the voices of the performers in the video of the video data D0, thereby generating performer information in which the conversations of the performers are converted into text.
[0139] Next, the control unit 112 (information acquisition unit 122) of the server 110 acquires performer action information indicating performer actions based on the performer information (step S52). Specifically, the control unit 112 (information acquisition unit 122) of the server 110 performs the following processes (a) to (c).
[0140] (a) Opening The control unit 112 (information acquisition unit 122) of the server 110 extracts keywords related to greetings from the performer information to generate performer action information indicating that a performer has given an opening greeting. Keywords related to opening greetings are, for example, "hello" or "good morning." Furthermore, the control unit 112 (information acquisition unit 122) of the server 110 identifies the action occurrence time at which a performer action indicating that a performer has given an opening greeting occurred.
[0141] (b) Topic The control unit 112 (information acquisition unit 122) of the server 110 extracts keywords related to the topic from the performer information to generate performer action information indicating that the performer talked about the topic. Keywords related to the topic are, for example, words determined by the performer, such as the title of an anime, the title of a piece of music, or a specific character. If the keyword is the title of an anime, the topic portion A12 is the portion where the performer talks about the anime. If the keyword is the title of a piece of music, the topic portion A12 is the portion where the performer talks about the music. If the keyword is a specific character, the topic portion A12 is the portion where the performer talks about the specific character. The keyword may also be a scream made by a performer watching a horror movie, a scream made by a performer playing a horror game, or a word like "cute." If the keyword is a scream, the topic portion is the portion where the broadcaster is excited by the fear in a horror movie or a horror game. If the keyword is "cute," the topic portion is the portion where the broadcaster is happy to see a cute animal, etc. In this case, the keyword is included in the data generation instruction. The keyword may also be, for example, a word included in the title of the video data D0. Furthermore, the control unit 112 (information acquisition unit 122) of the server 110 identifies the action occurrence time when a performer action indicating that the performer talked about the topic occurred.
[0142] (c) Closing The control unit 112 (information acquisition unit 122) of the server 110 extracts keywords related to closing greetings from the performer information, thereby generating performer action information indicating that the performer has given a closing greeting. Keywords related to closing greetings are, for example, "to wrap up" or "goodbye." Furthermore, the control unit 112 (information acquisition unit 122) of the server 110 identifies the action occurrence time at which the performer action indicating that the performer has given a closing greeting occurred. As shown in (a) to (c), the control unit 112 (information acquisition unit 122) of the server 110 generates performer action information by extracting keywords from sounds included in the video of the video data D0.
[0143] Next, the control unit 112 of the server 110 determines a start point SP11 and an end point EP11 of the opening portion A11 of the video data D0 based on the performer action information (step S53). Specifically, the control unit 112 (start and end point determination unit 124) of the server 110 determines a start point SP11 and an end point EP11 of the opening portion A11 of the video data D0 based on the performer action information indicating that the performer has given an opening greeting and the action occurrence time information. In this embodiment, the control unit 112 (start and end point determination unit 124) of the server 110 determines the start point of the action occurrence time indicated by the action occurrence time information as the start point SP11, and determines the end point of the action occurrence time indicated by the action occurrence time information as the end point EP11.
[0144] Next, the control unit 112 (start and end point determination unit 124) of the server 110 corrects the start point SP11 and end point EP11 of the opening portion A11 of the video data D0 (step S54). More specifically, the start point SP11 of the opening portion A11 may not coincide with the start of the performers' conversation. Similarly, the end point EP11 of the opening portion A11 may not coincide with the end of the performers' conversation. Therefore, in step S54, the control unit 112 (start and end point determination unit 124) of the server 110 corrects the start point SP11 of the opening portion A11 of the video data D0 so that the start point SP11 coincides with the start of the performers' conversation. Similarly, the control unit 112 (start and end point determination unit 124) of the server 110 corrects the end point EP11 of the opening portion A11 of the video data D0 so that the start point SP11 coincides with the end of the performers' conversation. However, the details of step S54 are the same as those of step S103 in FIG. 11, and therefore will not be described here.
[0145] Next, control unit 112 (extraction unit 126) of server 110 extracts opening edited video data DX11 corresponding to opening portion A11 between corrected start point SPa11 and corrected end point EPa11 from video data D0 (step S55).
[0146] Next, the control unit 112 of the server 110 determines the start point SP12 and end point EP12 of the topic portion A12 of the video data D0 based on the performer action information (step S56). Note that the details of step S56 are the same as those of step S53, and therefore will not be described here.
[0147] Next, the control unit 112 (start and end point determining unit 124) of the server 110 corrects the start point SP11 and end point EP11 of the opening portion A11 of the video data D0 (step S57). Note that the details of step S57 are the same as those of step S54, and therefore will not be described again.
[0148] Next, control unit 112 (extraction unit 126) of server 110 extracts topic-edited video data DX12 corresponding to topic portion A12 between corrected start point SPa12 and corrected end point EPa12 from video data D0 (step S58).
[0149] Next, the control unit 112 of the server 110 determines the start point SP13 and end point EP13 of the closing portion A13 of the video data D0 based on the performer action information (step S59). Note that the details of step S59 are the same as those of step S53, and therefore will not be described here.
[0150] Next, the control unit 112 (start and end point determination unit 124) of the server 110 corrects the start point SP13 and end point EP13 of the closing portion A13 of the video data D0 (step S60). Note that the details of step S60 are the same as those of step S54, and therefore will not be described again.
[0151] Next, the control unit 112 (extraction unit 126) of the server 110 extracts, from the video data D0, closing-edited video data DX13 corresponding to the closing portion A13 between the corrected start point SPa13 and the corrected end point EPa13 (step S61).
[0152] Next, the control unit 112 of the server 110 generates edited video data DX by combining the opening edited video data DX11, the topic edited video data DX12, and the closing edited video data DX13 (step S62). This ends this process. As a result, the distributor terminal 10 and the viewer terminals 210-1 to 210-n can obtain the edited video data DX from the server 110. As a result, the distributor and multiple viewers can view the edited video data DX.
[0153] [effect] In the system 1c, the control unit 112 of the server 110 generates performer action information by extracting keywords from sounds included in the video of the video data D0. This allows the control unit 112 of the server 110 to extract portions related to the keywords from the video data D0. Therefore, by selecting keywords preferred by the distributor, edited video data DX tailored to the distributor's preferences is generated.
[0154] (Fourth Modification) A system 1d according to the fourth modification will be described below with reference to the drawings.
[0155] System 1d differs from system 1 in that the control unit 112 of the server 110 generates edited video data DX of a length specified by the distributor. In this embodiment, the distributor can select whether to have the server 110 generate edited video data DX having a length of 15 seconds or edited video data DX having a length of 30 seconds.
[0156] [System 1d operation overview] In the system 1d, the following two operations are performed. (1) Watching videos (2) Generation of edited video data
[0157] (1) Watching videos Viewing of videos in system 1d is the same as viewing of videos in system 1, so a description thereof will be omitted.
[0158] (2) Generation of edited video data In generating edited video data, the server 110 extracts a portion of the video data D0 to generate edited video data DX and transmits the edited video data DX to the distributor terminal 10. Here, when generating edited video data DX having a length of 15 seconds, as shown in FIG. 19, the server 110 extracts a first portion A101 having an excitement parameter P greater than a predetermined value px. The server 110 determines whether the length of the first portion A101 is 15 seconds. The server 110 adjusts the value of the predetermined value px so that the length of the first portion A101 is 15 seconds. Thereafter, the server 110 generates edited video data DX having a length of 15 seconds.
[0159] On the other hand, when generating edited video data DX having a length of 30 seconds, as shown in FIG. 19, the server 110 extracts the first portion A111 and the second portion A112 having an excitement parameter P greater than a predetermined value py. The predetermined value py is smaller than a predetermined value px. The server 110 determines whether the total length of the first portion A111 and the second portion A112 is 30 seconds. The server 110 adjusts the predetermined value py so that the total length of the first portion A111 and the second portion A112 is 30 seconds. After this, the server 110 generates edited video data DX having a length of 30 seconds.
[0160] [System 1d Operation] Next, the operation of the system 1d will be described with reference to the drawings. Figures 20 and 21 are flowcharts showing the operation of the control unit 112 of the server 110.
[0161] This process is started when the distributor operates the operation unit 26 of the distributor terminal 10. At this time, the distributor selects whether to have the server 110 generate edited video data DX having a length of 15 seconds or edited video data DX having a length of 30 seconds. The control unit 12 of the distributor terminal 10 transmits a data generation instruction to the server 110 via the network interface 16, indicating that edited video data DX is to be generated. The data generation instruction includes information regarding the length of the edited video data DX. In response, the network interface 116 of the server 110 receives the data generation instruction and outputs the data generation instruction to the control unit 112. As a result, the control unit 112 of the server 110 acquires the data generation instruction (step S11).
[0162] Next, the control unit 112 of the server 110 reads out the video data D0 stored in the storage unit 114. As a result, the control unit 112 of the server 110 acquires the video data D0 distributed by the distributor to multiple viewers (step S12).
[0163] Next, the control unit 112 of the server 110 determines whether to generate 15-second edited video data DX or 30-second edited video data DX based on the data generation instruction (step S301). If 15-second edited video data DX is to be generated, the process proceeds to step S302. If 30-second edited video data DX is to be generated, the process proceeds to step S307.
[0164] When generating 15-second edited video data DX, the control unit 112 of the server 110 sets the predetermined value px to a predetermined value (step S302). Next, the control unit 112 of the server 110 sets N to 1 (step S13), where N is a natural number. Then, the control unit 112 of the server 110 determines the start point SPN and end point EPN of the Nth portion AN of the video data D0 based on the viewer action information (step S14). Note that step S14 of the system 1d differs from step S14 of the system 1 only in that the predetermined value px is used instead of the predetermined value p0. Furthermore, steps S15 to S17 of the system 1d are the same as steps S15 to S17 of the system 1. Therefore, a description of steps S14 to S17 will be omitted.
[0165] If there is no portion having an excitement parameter P greater than the predetermined value px in step S16, control unit 112 of server 110 determines whether the total length of the first edited video data DX1 through the Nth edited video data DXN extracted in steps S13 to S17 is longer than 15 seconds (step S303). If the total length of the first edited video data DX1 through the Nth edited video data DXN is longer than 15 seconds, the process proceeds to step S304. If the total length of the first edited video data DX1 through the Nth edited video data DXN is not longer than 15 seconds, the process proceeds to step S305.
[0166] If the total length of the first edited video data DX1 through the Nth edited video data DXN is longer than 15 seconds, the control unit 112 of the server 110 increases the predetermined value px (step S304), after which the process returns to step S13.
[0167] If the total length of the first edited video data DX1 through the Nth edited video data DXN is not longer than 15 seconds, it is determined whether the total length of the first edited video data DX1 through the Nth edited video data DXN extracted in steps S13 to S17 is shorter than 15 seconds (step S305). If the total length of the first edited video data DX1 through the Nth edited video data DXN is shorter than 15 seconds, the process proceeds to step S306. If the total length of the first edited video data DX1 through the Nth edited video data DXN is not shorter than 15 seconds, the process proceeds to step S18.
[0168] If the total length of the first edited video data DX1 through the Nth edited video data DXN is shorter than 15 seconds, the control unit 112 of the server 110 decreases the predetermined value px (step S306). After this, the process returns to step S13.
[0169] If the total length of the first edited video data DX1 through the Nth edited video data DXN is not less than 15 seconds, the control unit 112 of the server 110 determines that the total length of the first edited video data DX1 through the Nth edited video data DXN is 15 seconds. The control unit 112 of the server 110 then combines the first edited video data DX1 through the Nth edited video data DXN to generate edited video data DX (step S18). This process then ends.
[0170] When generating 30-second edited video data DX (Yes in step S301), the control unit 112 of the server 110 sets the predetermined value py to a predetermined value (step S307). Next, the control unit 112 of the server 110 sets N to 1 (step S13), where N is a natural number. Then, the control unit 112 of the server 110 determines the start point SPN and end point EPN of the Nth portion AN of the video data D0 based on the viewer action information (step S14). Note that step S14 of the system 1d differs from step S14 of the system 1 only in that a predetermined value py is used instead of the predetermined value p0. Furthermore, steps S15 to S17 of the system 1d are the same as steps S15 to S17 of the system 1. Therefore, a description of steps S14 to S17 will be omitted.
[0171] If there is no portion having an excitement parameter P greater than the predetermined value py in step S16, control unit 112 of server 110 determines whether the total length of the first edited video data DX1 through the Nth edited video data DXN extracted in steps S13 to S17 is longer than 30 seconds (step S308). If the total length of the first edited video data DX1 through the Nth edited video data DXN is longer than 30 seconds, the process proceeds to step S309. If the total length of the first edited video data DX1 through the Nth edited video data DXN is not longer than 30 seconds, the process proceeds to step S310.
[0172] If the total length of the first edited video data DX1 through the Nth edited video data DXN is longer than 30 seconds, the control unit 112 of the server 110 increases the predetermined value py (step S309), after which the process returns to step S13.
[0173] If the total length of the first edited video data DX1 through the Nth edited video data DXN is not longer than 30 seconds, it is determined whether the total length of the first edited video data DX1 through the Nth edited video data DXN extracted in steps S13 to S17 is shorter than 30 seconds (step S310). If the total length of the first edited video data DX1 through the Nth edited video data DXN is shorter than 30 seconds, the process proceeds to step S311. If the total length of the first edited video data DX1 through the Nth edited video data DXN is not shorter than 30 seconds, the process proceeds to step S18.
[0174] If the total length of the first edited video data DX1 through the Nth edited video data DXN is shorter than 30 seconds, the control unit 112 of the server 110 decreases the predetermined value py (step S311). After this, the process returns to step S13.
[0175] If the total length of the first edited video data DX1 through the Nth edited video data DXN is not less than 30 seconds, the control unit 112 of the server 110 determines that the total length of the first edited video data DX1 through the Nth edited video data DXN is 30 seconds. The control unit 112 of the server 110 then combines the first edited video data DX1 through the Nth edited video data DXN to generate edited video data DX (step S18). This process then ends.
[0176] [effect] The system 1d can achieve the same effects as the system 1.
[0177] In system 1d, control unit 112 of server 110 adjusts the predetermined values px and py to adjust the total length of first edited video data DX1 through Nth edited video data DXN, thereby enabling the distributor to obtain edited video data DX of the desired length.
[0178] (Other embodiments) The various control means and processing procedures described in the above embodiments are merely examples and are not intended to limit the scope of the present invention, its applications, or its uses. The various control means and processing procedures can be appropriately modified in design within the scope that does not change the gist of the present invention.
[0179] In the systems 1, 1a to 1d, the video of the video data D0 may display the performers themselves instead of their avatars.
[0180] In the systems 1, 1a to 1d, the control unit 112 of the server 110 may acquire action information, which is viewer action information and distributor action information. Then, the control unit 112 of the server 110 may determine the start point SPN and end point EPN of the Nth portion AN of the video data D0 based on the action information. Therefore, the control unit 112 of the server 110 may extract the exciting parts (first portion A1 and second portion A2 in FIG. 5), the topic parts (topic parts in FIG. 17), and the standard parts (opening portion A11 and closing portion A13 in FIG. 17), and combine the exciting parts, topic parts, and standard parts to generate edited video data DX.
[0181] 9 to 11 of the system 1 may be performed in the distributor terminal 10. For example, the server 110 generates a counting result (step S101). Then, the distributor terminal 10 determines the start SPN and end EPN of the Nth part AN (step S102). The server 110 may extract the Nth edited video data DXN based on the start SPN and end EPN, and may combine the first edited video data DX1 through the Nth edited video data DXN (steps S15 and S18). In this way, the process of acquiring the video data D0, the process of acquiring performer action information and / or viewer action information, and the process of determining the start SPN and end EPN of the Nth part AN may be performed by one or more computers. The one or more computers are the distributor terminal 10 and the server 110.
[0182] The one or more computers may be one or more servers, that is, the systems 1a to 1d may include a plurality of servers.
[0183] In the systems 1a to 1d, part of the processing in the flowchart may be performed in the distributor terminal 10.
[0184] Note that the control unit 112 of the server 110 generates performer action information based on sounds included in the video of the video data D0 or on the actions of the performers included in the video of the video data D0. However, the control unit 112 of the server 110 may also generate performer action information based on sounds included in the video of the video data D0 and on the actions of the performers included in the video of the video data D0.
[0185] The control unit 112 of the server 110 generates time distribution information indicating the relationship between the excitement parameter P and time based on the action information and the action occurrence time information, and determines the start SPN and end EPN of the Nth portion AN of the video data D0 based on the time distribution information. However, the control unit 112 of the server 110 may also generate time distribution information indicating the relationship between the number of performer actions and time based on the action information and the action occurrence time information, and determine the start SPN and end EPN of the Nth portion AN of the video data D0 based on the time distribution information. In this case, the action information is associated with action occurrence time information indicating the action occurrence times at which performer actions and viewer actions occurred. For example, in system 1c, the control unit 112 of the server 110 may determine the portion where a large number of performer actions in which performers utter keywords occur as the topic portion A12. In this case, the control unit 112 of the server 110 may perform a weighting process on the performer actions.
[0186] In the systems 1, 1a to 1d, the control unit 112 of the server 110 does not have to extract the Nth portion AN and combine the first edited video data DX1 through the Nth edited video data DXN. That is, in the systems 1, 1a to 1d, the control unit 112 of the server 110 may end the process by determining the start point SPN and end point EPN of the Nth portion AN. In this case, the control unit 112 of the server 110 can determine the start points and end points of the first portion A1 through the Nth portion AN of the video data D0. That is, the control unit 112 of the server 110 can determine the start points and end points of multiple chapters included in the video data D0.
[0187] In the systems 1, 1a to 1d, the control unit 112 of the server 110 may determine only the start point SP1 and end point EP1 of the first part A1 and extract the first edited video data DX1 of the first part A1. Therefore, the control unit 112 of the server 110 does not need to extract the second edited video data DX2 through the Nth edited video data DXN.
[0188] In the systems 1, 1a to 1d, the video data D0 is distributed live, but pre-recorded video data D0 may also be distributed.
[0189] In systems 1, 1a to 1d, control unit 112 of server 110 may generate thumbnail image data based on action information. The thumbnail image is an image showing the contents of video data D0 or edited video data DX. The thumbnail image is included in a web page transmitted from server 110 to viewer terminals 210-1 to 210-n when viewer terminals 210-1 to 210-n download video data D0 or edited video data DX. Specifically, control unit 112 of server 110 may generate thumbnail images of video data D0 as shown in FIG. 2 based on viewer action information and / or performer action information. The thumbnail image may be, for example, an image of the time when excitement parameter P in video data D0 was largest, or an image of the time when a performer uttered a keyword the most frequently in video data D0.
[0190] When generating thumbnail images, the control unit 112 of the server 110 may combine the text of a keyword issued by the distributor or the text of the title of the video with the thumbnail images.
[0191] It should be noted that the control unit 112 of the server 110 may generate thumbnail video data by combining videos of the video data D0 at multiple times.
[0192] It is not necessary for the control unit 112 of the server 110 to perform weighting processing. In this case, the excitement parameter P is the frequency of occurrence of a viewer action.
[0193] The control unit 112 of the server 110 associates the viewer action information with the action occurrence time information. However, the control unit 12 of the distributor terminal 10 may associate the viewer action information with the action occurrence time information and transmit the viewer action information and the action occurrence time information to the server 110.
[0194] Server 110 combines the viewer action display with the video of video data D0. However, viewer terminals 210-1 to 210-n may combine the viewer action display with the video of video data D0. In this case, server 110 distributes the viewer actions and the video data of the video to viewer terminals 210-1 to 210-N.
[0195] In step S32 of the system 1a, the control unit 112 (information acquisition unit 122) of the server 110 does not need to specify whether the performer action is a performer screaming, laughing loudly, or shouting loudly. The performer action may indicate that the performer's voice volume has exceeded a predetermined value p2.
[0196] Note that the control unit 112 of the server 110 may obtain the video data D0 from a computer other than the server 110, rather than from the storage unit 114 of the server 110. The computer other than the server 110 is, for example, the distributor terminal 10.
[0197] The distributor terminal 10 may generate the distributor action information by analyzing the voices of the performers in the video data D0 or the movements of the performers in the video data D0. In this case, the distributor terminal 10 transmits the distributor action information to the server 110. As a result, the control unit 112 of the server 110 may acquire the distributor action information.
[0198] 18, the control unit 112 of the server 110 may determine the start and end points of the opening and closing portions based on time. For example, the control unit 112 of the server 110 may identify one minute from the start of the video data D0 as the opening portion. Furthermore, the control unit 112 of the server 110 may identify one minute until the end of the video data D0 as the closing portion.
[0199] Note that a viewer performs a viewer action after considering the viewer action. Therefore, the timing at which the viewer action occurs is slightly delayed from the timing at which the viewer considers performing the viewer action. Therefore, the control unit 112 of the server 110 may determine the start point SPN of the Nth portion AN to be a time slightly before the moment at which the excitement parameter P becomes greater than a predetermined value p0.
[0200] Note that the video data D0 may be data of a video that does not include performers. Examples of videos that do not include performers include animation, natural scenery, etc. In this case, the control unit 112 of the server 110 determines the start SPN and end EPN of the Nth part AN based on viewer actions. Also, the video that does not include performers may be, for example, a video that does not include video of the performers but includes audio of the performers. In this case, the control unit 112 of the server 110 may determine the start SPN and end EPN of the Nth part AN based on viewer actions, or may determine the start SPN and end EPN of the Nth part AN based on distributor actions.
[0201] The performer and the distributor may be the same person.
[0202] Although there are first to third viewer actions, there may be three or more types of viewer actions. In other words, there may be first to M-th viewer actions, where M is a natural number.
[0203] The number of viewers may be one or more.
[0204] If there are multiple types of gifts, the weighting table shown in FIG. 4 may record multiple types of gifts and points corresponding to the multiple types of gifts. The points corresponding to the multiple types of gifts have values according to the type of gift. In other words, the points corresponding to the multiple types of gifts represent the degree of excitement of viewer actions according to the type of gift. For example, points corresponding to expensive gifts have large values, and points corresponding to inexpensive gifts have small values.
[0205] In system 1c, server 110 extracts the opening, topic, and closing parts of video data D0 and combines them. However, server 110 may extract and combine multiple parts other than the opening, topic, and closing parts. The multiple parts may be, for example, multiple exciting parts in system 1, and may be parts whose start and end points are determined based on action information.
[0206] In system 1c, the control unit 112 (information acquisition unit 122) of the server 110 may extract a topic portion based on viewer action information instead of performer action information. In this case, the control unit 112 (information acquisition unit 122) of the server 110 may identify a topic portion included in a video by determining whether a viewer comment included in the viewer action information is a keyword related to the topic. Keywords include, for example, words such as "congratulations" used by viewers to congratulate a performer, or words such as "cute" used by viewers to praise a performer. If the keyword is "congratulations," the topic portion is the portion where the performer is congratulated. If the keyword is "cute," the topic portion is the portion where the performer is praised.
[0207] The effects of this embodiment can be achieved even when these other embodiments are adopted. Furthermore, this embodiment can be combined with other embodiments, and other embodiments can be combined with each other as appropriate. [Explanation of symbols]
[0208] 1, 1a to 1d: System 10: Streamer terminal 12: Control unit 14: Storage part 26:Operation section 110: Server 112: Control unit 114: Storage section 116: Network interface 120: Data acquisition unit 122: Information acquisition department 124: Start and end point determination unit 126: Extraction part 128:Joining part 210-1~210-N: Viewer terminals 212: Control unit A1,A101,A111: 1st part A2,A112:Second part A101: Part 1 A11: Opening part A12: Topic section A13: Closing part AN: Part N D0: Video data DX: Edited video data DX1: First edited video data DX11: Opening edited video data DX12: Topic edited video data DX13: Closing edited video data DX2: Second edited video data DX3: Third edited video data DXN: Edited video data N P: Excitement parameter SP1,SP11,SP12,SP13,SPN,SPa11,SPa12,SPa13,SPaN:Starting point EP1,EP11,EP12,EP13,EPN,EPa11,EPa12,EPa13,EPaN:End point p0,p1,px,py: predetermined value
Claims
1. A system comprising one or more computers, The one or more computers Obtaining video data distributed by a distributor to one or more viewers; acquire action information indicating performer actions by performers when the video data is distributed and / or viewer actions by the one or more viewers with respect to the video data when the video data is distributed; determining a start point and an end point of a first portion of the video data based on the action information; system.
2. the one or more computers are one or more servers; The system of claim 1 .
3. The first portion is an exciting portion in the video in which the performers or the one or more viewers are excited, a topic portion in the video containing content related to a specific topic, or a standard portion in the video containing standard content.
3. A system according to claim 1 or claim 2.
4. the action information is viewer action information indicating the viewer actions taken by the one or more viewers, the viewer action is an action of each of the one or more viewers operating one or more viewer terminals, a video of the video data including a viewer action indication indicating the viewer actions by the one or more viewers; the one or more computers determine a start point and an end point of the first portion of the video data based on the viewer action information; 3. The system according to claim 1 or claim 2.
5. The viewer action information is information showing the viewer's goodwill toward the performer, information showing a comment from the viewer to the performer, or information showing a gift given by the viewer to the performer. The system of claim 4.
6. In the process of determining the start point and the end point of the first portion of the video data, the one or more computers generate a count result of counting the frequency of occurrence of the viewer action based on the viewer action information, and determine the start point and the end point of the first portion of the video data based on the count result. The system of claim 4.
7. the action information is performer action information indicating the performer action by the performer, In the process of acquiring the action information, the one or more computers generate the performer action information based on sounds included in the video of the video data and / or actions of the performers included in the video of the video data; the one or more computers determine a start point and an end point of the first portion of the video data based on the performer action information; 3. The system according to claim 1 or claim 2.
8. In the process of acquiring the action information, the one or more computers generate the performer action information by extracting keywords from sounds included in the video of the video data. The system of claim 7.
9. The action information is associated with action occurrence time information indicating the time at which the performer action and / or the viewer action occurred, In the process of determining the start point and the end point of the first portion of the video data, the one or more computers determine the start point and the end point of the first portion of the video data based on the action information and the action occurrence time information.
3. The system according to claim 1 or claim 2.
10. In the process of determining the start point and the end point of the first portion of the video data, the one or more computers generate time distribution information indicating a relationship between time and a parameter related to the number of occurrences of the performer actions or the number of occurrences of the viewer actions, based on the action information and the action occurrence time information, and determine the start point and the end point of the first portion of the video data based on the time distribution information. The system of claim 9.
11. In the process of determining the start point and the end point of the first portion of the video data, the one or more computers determine the start point and the end point of a period in which the parameter of the time distribution information is greater than a predetermined value as the start point and the end point of the first portion of the video data, respectively. The system of claim 10.
12. M is a natural number, The viewer actions include a first viewer action to an M-th viewer action, In the process of determining the start point and the end point of the first portion of the video data, the one or more computers calculate the parameters by weighting the first viewer action through the Mth viewer action, respectively, according to the contents of the first viewer action through the Mth viewer action. The system of claim 10.
13. the one or more computers extract first edited video data corresponding to the first portion from the video data; 3. The system according to claim 1 or claim 2.
14. N is a natural number, In the process of determining the start point and the end point of the N-th portion of the video data, the one or more computers determine the start point and the end point of the N-th portion of the video data based on the action information; The one or more computers extracting N-th edited video data corresponding to the N-th portion from the video data; generating edited video data by combining the first edited video data through the Nth edited video data; The system of claim 13.
15. 1. A data production method, comprising: The one or more computers Acquire video data distributed by a distributor to viewers, acquire action information indicating performer actions by performers when the video data is distributed and / or viewer actions by the one or more viewers with respect to the video data when the video data is distributed; extracting a portion of the video data based on the action information, thereby generating edited video data having a structure that allows a computer to reproduce the portion of the video data; Data production method.
Citation Information
Patent Citations
Identification of the intro part of a video content
EP3772856A1
Retrieving device
JP2008276340A
Device, method, system and program for content distribution
JP2021141564A
Dynamic chapter generation for a communication session
US20230326454A1
Program, information processing method, terminal and server
JP7128338B1