Content processing device, content processing program, and content processing method
Patent Information
- Application Number
- JP2025508531
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-12-11
- Estimated Expiration
- 2044-06-04
AI Technical Summary
Conventional video processing technologies require multiple cameras and acoustic analysis to handle content, limiting their ability to display a video showing only a tracking target.
A content processing device that acquires video data, detects and generates position and size information of objects, cuts out partial areas for tracking, and outputs a tracking video based on viewer selection, allowing display of a single tracking target using a simple mechanism.
Enables the display of a video showing only the tracking target, even when multiple objects are present, using a straightforward approach.
Abstract
Description
[Technical Field]
[0001] The present invention relates to a content processing device, a content processing program, and a content processing method. [Background technology]
[0002] Various video processing technologies have been proposed for providing live video customized for each viewer. For example, Patent Document 1 describes a video data processing device that includes an input means for inputting video captured by multiple cameras placed in a venue, an analysis data acquisition means for acquiring music analysis data that is the result of an acoustic analysis of the music played in the venue, a switching means for switching the video to be output based on the music analysis data, and an output means for outputting the video switched by the switching means. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2021-153257 Summary of the Invention [Problem to be solved by the invention]
[0004] However, conventional technologies such as those described in Patent Document 1 require multiple cameras to capture images of the target. Also, because switching is performed using acoustic analysis results, they cannot handle content that is only video.
[0005] One aspect of the present invention has been made in view of the above-mentioned problems, and its purpose is to make it possible to display a video showing only a tracking target using a simple mechanism. [Means for solving the problem]
[0006] In order to solve the above problems, a content processing device according to one aspect of the present invention comprises a video acquisition unit that acquires content data including a video; an information generation unit that detects a plurality of objects to be viewed from the video and generates position information indicating the position of each characteristic part of the plurality of objects to be viewed at multiple points in time in the video and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation unit that, based on the position information and the size information, cuts out the partial areas corresponding to the plurality of objects to be viewed from each frame corresponding to multiple points in time in the video and generates a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; an information acquisition unit that acquires selection information indicating which of the plurality of objects to be viewed has been selected as the tracking object; and an output unit that, based on the selection information, outputs the tracking video from the plurality of tracking videos in which the tracking object appears.
[0007] In addition, a content processing program according to another aspect of the present invention causes a computer to execute a video acquisition process for acquiring content data including a video; a generation process for detecting each of a plurality of objects to be viewed from the video and generating position information indicating the position of each characteristic part of the plurality of objects to be viewed at multiple points in time in the video, and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation process for cutting out the partial areas corresponding to the plurality of objects to be viewed from each frame corresponding to multiple points in time in the video based on the position information and the size information, and generating a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; a result acquisition process for acquiring selection information indicating which of the plurality of objects to be viewed has been selected as the tracking object; and an output process for outputting the tracking video in which the tracking object appears from among the plurality of tracking videos based on the selection information.
[0008] In addition, a content processing method according to another aspect of the present invention includes a generation step in which a computer detects each of a plurality of objects to be viewed from a video included in the content, and generates position information indicating the position of each characteristic part of the plurality of objects to be viewed at multiple points in time in the video, and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation step in which the computer cuts out each of the partial areas corresponding to the plurality of objects to be viewed from each frame corresponding to multiple points in time in the video based on the position information and the size information, and generates a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; a result acquisition step in which the computer acquires selection information indicating which of the plurality of objects to be viewed has been selected as the tracking object; and an output step in which the computer outputs, from the plurality of tracking videos, the tracking video in which the tracking object is shown based on the selection information. [Effects of the Invention]
[0009] According to one aspect of the present invention, a video showing only the tracking target can be displayed using a simple mechanism. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram showing an example of a functional configuration of a content playback system according to a first embodiment of the present invention. [Figure 2] 10 is an image diagram showing an example of the functions of a reception unit included in a terminal device of the system. FIG. [Figure 3] 10 is an image diagram showing an example of the function of an information generating unit included in a content processing device of the system. FIG. [Figure 4] FIG. 10 is a schematic diagram showing a state in which a terminal device of the system displays a tracking image. [Figure 5] FIG. 10 is a block diagram showing an example of a functional configuration of a content playback system according to a second embodiment of the present invention. [Figure 6] 1 is a flowchart showing an example of the flow of a content processing method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0011] <First embodiment of content playback system> Hereinafter, one embodiment of the present invention will be described in detail.
[0012] [Configuration of Content Playback System 100] 1, the content playback system 100 includes a content server 1, a content processing device 2, and a terminal device 3. The content processing device 2 may be configured integrally with the content server 1.
[0013] [Content Server 1] The content server 1 communicates with the terminal device 3. The content server 1 according to this embodiment communicates with the terminal device 3 wirelessly. The content server 1 according to this embodiment also communicates with the content processing device 2. The content server 1 according to this embodiment also communicates with the content processing device 2 wirelessly. Note that the content server 1 may be configured to communicate with the terminal device 3 and the content processing device 2 via a wired connection.
[0014] The content server 1 also stores video content data to be transmitted to the content processing device 2 and the terminal device 3. The content data stored in the content server 1 according to this embodiment includes audio. The content data stored in the content server 1 according to this embodiment also includes data obtained by filming a music event.
[0015] Furthermore, when the content server 1 receives a transmission request signal from the content processing device 2, it transmits content data corresponding to the transmission request signal to the content processing device 2 and the terminal device 3. Upon receiving the content data, the terminal device 3 plays back a video based on the data. The operation of the content processing device 2 upon receiving the content data will be described later.
[0016] [Content processing device 2] The content processing device 2 includes a communication unit 21, a storage unit 22, and a calculation unit .
[0017] (Communications Department 21) The communication unit 21 is the output unit in the first embodiment. The communication unit 21 communicates with the terminal device 3. The communication unit 21 according to this embodiment communicates wirelessly with the terminal device 3. The communication unit 21 according to this embodiment also communicates with the content server 1. The communication unit 21 according to this embodiment also communicates wirelessly with the content server 1. The communication unit 21 according to this embodiment is configured with a communication module. Note that the communication unit 21 may be configured with a terminal for connecting a cable for wired communication with the terminal device 3 or the content server 1.
[0018] (Storage unit 22) The storage unit 22 stores a content processing program. The content processing program is used to cause a computer to function as the content processing device 2. The storage unit 22 according to this embodiment is configured with a semiconductor memory, a hard disk drive, and the like.
[0019] (Computation unit 23) The calculation unit 23 includes a video acquisition unit 231, an information generation unit 232, a video generation unit 233, and an information acquisition unit 234. The calculation unit 23 according to this embodiment further includes an output processing unit 235. The calculation unit 23 according to this embodiment is configured with a processor. When this processor executes a content processing program stored in the storage unit 22, the processor functions as the video acquisition unit 231, the information generation unit 232, the video generation unit 233, the information acquisition unit 234, and the output processing unit 235.
[0020] Video Acquisition Unit 231 The video acquisition unit 231 executes video acquisition processing. In the video acquisition processing, the video acquisition unit 231 acquires content data including video. The video acquisition unit 231 according to this embodiment acquires content data received by the communication unit 21 from the content server 1. Note that if the content processing device 2 is configured integrally with the content server 1, the video acquisition unit 231 may acquire data read from the storage unit 22 or the like.
[0021] ·Information generation unit 232 When the video acquisition unit 231 acquires the video data, the information generation unit 232 executes information generation processing. In the information generation processing, the information generation unit 232 detects multiple objects of appreciation W from the video. Then, the information generation unit 232 generates position information and size information of the multiple objects of appreciation W at multiple points in time in the video. The "objects of appreciation W" are, for example, people. The objects of appreciation W may also be animals or objects (such as musical instruments). The "position information" is information indicating the position of each characteristic part of the multiple objects of appreciation W. The "size information" is information indicating the size of the partial region R. The "partial region R" is a region that includes at least the characteristic parts of the object of appreciation W. If the object of appreciation W is a person, the "characteristic part" is, for example, the person's face. The characteristic part may also be an object (such as a musical instrument) attached to the object of appreciation W (held by the person). If the object of appreciation W is a person, the partial region R may include a region that includes only the face, a region that includes the upper body, and a region that includes the entire body. The shape of the partial region R is preferably similar to the shape of the display region of the playback unit 32 of the terminal device 3 (for example, a rectangle with the same aspect ratio). The "time point" may be a time (x seconds after the start of playback of the video) or a frame. As shown in FIG. 3, the information generation unit 232 according to this embodiment generates size information such that the ratio b / a of the size b of the characteristic part of the tracking target T to the size a (width, height, or area) of the partial region R is within a predetermined range (lower limit c≦b / a≦upper limit d).
[0022] Furthermore, the information generation unit 232 according to this embodiment generates position information and size information using a trained model. The trained model is constructed so as to receive data of content showing the viewing object W as input and output position information and size information of the viewing object W. This makes it possible to detect the position of the tracking object T without constructing a complex algorithm.
[0023] Video Generation Unit 233 When the information generation unit 232 generates the position information and size information, the video generation unit 233 executes a video generation process. In the video generation process, the video generation unit 233 first cuts out partial regions R corresponding to multiple viewing objects W from each frame corresponding to multiple time points in the video based on the position information and size information. Then, the video generation unit 233 generates a tracking video for each viewing object W, with the cut-out partial regions R as frames. The video generation unit according to this embodiment generates a tracking video in which the partial regions R are enlarged so as to match the display area of the video.
[0024] The moving image generating unit 233 according to this embodiment executes a process of emphasizing, from among the sounds included in the content, components corresponding to the voice uttered by the person who has become the tracking target T or the sound emitted by the instrument played by that person. Specifically, the moving image generating unit 233 extracts, from the sound, frequency components corresponding to the voice uttered by the person who has become the tracking target T or the sound emitted by the instrument played by that person, amplifies the extracted frequency components, and combines them with the original sound.
[0025] ·Information acquisition unit 234 The information acquisition unit 234 executes information acquisition processing. In the information acquisition processing, the information acquisition unit 234 acquires selection information. The selection information is information indicating the viewing target W selected as the tracking target T from among a plurality of viewing targets W. The "tracking target T" is the viewing target W that the viewer particularly wants to view (only) from among the viewing targets W (candidates for tracking targets) appearing in the video, and is the viewing target W selected by the viewer via the terminal device 3 described below. The information acquisition unit 234 according to this embodiment acquires the selection information received by the communication unit 21 from the terminal device 3.
[0026] Output processing unit 235 When the information acquisition unit 234 acquires the selection information, the output processing unit 235 executes output processing. In the output processing, the output processing unit 235 controls the communication unit 21 (output unit). As a result, the communication unit 21 according to this embodiment transmits data of a tracking video in which the tracking target T appears among a plurality of tracking videos to the terminal device 3 (outputs the tracking video) based on the selection information. As described above, the video generation unit 233 according to this embodiment executes processing to emphasize audio components included in the content. Therefore, the communication unit 21 according to this embodiment outputs data of the tracking video together with data of audio whose components have been emphasized.
[0027] [Terminal device 3] The terminal device 3 includes a terminal communication unit 31, a playback unit 32, and an operation unit 33. The terminal device 3 according to this embodiment further includes a terminal storage unit 34 and a terminal calculation unit 35.
[0028] (Terminal communication unit 31) The terminal communication unit 31 communicates with the content processing device 2. The terminal communication unit 31 according to this embodiment communicates with the content processing device 2 wirelessly. The terminal communication unit 31 according to this embodiment also communicates with the content server 1. The terminal communication unit 31 according to this embodiment also communicates with the content server 1 wirelessly. The terminal communication unit 31 according to this embodiment is configured with a communication module. Note that the terminal communication unit 31 may be configured to communicate with the content server 1 and the content processing device 2 via a wired connection.
[0029] (Playback section 32) The playback unit 32 is configured to be able to display various moving images. The playback unit 32 according to this embodiment is configured with a display device (for example, a liquid crystal display, an organic EL display, etc.).
[0030] (Operation unit 33) The operation unit 33 is operated by a user. The operations include a transmission request operation, a tracking target selection operation, etc. The transmission request operation is an operation for requesting the content server 1 to transmit data of content desired by the user. When a transmission request operation is performed, the terminal device 3 transmits a transmission request signal corresponding to the transmission request operation to the content server 1. The tracking target T selection operation will be described later. The operation unit 33 includes, for example, a physical button, a touch panel, a pointing device (e.g., a mouse), etc.
[0031] (Terminal storage unit 34) The terminal storage unit 34 stores a content processing program (application program). The content processing program is used to cause a computer to function as the terminal device 3. The terminal storage unit 34 according to this embodiment is configured with a semiconductor memory, a hard disk drive, etc.
[0032] (Terminal Calculation Unit 35) The terminal calculation unit 35 includes a terminal video acquisition unit 351, a reception unit 352, a transmission control unit 353, and a playback control unit 354. The terminal calculation unit 35 according to this embodiment is configured with a processor. When the processor executes a content processing program stored in the terminal storage unit 34, the processor functions as the terminal video acquisition unit 351, the reception unit 352, the transmission control unit 353, and the playback control unit 354.
[0033] ·Device video acquisition unit 351 The terminal video acquisition unit 351 is the video acquisition unit in embodiment 2. The terminal video acquisition unit 351 executes video acquisition processing. In the video acquisition processing, the terminal video acquisition unit 351 acquires content data. The video included in the content may be unprocessed (uncut), or may be a tracking video from which the partial region R has been cut out. The terminal video acquisition unit 351 in this embodiment acquires content data received by the terminal communication unit 31 from the content server 1.
[0034] Reception Department 352 When a tracking target selection operation is performed on the operation unit 33, the reception unit 352 executes a reception process. In the reception process, the reception unit 352 receives a selection of a tracking target T in a moving image based on an operation performed on the operation unit 33, and generates selection information. The reception unit 352 according to this embodiment receives, as a selection of a tracking target T, a tracking target selection operation performed on the operation unit 33 to select one of multiple viewing targets W displayed in the moving image. If the operation unit 33 is configured as a touch panel, touching the viewing target W in the moving image constitutes the tracking target selection operation. If the operation unit 33 is configured as a pointing device, placing a cursor on the viewing target W in the moving image and clicking constitutes the tracking target selection operation. The reception unit 352 according to this embodiment displays a frame F surrounding selectable viewing targets W displayed in the moving image, as shown in FIG. 2 . This allows the viewer to easily select a tracking target T.
[0035] In the reception process, the reception unit 352 may be configured to display a list of multiple names on the reproduction unit 32. The "name" includes the name of the appreciation object W, the name of an object attached to the appreciation object W (such as a musical instrument held by a person), the name of a group to which the appreciation object W belongs (such as a performance part), and the like. The reception unit 352 may be configured to receive, as a selection of the tracking object T, a tracking object selection operation to select one of the names from the list performed on the operation unit 33. In this case, the reception unit 352 refers to the correspondence between the appreciation object W in the video and the names in the list, which is stored in the storage unit 22, for example, and generates selection information indicating the appreciation object W corresponding to the selected name.
[0036] Transmission control unit 353 When the reception unit 352 generates the selection information, the transmission control unit 353 executes a transmission control process. In the transmission control process, the transmission control unit 353 controls the terminal communication unit 31. As a result, the terminal communication unit 31 transmits the selection information to the content processing device 2.
[0037] Playback control unit 354 When the terminal video acquisition unit 351 acquires content data, the playback control unit 354 executes playback control processing. In the terminal display control processing, the playback control unit 354 controls the playback unit 32. As a result, the playback unit 32 displays the video (including the tracking video) in the display area. As described above, the video generation unit 233 of the content processing device 2 generates a tracking video in which the partial region R is enlarged so that it matches the display area of the video. Therefore, the playback unit 32 displays the tracking video using the entire display area. As a result, even in a scene in which the tracking target T appears small due to the camera pulling back, the viewer can enjoy a tracking video in which the tracking target T is greatly enlarged, as shown in FIG. 4.
[0038] [Effects of the content processing device 2] In the content processing device 2 described above, the information generation unit 232 generates position information and size information, and the video generation unit 233 generates a tracking video for each of the multiple tracking targets T based on the position information and size information. Then, the communication unit 21 (output unit) transmits data of the tracking video corresponding to the selection information acquired by the information acquisition unit 234 to the terminal device 3. Therefore, according to the content processing device 2 or the content reproduction system 100 including the content processing device 2, the terminal device 3 can display in the display area a tracking video that shows only the tracking target T (partial region R) selected by the viewer. As a result, even if multiple viewing targets W are shown in the original video, the viewer can view only the tracking target T.
[0039] <Second embodiment of content playback system> Next, another embodiment of the present invention will be described below. For the sake of convenience, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.
[0040] [Configuration of Content Playback System 100A] As shown in FIG. 5, the content reproduction system 100A includes a content server 1 similar to that of the content reproduction system 100 according to the first embodiment, and also includes a terminal device 3A.
[0041] [Terminal Device 3A] The terminal device 3A is a content processing device in embodiment 2. The terminal device 3A includes a terminal communication unit 31, a playback unit 32, and an operation unit 33 similar to those of the terminal device 3 in embodiment 1, as well as a terminal storage unit 34A and a terminal calculation unit 35A.
[0042] (Terminal storage unit 34A) The terminal storage unit 34A stores a content processing program (application program). The content processing program is used to cause a computer to function as the terminal device 3A. The terminal storage unit 34A according to this embodiment is configured with a semiconductor memory, a hard disk drive, etc.
[0043] (Terminal Calculation Unit 35A) The terminal computing unit 35A includes a terminal video acquisition unit 351 and a reception unit 352 similar to those of the terminal computing unit 35 according to the first embodiment, an information generation unit 232 and a video generation unit 233 similar to those of the computing unit 23 of the content processing device 2 according to the first embodiment, as well as an information acquisition unit 234A and an output processing unit 235A. The terminal computing unit 35A according to the present embodiment is configured with a processor. When this processor executes a content processing program stored in the terminal storage unit 34A, the processor functions as the terminal video acquisition unit 351, the reception unit 352, the information generation unit 232, the video generation unit 233, the information acquisition unit 234A, and the output processing unit 235A.
[0044] ·Information acquisition section 234A The information acquisition unit 234A according to this embodiment acquires the selection information generated by the reception unit 352.
[0045] Output processing unit 235A The output processing unit 235A according to the present embodiment controls the playback unit 32 when the information acquisition unit 234A acquires selection information. As a result, the playback unit 32 according to the present embodiment displays data of a tracking video in which the tracking target T appears, among a plurality of tracking videos, based on the selection information. The video generation unit 233 according to the present embodiment also executes processing to emphasize audio components included in content, similar to the video generation unit 233 according to the first embodiment. Therefore, the playback unit 32 according to the present embodiment displays the tracking video together with audio whose components have been emphasized. Furthermore, the output processing unit 235A according to the present embodiment also controls the playback unit 32 (output unit) when the terminal video acquisition unit 351 acquires content data. As a result, the playback unit 32 plays back video that has not been processed (cut out).
[0046] [Operation and effect of terminal device 3A] In the terminal device 3A described above, the information generation unit 232 generates position information and size information, and the video generation unit 233 generates a tracking video for each of the multiple tracking targets T based on the position information and size information. Then, the playback unit 32 (output unit) displays the tracking video corresponding to the selection information acquired by the information acquisition unit 234. Therefore, according to the terminal device 3A or the content playback system 100A including the terminal device 3A, even if multiple viewing targets W are shown in the original video, the viewer can view only the tracking target T.
[0047] <Embodiment of Content Processing Method> Next, other embodiments of the present invention will be described below. For the sake of convenience, the same reference numerals will be used to designate components having the same functions as those described in the above embodiments, and the description thereof will not be repeated.
[0048] [Flow of content processing method S100] As shown in FIG. 6, the content processing method S100 includes an information generating step S1, a moving image generating step S2, an information obtaining step S3, and an outputting step S4.
[0049] [Information generation step S1] In the initial information generation step S1, multiple viewing objects W are detected from a video. Then, position information indicating the position of each characteristic part of the multiple viewing objects W at multiple points in time in the video, and size information indicating the size of a partial region R including at least the characteristic part of the viewing object W, are generated. The detection of the viewing objects W and the generation of the position information and size information may be performed by the content processing device 2 according to embodiment 1 or the terminal device 3A according to embodiment 2, or may be performed by another device. Furthermore, the device that detects the viewing objects W may be different from the device that generates the position information and size information.
[0050] [Video generation step S2] After generating the position information and size information, a video generation step S2 is performed. In the video generation step S2, partial regions R corresponding to multiple viewing targets W are respectively cut out from each frame corresponding to multiple time points in the video based on the position information and size information. Then, a tracking video is generated for each viewing target W, with the cut-out partial regions R as frames. The cutting out of the partial regions R and the generation of the tracking video may be performed by the content processing device 2 according to the first embodiment or the terminal device 3A according to the second embodiment, or by another device. Furthermore, the device that cuts out the partial regions R may be different from the device that generates the tracking video.
[0051] [Information acquisition step S3] After generating the tracking video, an information acquisition step S3 is performed. In the information acquisition step S3, selection information is acquired that indicates the viewing object W selected as the tracking object T from among the multiple viewing objects. The selection information may be acquired by the content processing device 2 according to the first embodiment, the terminal device 3A according to the second embodiment, or another device.
[0052] [Output step S4] After receiving the selection of the tracking target T, an output step S4 is performed. In the output step S4, a tracking video showing the tracking target T is output from among a plurality of tracking videos based on the selection information. The tracking video may be output by the content processing device 2 according to the first embodiment or the terminal device 3A according to the second embodiment, or by another device.
[0053] [Effects of the content processing method S100] In the content processing method S100 described above, position information and size information are generated in the information generating step S1, and a tracking video of each of the multiple tracking targets T is generated based on the position information and size information in the video generating step S2. Then, the tracking video corresponding to the selection information acquired in the information acquiring step S3 is output in the output step S4. Therefore, according to the content processing method S100, it becomes possible to display in the display area a tracking video that shows only the tracking target T (partial region R) selected by the viewer. As a result, even if multiple viewing targets W are shown in the original video, the viewer can view only the tracking target T.
[0054] <Modification> The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.
[0055] For example, the content processing device 2 or the terminal device 3A may be configured so that the information acquisition unit 234 acquires the selection information before the video generation unit 233 generates the tracking video. Then, the video generation unit 233 may be configured to generate only the tracking video in which the tracking target T appears, based on the selection information.
[0056] Furthermore, the content processing device 2 or the terminal device 3A may be configured so that the information generation unit 232 generates the position information and size information, and the video generation unit 233 generates the tracking video alternately for each frame, and at the same cycle as the frame rate of the video. This shortens the time lag between acquiring content data and playing the tracking video, making it possible to handle live video as well.
[0057] Furthermore, the content playback system 100 (content playback system 100A) may include an authentication device that authenticates identification information. In this case, the terminal device 3, 3A may be configured to accept input of identification information and transmit it to the authentication device. The authentication device may be configured to transmit connection information for accessing the content server 1 to the terminal device 3, 3A when playback of content on the terminal device 3, 3A is permitted.
[0058] The content server 1 may be configured to encrypt content data and transmit the encrypted data to the content processing device 2 (terminal device 3A). In this case, the content playback system 100 (content playback system 100A) may include a key server that generates a decryption key for decrypting the encrypted data. The content processing device 2 (terminal device 3A) may be configured to obtain the decryption key from the key server and decrypt the encrypted data using the decryption key.
[0059] The content processing program may be stored non-transitory on one or more computer-readable storage media. The storage media may or may not be included in the device. In the latter case, the content processing program may be supplied to the device via any wired or wireless transmission medium.
[0060] Furthermore, some or all of the functions of the control blocks can be realized by logic circuits. For example, an integrated circuit in which a logic circuit that functions as each of the control blocks is formed is also included in the scope of the present invention. In addition, the functions of the control blocks can also be realized by, for example, a quantum computer.
[0061] <Summary> A content processing device according to aspect 1 of the present invention comprises a video acquisition unit that acquires content data including a video; an information generation unit that detects a plurality of objects to be viewed from the video and generates position information indicating the position of each characteristic part of the plurality of objects to be viewed at a plurality of points in time in the video and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation unit that, based on the position information and the size information, cuts out the partial areas corresponding to the plurality of objects to be viewed from each frame of the video corresponding to a plurality of points in time, and generates a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; an information acquisition unit that acquires selection information indicating which of the plurality of objects to be viewed has been selected as the tracking object; and an output unit that, based on the selection information, outputs the tracking video in which the tracking object is shown from among the plurality of tracking videos.
[0062] A content processing device according to aspect 2 of the present invention may be configured in the above aspect 1 such that the information generation unit receives data of content showing the object of viewing as input and generates the position information and size information of the object of viewing using a trained model that outputs the position information and size information of the object of viewing.
[0063] A content processing device according to a third aspect of the present invention may be configured in the first or second aspect above, where the object of appreciation is a person, and the characteristic part is a face of the person.
[0064] A content processing device according to aspect 4 of the present invention may be configured such that, in any of aspects 1 to 3 above, the information generation unit generates size information such that the ratio of the size of the characteristic portion to the size of the partial region is within a predetermined range.
[0065] A content processing device according to aspect 5 of the present invention may be configured in the above-mentioned aspect 4 such that the video generation unit generates the tracking video in which the partial area is enlarged so as to match the display area of the video.
[0066] A content processing device according to a sixth aspect of the present invention may be configured in any one of the first to fifth aspects above, wherein the content includes audio.
[0067] A content processing device according to aspect 7 of the present invention may be configured in the above-mentioned aspect 6 such that the object of appreciation is a person and the content data includes data obtained by filming a music event.
[0068] A content processing device according to aspect 8 of the present invention may be configured such that, in aspect 7 above, the video generation unit performs processing to emphasize components of the audio corresponding to the voice of the person being tracked or the sound of an instrument played by that person, and the output unit outputs the tracking video together with the audio in which the components have been emphasized.
[0069] A content processing program according to a ninth aspect of the present invention causes a computer to execute a video acquisition process for acquiring data of content including a video; a generation process for detecting each of a plurality of objects to be viewed from the video and generating position information indicating the position of each characteristic part of the plurality of objects to be viewed at multiple points in time in the video and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation process for cutting out the partial areas corresponding to the plurality of objects to be viewed from each frame corresponding to multiple points in time in the video based on the position information and the size information and generating a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; a result acquisition process for acquiring selection information indicating which of the plurality of objects to be viewed has been selected as the tracking target; and an output process for outputting the tracking video in which the tracking target is shown from among the plurality of tracking videos based on the selection information.
[0070] A content processing method according to aspect 10 of the present invention includes a generation step in which a computer detects each of a plurality of objects to be viewed from a video included in the content, and generates position information indicating the position of each characteristic part of the plurality of objects to be viewed at multiple points in time in the video, and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation step in which the computer cuts out each of the partial areas corresponding to the plurality of objects to be viewed from each frame corresponding to multiple points in time in the video based on the position information and the size information, and generates a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; a result acquisition step in which the computer acquires selection information indicating which of the plurality of objects to be viewed has been selected as the tracking object; and an output step in which the computer outputs, from the plurality of tracking videos, the tracking video in which the tracking object is shown based on the selection information. [Explanation of symbols]
[0071] 100,100A Content Playback System 1 Content Server 2 Content Processing Device 21 Communication unit (output unit) 22 Memory section 23 Arithmetic section 231 Video Acquisition Unit 232 Information generation section 233 Video Generation Unit 234, 234A Information acquisition section 235, 235A output processing section 3,3A terminal equipment 31 Terminal communication unit 32 Playback section (output section) 33 Operation section 34, 34A Terminal memory section 35, 35A Terminal calculation unit 351 Terminal video acquisition unit (video acquisition unit) 352 Reception Department 353 Transmission control section 354 Playback control unit S100 Content Processing Method S1 Information generation step S2 Video generation step S3 information acquisition step S4 Output Step
Claims
1. a video acquisition unit that acquires content data including videos; An information generating unit that detects each of a plurality of objects to be viewed from the video and generates position information indicating the position of each characteristic part of the plurality of objects to be viewed at a plurality of points in time in the video, and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation unit that, based on the position information and the size information, cuts out the partial regions corresponding to the plurality of objects of appreciation from each frame of the video corresponding to a plurality of time points so that the ratio of the size of the characteristic portion to the partial region falls within a predetermined range, and generates tracking video data for each object of appreciation, with the plurality of cut-out partial regions as frames; an information acquisition unit that acquires selection information indicating an object selected as a tracking object from among the plurality of objects; an output unit that outputs the tracking video in which the tracking target appears, from among the plurality of tracking videos, based on the selection information; Equipped with Content processing device.
2. The information generation unit receives data of content in which the object of appreciation is displayed as an input, and generates the position information and the size information of the object of appreciation using a trained model that outputs the position information and the size information of the object of appreciation. The content processing device according to claim 1 .
3. The object of appreciation is a person, The feature is the face of the person. The content processing device according to claim 1 .
4. the information generation unit generates size information such that the ratio of the size of the characteristic portion to the size of the partial region falls within a predetermined range. The content processing device according to claim 1 .
5. the video generation unit generates the tracking video in which the partial region is enlarged so as to coincide with a display region of the video. The content processing device according to claim 4 .
6. the content includes audio; The content processing device according to claim 1 .
7. The object of appreciation is a person, The content data includes data obtained by filming a music event. The content processing device according to claim 6 .
8. the video generation unit executes a process of emphasizing a component of the audio corresponding to a voice uttered by the person who is the tracked target or a sound emitted by an instrument played by the person, The output unit outputs the tracking video together with the audio in which the component is emphasized. The content processing device according to claim 7 .
9. On the computer, A video acquisition process for acquiring data of content including video; A generation process of detecting each of a plurality of objects to be viewed from the video and generating position information indicating the position of each characteristic part of the plurality of objects to be viewed at a plurality of points in time of the video, and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation process for extracting the partial regions corresponding to the plurality of objects of appreciation from each frame of the video corresponding to a plurality of time points based on the position information and the size information so that the ratio of the size of the characteristic portion to the partial region falls within a predetermined range, and generating tracking video data for each object of appreciation, the tracking video data having the extracted partial regions as frames; A result acquisition process for acquiring selection information indicating an object selected as a tracking object from among the plurality of objects; an output process of outputting the tracking video in which the tracking target appears among the plurality of tracking videos based on the selection information; Execute Content processing program.
10. A generation step in which a computer detects each of a plurality of objects to be viewed from a video included in the content, and generates position information indicating the position of each characteristic part of the plurality of objects to be viewed at a plurality of points in time in the video, and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation step in which the computer cuts out the partial regions corresponding to the plurality of objects of appreciation from each frame of the video corresponding to a plurality of time points based on the position information and the size information so that the ratio of the size of the characteristic part to the partial region falls within a predetermined range, and generates tracking video data for each object of appreciation, with the plurality of cut-out partial regions as frames; a result acquisition step in which the computer acquires selection information indicating an object selected as a tracking object from among the plurality of objects; an output step in which the computer outputs, based on the selection information, one of the plurality of tracking videos in which the tracking target is captured; Including, Content processing methods.