Content processing device, content processing program, and content processing method

The content processing device addresses the limitation of conventional technologies by detecting and generating position and size information for objects in video data, enabling the display of a tracking target through viewer selection, thus overcoming the need for multiple cameras and acoustic analysis.

WO2025253476A1PCT designated stage Publication Date: 2025-12-11RADIUS CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/020297
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-04
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Conventional video processing technologies require multiple cameras and acoustic analysis to handle content, limiting their ability to display a video showing only a tracking target.

Method used

A content processing device that acquires video data, detects and generates position and size information for objects, cuts out partial areas corresponding to these objects, and outputs a tracking video based on viewer selection, allowing for a simple mechanism to display the tracking target.

Benefits of technology

Enables the display of a video showing only the tracking target, even when multiple objects are present, using a simple mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024020297_11122025_PF_FP_ABST
    Figure JP2024020297_11122025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention makes it possible to display a moving image in which only a tracking target T appears using a simple mechanism. A content processing device (2) comprises: a moving image acquisition unit (231) that acquires data of content including a moving image; an information generation unit (232) that detects each of a plurality of viewing targets from the moving image and generates, at a plurality of time points of the moving image, position information indicating the position of individual feature portions of the plurality of viewing targets and size information indicating the size of a partial region including at least the feature portion of the viewing target; a moving image generation unit (233) that, on the basis of the position information and the size information, cuts out respective partial regions corresponding to the plurality of viewing targets from individual frames corresponding to the plurality of time points of the moving image and generates, for each viewing target, a tracking moving image in which the plurality of cut-out partial regions serve as frames; an information acquisition unit that acquires selection information indicating a viewing target selected as the tracking target from among the plurality of viewing targets; and an output unit (21) that, on the basis of the selection information, outputs a tracking moving image in which the tracking target appears from among the plurality of tracking moving images.
Need to check novelty before this filing date? Find Prior Art

Description

Content processing device, content processing program, and content processing method

[0001] The present invention relates to a content processing device, a content processing program, and a content processing method.

[0002] Various video processing technologies have been proposed for providing live video customized for each viewer. For example, Patent Document 1 describes a video data processing device that includes an input means for inputting video captured by multiple cameras arranged at a venue, an analysis data acquisition means for acquiring music analysis data that is the result of an acoustic analysis of the music played at the venue, a switching means for switching the video to be output based on the music analysis data, and an output means for outputting the video switched by the switching means.

[0003] Japanese Patent Application Publication No. 2021-153257

[0004] However, in the conventional technology described in Patent Document 1, the target must be photographed with multiple cameras, and since switching is performed using the results of acoustic analysis, it is not possible to handle content that is only video.

[0005] One aspect of the present invention has been made in view of the above-mentioned problems, and its purpose is to make it possible to display a video showing only a tracking target using a simple mechanism.

[0006] In order to solve the above problems, a content processing device according to one aspect of the present invention comprises a video acquisition unit that acquires content data including a video; an information generation unit that detects a plurality of objects to be viewed from the video and generates position information indicating the position of each characteristic part of the plurality of objects to be viewed at multiple points in time in the video and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation unit that, based on the position information and the size information, cuts out the partial areas corresponding to the plurality of objects to be viewed from each frame corresponding to multiple points in time in the video and generates a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; an information acquisition unit that acquires selection information indicating which of the plurality of objects to be viewed has been selected as the tracking object; and an output unit that, based on the selection information, outputs the tracking video from the plurality of tracking videos in which the tracking object appears.

[0007] In addition, a content processing program according to another aspect of the present invention causes a computer to execute a video acquisition process for acquiring content data including a video; a generation process for detecting each of a plurality of objects to be viewed from the video and generating position information indicating the position of each characteristic part of the plurality of objects to be viewed at multiple points in time in the video, and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation process for cutting out the partial areas corresponding to the plurality of objects to be viewed from each frame corresponding to multiple points in time in the video based on the position information and the size information, and generating a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; a result acquisition process for acquiring selection information indicating which of the plurality of objects to be viewed has been selected as the tracking object; and an output process for outputting the tracking video in which the tracking object appears from among the plurality of tracking videos based on the selection information.

[0008] In addition, a content processing method according to another aspect of the present invention includes a generation step in which a computer detects each of a plurality of objects to be viewed from a video included in the content, and generates position information indicating the position of each characteristic part of the plurality of objects to be viewed at multiple points in time in the video, and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation step in which the computer cuts out each of the partial areas corresponding to the plurality of objects to be viewed from each frame corresponding to multiple points in time in the video based on the position information and the size information, and generates a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; a result acquisition step in which the computer acquires selection information indicating which of the plurality of objects to be viewed has been selected as the tracking object; and an output step in which the computer outputs, from the plurality of tracking videos, the tracking video in which the tracking object is shown based on the selection information.

[0009] According to one aspect of the present invention, a video showing only the tracking target can be displayed using a simple mechanism.

[0010] FIG. 1 is a block diagram showing an example of the functional configuration of a content reproduction system according to a first embodiment of the present invention. FIG. 2 is an image diagram showing an example of the functions of a reception unit included in a terminal device of the system. FIG. 3 is an image diagram showing an example of the functions of an information generation unit included in a content processing device of the system. FIG. 4 is a schematic diagram showing a state in which a terminal device of the system is displaying a tracking image. FIG. 5 is a block diagram showing an example of the functional configuration of a content reproduction system according to a second embodiment of the present invention. FIG. 6 is a flowchart showing an example of the flow of a content processing method according to an embodiment of the present invention.

[0011] <First Embodiment of Content Reproduction System> Hereinafter, a first embodiment of the present invention will be described in detail.

[0012] 1, the content playback system 100 includes a content server 1, a content processing device 2, and a terminal device 3. The content processing device 2 may be configured integrally with the content server 1.

[0013] [Content Server 1] The content server 1 communicates with the terminal device 3. The content server 1 according to this embodiment communicates with the terminal device 3 wirelessly. The content server 1 according to this embodiment also communicates with the content processing device 2. The content server 1 according to this embodiment also communicates with the content processing device 2 wirelessly. Note that the content server 1 may be configured to communicate with the terminal device 3 and the content processing device 2 via a wired connection.

[0014] The content server 1 also stores video content data to be transmitted to the content processing device 2 and the terminal device 3. The content data stored in the content server 1 according to this embodiment includes audio. The content data stored in the content server 1 according to this embodiment also includes data obtained by filming a music event.

[0015] Furthermore, when the content server 1 receives a transmission request signal from the content processing device 2, it transmits content data corresponding to the transmission request signal to the content processing device 2 and the terminal device 3. Upon receiving the content data, the terminal device 3 plays a video based on the data. The operation of the content processing device 2 upon receiving the content data will be described later.

[0016] [Content Processing Apparatus 2] The content processing apparatus 2 includes a communication unit 21, a storage unit 22, and a calculation unit 23.

[0017] (Communication unit 21) The communication unit 21 is the output unit in embodiment 1. The communication unit 21 communicates with the terminal device 3. The communication unit 21 in this embodiment communicates with the terminal device 3 wirelessly. The communication unit 21 in this embodiment also communicates with the content server 1. The communication unit 21 in this embodiment also communicates with the content server 1 wirelessly. The communication unit 21 in this embodiment is configured with a communication module. Note that the communication unit 21 may be configured with a terminal for connecting a cable for wired communication with the terminal device 3 or the content server 1.

[0018] (Storage Unit 22) The storage unit 22 stores a content processing program. The content processing program is used to cause a computer to function as the content processing device 2. The storage unit 22 according to this embodiment is configured with a semiconductor memory, a hard disk drive, etc.

[0019] (Calculation unit 23) The calculation unit 23 includes a video acquisition unit 231, an information generation unit 232, a video generation unit 233, and an information acquisition unit 234. The calculation unit 23 according to this embodiment further includes an output processing unit 235. The calculation unit 23 according to this embodiment is configured with a processor. When this processor executes a content processing program stored in the storage unit 22, the processor functions as the video acquisition unit 231, the information generation unit 232, the video generation unit 233, the information acquisition unit 234, and the output processing unit 235.

[0020] Video acquisition unit 231 The video acquisition unit 231 executes video acquisition processing. In the video acquisition processing, the video acquisition unit 231 acquires content data including videos. The video acquisition unit 231 according to this embodiment acquires content data received by the communication unit 21 from the content server 1. Note that if the content processing device 2 is configured integrally with the content server 1, the video acquisition unit 231 may acquire data read from the storage unit 22 or the like.

[0021] Information Generation Unit 232 When the video acquisition unit 231 acquires the video acquisition unit 231, the information generation unit 232 executes information generation processing. In the information generation processing, the information generation unit 232 detects multiple objects of appreciation W from the video. Then, the information generation unit 232 generates position information and size information for the multiple objects of appreciation W at multiple points in time in the video. The "objects of appreciation W" are, for example, people. Note that the objects of appreciation W may also be animals or objects (such as musical instruments). The "position information" is information indicating the position of each characteristic part of the multiple objects of appreciation W. The "size information" is information indicating the size of the partial region R. The "partial region R" is a region that includes at least the characteristic part of the object of appreciation W. If the object of appreciation W is a person, the "characteristic part" is, for example, the person's face. Note that the characteristic part may also be an object (such as a musical instrument) attached to the object of appreciation W (held by the person) that is attached to the object of appreciation W. Furthermore, if the viewing target W is a person, the partial region R may include a region including only the face, a region including the upper body, and a region including the entire body. The shape of the partial region R is preferably similar to the shape of the display area of ​​the playback unit 32 of the terminal device 3 (e.g., a rectangle with the same aspect ratio). The "time point" may be a time (x seconds after the start of video playback) or a frame. As shown in FIG. 3 , the information generation unit 232 according to this embodiment generates size information such that the ratio b / a of the size b of the characteristic portion of the tracking target T to the size a (width, height, or area) of the partial region R falls within a predetermined range (lower limit c≦b / a≦upper limit d).

[0022] Furthermore, the information generation unit 232 according to this embodiment generates position information and size information using a trained model. The trained model is constructed so as to receive data of content showing the viewing object W as input and output position information and size information of the viewing object W. This makes it possible to detect the position of the tracking object T without constructing a complex algorithm.

[0023] Video Generation Unit 233 When the information generation unit 232 generates the position information and size information, the video generation unit 233 executes a video generation process. In the video generation process, the video generation unit 233 first cuts out partial regions R corresponding to multiple viewing objects W from each frame corresponding to multiple time points in the video based on the position information and size information. The video generation unit 233 then generates a tracking video for each viewing object W, with the cut-out partial regions R as frames. The video generation unit according to this embodiment generates a tracking video in which the partial regions R are enlarged so as to match the display area of ​​the video.

[0024] The moving image generating unit 233 according to this embodiment executes a process of emphasizing, from among the sounds included in the content, components corresponding to the voice uttered by the person who has become the tracking target T or the sound emitted by the instrument played by that person. Specifically, the moving image generating unit 233 extracts, from the sound, frequency components corresponding to the voice uttered by the person who has become the tracking target T or the sound emitted by the instrument played by that person, amplifies the extracted frequency components, and combines them with the original sound.

[0025] Information Acquisition Unit 234 The information acquisition unit 234 executes information acquisition processing. In the information acquisition processing, the information acquisition unit 234 acquires selection information. The selection information is information indicating the viewing target W selected as the tracking target T from among multiple viewing targets W. The "tracking target T" is the viewing target W that the viewer particularly wants to view (only) from among the viewing targets W (candidates for tracking targets) appearing in the video, and is the viewing target W selected by the viewer via the terminal device 3, which will be described later. The information acquisition unit 234 according to this embodiment acquires the selection information received by the communication unit 21 from the terminal device 3.

[0026] Output Processing Unit 235 When the information acquisition unit 234 acquires the selection information, the output processing unit 235 executes output processing. In the output processing, the output processing unit 235 controls the communication unit 21 (output unit). As a result, the communication unit 21 according to this embodiment transmits data of a tracking video in which the tracking target T appears among a plurality of tracking videos to the terminal device 3 (outputs the tracking video) based on the selection information. As described above, the video generation unit 233 according to this embodiment executes processing to emphasize audio components included in the content. Therefore, the communication unit 21 according to this embodiment outputs the tracking video data together with audio data in which the components have been emphasized.

[0027] [Terminal Device 3] The terminal device 3 includes a terminal communication unit 31, a playback unit 32, and an operation unit 33. The terminal device 3 according to this embodiment further includes a terminal storage unit 34 and a terminal calculation unit 35.

[0028] (Device communication unit 31) The device communication unit 31 communicates with the content processing device 2. The device communication unit 31 according to this embodiment communicates with the content processing device 2 wirelessly. The device communication unit 31 according to this embodiment also communicates with the content server 1. The device communication unit 31 according to this embodiment also communicates with the content server 1 wirelessly. The device communication unit 31 according to this embodiment is configured with a communication module. Note that the device communication unit 31 may be configured to communicate with the content server 1 and the content processing device 2 via a wired connection.

[0029] (Playback Unit 32) The playback unit 32 is configured to be able to display various moving images. The playback unit 32 according to this embodiment is configured with a display device (for example, a liquid crystal display, an organic EL display, etc.).

[0030] (Operation Unit 33) The operation unit 33 is operated by the user. The operations include a transmission request operation, a tracking target selection operation, etc. The transmission request operation is an operation for requesting the content server 1 to transmit data of content desired by the user. When a transmission request operation is performed, the terminal device 3 transmits a transmission request signal corresponding to the transmission request operation to the content server 1. The tracking target T selection operation will be described later. The operation unit 33 includes, for example, a physical button, a touch panel, a pointing device (e.g., a mouse), etc.

[0031] (Terminal Storage Unit 34) The terminal storage unit 34 stores a content processing program (application program). The content processing program is used to cause a computer to function as the terminal device 3. The terminal storage unit 34 according to this embodiment is configured with a semiconductor memory, a hard disk drive, etc.

[0032] (Device Computing Unit 35) The device computing unit 35 includes a device video acquisition unit 351, a reception unit 352, a transmission control unit 353, and a playback control unit 354. The device computing unit 35 according to this embodiment is configured with a processor. When the processor executes a content processing program stored in the device storage unit 34, the processor functions as the device video acquisition unit 351, the reception unit 352, the transmission control unit 353, and the playback control unit 354.

[0033] Terminal video acquisition unit 351 The terminal video acquisition unit 351 is the video acquisition unit in embodiment 2. The terminal video acquisition unit 351 executes video acquisition processing. In the video acquisition processing, the terminal video acquisition unit 351 acquires content data. The video included in the content may be unprocessed (uncut), or may be a tracking video from which the partial region R has been cut out. The terminal video acquisition unit 351 in this embodiment acquires content data received by the terminal communication unit 31 from the content server 1.

[0034] Receiving Unit 352 When a tracking target selection operation is performed on the operation unit 33, the receiving unit 352 executes a reception process. In the reception process, the receiving unit 352 receives a selection of a tracking target T in a video based on the operation performed on the operation unit 33 and generates selection information. The receiving unit 352 according to the present embodiment receives, as a selection of a tracking target T, a tracking target selection operation performed on the operation unit 33 to select one of multiple viewing targets W displayed in the video. If the operation unit 33 is configured as a touch panel, touching the viewing target W in the video constitutes the tracking target selection operation. If the operation unit 33 is configured as a pointing device, placing a cursor on the viewing target W in the video and clicking constitutes the tracking target selection operation. The receiving unit 352 according to the present embodiment displays a frame F surrounding selectable viewing targets W displayed in the video, as shown in FIG. 2 . This allows the viewer to easily select a tracking target T.

[0035] In the reception process, the reception unit 352 may be configured to display a list of multiple names on the reproduction unit 32. The "name" includes the name of the appreciation object W, the name of an object attached to the appreciation object W (such as a musical instrument held by a person), the name of a group to which the appreciation object W belongs (such as a performance part), etc. The reception unit 352 may be configured to receive, as a selection of the tracking object T, a tracking object selection operation performed on the operation unit 33 to select one of the names from the list. In this case, the reception unit 352 refers to the correspondence between the appreciation object W in the video and the names in the list, which is stored in the storage unit 22, for example, and generates selection information indicating the appreciation object W corresponding to the selected name.

[0036] Transmission control unit 353 When the reception unit 352 generates selection information, the transmission control unit 353 executes a transmission control process. In the transmission control process, the transmission control unit 353 controls the terminal communication unit 31. As a result, the terminal communication unit 31 transmits the selection information to the content processing device 2.

[0037] Playback control unit 354 When the terminal video acquisition unit 351 acquires content data, the playback control unit 354 executes playback control processing. In the terminal display control processing, the playback control unit 354 controls the playback unit 32. As a result, the playback unit 32 displays the video (including the tracking video) in the display area. As described above, the video generation unit 233 of the content processing device 2 generates a tracking video in which the partial region R is enlarged so that it matches the display area of ​​the video. Therefore, the playback unit 32 displays the tracking video using the entire display area. As a result, even in a scene in which the tracking target T appears small due to the camera being pulled back, the viewer can enjoy a tracking video in which the tracking target T is greatly enlarged, as shown in FIG. 4.

[0038] [Operation and Effect of Content Processing Device 2] In the content processing device 2 described above, the information generation unit 232 generates position information and size information, and the video generation unit 233 generates a tracking video for each of the multiple tracking targets T based on the position information and size information. Then, the communication unit 21 (output unit) transmits data of the tracking video corresponding to the selection information acquired by the information acquisition unit 234 to the terminal device 3. Therefore, according to the content processing device 2 or the content playback system 100 including the content processing device 2, the terminal device 3 can display in the display area a tracking video that shows only the tracking target T (partial region R) selected by the viewer. As a result, even if multiple viewing targets W are shown in the original video, the viewer can view only the tracking target T.

[0039] <Second Embodiment of Content Reproduction System> Next, another embodiment of the present invention will be described below. For convenience of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.

[0040] [Configuration of Content Reproduction System 100A] As shown in FIG. 5, the content reproduction system 100A includes a content server 1 similar to that of the content reproduction system 100 according to the first embodiment, and also includes a terminal device 3A.

[0041] [Terminal Device 3A] The terminal device 3A is a content processing device in embodiment 2. The terminal device 3A includes a terminal communication unit 31, a playback unit 32, and an operation unit 33 similar to those of the terminal device 3 in embodiment 1, as well as a terminal storage unit 34A and a terminal calculation unit 35A.

[0042] (Terminal storage unit 34A) The terminal storage unit 34A stores a content processing program (application program). The content processing program is used to cause a computer to function as the terminal device 3A. The terminal storage unit 34A according to this embodiment is configured with a semiconductor memory, a hard disk drive, etc.

[0043] (Device Computing Unit 35A) The device computing unit 35A includes a device video acquisition unit 351 and a receiving unit 352 similar to those of the device computing unit 35 according to the first embodiment, an information generation unit 232 and a video generation unit 233 similar to those of the computing unit 23 of the content processing device 2 according to the first embodiment, as well as an information acquisition unit 234A and an output processing unit 235A. The device computing unit 35A according to the present embodiment is configured with a processor. When this processor executes a content processing program stored in the device storage unit 34A, the processor functions as the device video acquisition unit 351, the receiving unit 352, the information generation unit 232, the video generation unit 233, the information acquisition unit 234A, and the output processing unit 235A.

[0044] Information Acquisition Unit 234A The information acquisition unit 234A according to this embodiment acquires the selection information generated by the reception unit 352.

[0045] Output Processing Unit 235A The output processing unit 235A according to this embodiment controls the playback unit 32 when the information acquisition unit 234A acquires selection information. As a result, the playback unit 32 according to this embodiment displays data of a tracking video in which the tracking target T appears, out of multiple tracking videos, based on the selection information. The video generation unit 233 according to this embodiment also executes processing to emphasize audio components included in content, similar to the video generation unit 233 according to the first embodiment. Therefore, the playback unit 32 according to this embodiment displays the tracking video together with audio whose components have been emphasized. Furthermore, the output processing unit 235A according to this embodiment also controls the playback unit 32 (output unit) when the terminal video acquisition unit 351 acquires content data. As a result, the playback unit 32 plays back video that has not been processed (cut out).

[0046] [Operation and Effect of Terminal Device 3A] In the terminal device 3A described above, the information generation unit 232 generates position information and size information, and the video generation unit 233 generates a tracking video for each of the multiple tracking targets T based on the position information and size information. Then, the playback unit 32 (output unit) displays the tracking video corresponding to the selection information acquired by the information acquisition unit 234. Therefore, with the terminal device 3A or the content playback system 100A including the terminal device 3A, even if multiple viewing targets W are shown in the original video, the viewer can view only the tracking target T.

[0047] <Embodiment of Content Processing Method> Next, another embodiment of the present invention will be described below. For convenience of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.

[0048] [Flow of Content Processing Method S100] As shown in FIG. 6, the content processing method S100 includes an information generating step S1, a moving image generating step S2, an information acquiring step S3, and an outputting step S4.

[0049] [Information Generation Step S1] In the initial information generation step S1, multiple viewing objects W are detected from a video. Then, position information indicating the position of each characteristic part of the multiple viewing objects W at multiple points in time in the video, and size information indicating the size of a partial region R including at least the characteristic part of the viewing object W, are generated. The detection of the viewing objects W and the generation of the position information and size information may be performed by the content processing device 2 according to the first embodiment or the terminal device 3A according to the second embodiment, or by another device. Furthermore, the device that detects the viewing objects W may be different from the device that generates the position information and size information.

[0050] [Video Generation Step S2] After generating the position information and size information, video generation step S2 is performed. In video generation step S2, partial regions R corresponding to multiple viewing targets W are respectively cut out from each frame corresponding to multiple time points in the video based on the position information and size information. Then, a tracking video is generated for each viewing target W, with the cut-out partial regions R as frames. The cutting out of the partial regions R and the generation of the tracking video may be performed by the content processing device 2 according to embodiment 1 or the terminal device 3A according to embodiment 2, or by another device. Furthermore, the device that cuts out the partial regions R may be different from the device that generates the tracking video.

[0051] [Information Acquisition Step S3] After generating the tracking video, information acquisition step S3 is performed. In information acquisition step S3, selection information is acquired that indicates the viewing target W selected as the tracking target T from among multiple viewing targets. The selection information may be acquired by the content processing device 2 according to embodiment 1 or the terminal device 3A according to embodiment 2, or by another device.

[0052] [Output Step S4] After receiving the selection of the tracking target T, output step S4 is performed. In output step S4, a tracking video showing the tracking target T is output from among a plurality of tracking videos based on the selection information. The tracking video may be output by the content processing device 2 according to the first embodiment or the terminal device 3A according to the second embodiment, or by another device.

[0053] [Effects of Content Processing Method S100] In the content processing method S100 described above, position information and size information are generated in information generation step S1, and a tracking video of each of multiple tracking targets T is generated based on the position information and size information in video generation step S2. Then, a tracking video corresponding to the selection information acquired in information acquisition step S3 is output in output step S4. Therefore, according to the content processing method S100, a tracking video showing only the tracking target T (partial region R) selected by the viewer can be displayed in the display area. As a result, even if multiple viewing targets W are shown in the original video, the viewer can view only the tracking target T.

[0054] <Modifications> The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.

[0055] For example, the content processing device 2 or the terminal device 3A may be configured so that the information acquisition unit 234 acquires the selection information before the video generation unit 233 generates the tracking video. Then, the video generation unit 233 may be configured to generate only the tracking video in which the tracking target T appears, based on the selection information.

[0056] Furthermore, the content processing device 2 or the terminal device 3A may be configured so that the information generation unit 232 generates the position information and size information, and the video generation unit 233 generates the tracking video alternately for each frame, at the same cycle as the frame rate of the video. This shortens the time lag between acquiring content data and playing the tracking video, making it possible to handle live video as well.

[0057] The content playback system 100 (content playback system 100A) may also include an authentication device that authenticates identification information. In this case, the terminal device 3, 3A may be configured to accept input of identification information and transmit it to the authentication device. The authentication device may then be configured to transmit connection information for accessing the content server 1 to the terminal device 3, 3A when playback of content on the terminal device 3, 3A is permitted.

[0058] The content server 1 may also be configured to encrypt content data and transmit the encrypted data to the content processing device 2 (terminal device 3A). In this case, the content playback system 100 (content playback system 100A) may include a key server that generates a decryption key for decrypting the encrypted data. The content processing device 2 (terminal device 3A) may also be configured to obtain a decryption key from the key server and decrypt the encrypted data using the decryption key.

[0059] The content processing program may be stored non-transitory on one or more computer-readable storage media. The storage media may or may not be included in the device. In the latter case, the content processing program may be supplied to the device via any wired or wireless transmission medium.

[0060] Furthermore, some or all of the functions of the control blocks can be realized by logic circuits. For example, an integrated circuit in which a logic circuit that functions as each of the control blocks is formed is also included in the scope of the present invention. In addition, the functions of the control blocks can also be realized by, for example, a quantum computer.

[0061] <Summary> A content processing device according to aspect 1 of the present invention comprises a video acquisition unit that acquires content data including a video; an information generation unit that detects each of a plurality of viewing objects from the video and generates position information that indicates the position of each characteristic part of the plurality of viewing objects at multiple time points in the video and size information that indicates the size of a partial region that includes at least the characteristic part of the viewing object; a video generation unit that cuts out each of the partial regions corresponding to the plurality of viewing objects from each frame of the video corresponding to multiple time points based on the position information and the size information, and generates a tracking video for each viewing object, the tracking video having the cut-out partial regions as frames; an information acquisition unit that acquires selection information that indicates which of the plurality of viewing objects has been selected as the tracking target; and an output unit that outputs the tracking video in which the tracking object appears from among the plurality of tracking videos based on the selection information.

[0062] A content processing device according to aspect 2 of the present invention may be configured in the above-described aspect 1 such that the information generation unit receives data of content showing the object of appreciation as input, and generates the position information and size information of the object of appreciation using a trained model that outputs the position information and size information of the object of appreciation.

[0063] A content processing device according to a third aspect of the present invention may be configured in the first or second aspect described above, where the object of appreciation is a person, and the characteristic part is a face of the person.

[0064] A content processing device according to aspect 4 of the present invention may be configured in any one of aspects 1 to 3 above, wherein the information generation unit generates size information such that the ratio of the size of the characteristic portion to the size of the partial region is within a predetermined range.

[0065] A content processing device according to aspect 5 of the present invention may be configured in the above-mentioned aspect 4 such that the video generation unit generates the tracking video in which the partial area is enlarged so as to match the display area of ​​the video.

[0066] A content processing device according to a sixth aspect of the present invention may be configured in any one of the first to fifth aspects above, wherein the content includes audio.

[0067] A content processing device according to aspect 7 of the present invention may be configured in the above-mentioned aspect 6 such that the object of appreciation is a person and the content data includes data obtained by filming a music event.

[0068] A content processing device according to aspect 8 of the present invention may be configured in the above-described aspect 7 such that the video generation unit performs processing to emphasize components of the audio corresponding to the voice of the person being tracked or the sound of an instrument played by that person, and the output unit outputs the tracking video together with the audio in which the components have been emphasized.

[0069] A content processing program according to a ninth aspect of the present invention causes a computer to execute a video acquisition process for acquiring content data including a video; a generation process for detecting each of a plurality of objects to be viewed from the video and generating position information indicating the position of each characteristic part of the plurality of objects to be viewed at multiple points in time in the video and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation process for cutting out the partial areas corresponding to the plurality of objects to be viewed from each frame corresponding to multiple points in time in the video based on the position information and the size information and generating a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; a result acquisition process for acquiring selection information indicating which of the plurality of objects to be viewed has been selected as the tracking object; and an output process for outputting the tracking video in which the tracking object appears from among the plurality of tracking videos based on the selection information.

[0070] A content processing method according to aspect 10 of the present invention includes a generation step in which a computer detects each of a plurality of objects to be viewed from a video included in the content, and generates position information indicating the position of each characteristic part of the plurality of objects to be viewed at multiple points in time in the video, and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation step in which the computer, based on the position information and the size information, cuts out the partial areas corresponding to the plurality of objects to be viewed from each frame of the video corresponding to multiple points in time, and generates a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; a result acquisition step in which the computer acquires selection information indicating which of the plurality of objects to be viewed has been selected as the tracking object; and an output step in which the computer, based on the selection information, outputs the tracking video in which the tracking object is shown from among the plurality of tracking videos.

[0071] 100, 100A Content reproduction system 1 Content server 2 Content processing device 21 Communication unit (output unit) 22 Storage unit 23 Calculation unit 231 Video acquisition unit 232 Information generation unit 233 Video generation unit 234, 234A Information acquisition unit 235, 235A Output processing unit 3, 3A Terminal device 31 Terminal communication unit 32 Reproduction unit (output unit) 33 Operation unit 34, 34A Terminal storage unit 35, 35A Terminal calculation unit 351 Terminal video acquisition unit (video acquisition unit) 352 Reception unit 353 Transmission control unit 354 Reproduction control unit S100 Content processing method S1 Information generation step S2 Video generation step S3 Information acquisition step S4 Output step

Claims

1. A content processing device comprising: a video acquisition unit that acquires data of content including video; an information generation unit that detects a plurality of objects to be viewed from the video and generates position information indicating the position of each characteristic part of the plurality of objects to be viewed at a plurality of time points in the video and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation unit that, based on the position information and the size information, cuts out the partial areas corresponding to the plurality of objects to be viewed from each frame corresponding to a plurality of time points in the video and generates a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; an information acquisition unit that acquires selection information indicating which of the plurality of objects to be viewed has been selected as a tracking target; and an output unit that, based on the selection information, outputs the tracking video in which the tracking object appears from among the plurality of tracking videos.

2. The content processing device according to claim 1, wherein the information generation unit receives data of content showing the object of appreciation as input and generates the position information and size information of the object of appreciation using a trained model that outputs the position information and size information of the object of appreciation.

3. The content processing device according to claim 1, wherein the object of appreciation is a person, and the characteristic part is the face of the person.

4. The content processing device according to claim 1, wherein the information generating section generates size information such that the ratio of the size of the characteristic portion to the size of the partial region falls within a predetermined range.

5. The content processing device according to claim 4, wherein the video generation unit generates the tracking video in which the partial region is enlarged so as to coincide with a display region of the video.

6. The content processing device according to claim 1, wherein the content includes audio.

7. The content processing device according to claim 6, wherein the object of appreciation is a person, and the content data includes data obtained by filming a music event.

8. A content processing device as described in claim 7, wherein the video generation unit performs processing to emphasize components of the audio that correspond to the voice of the person being tracked or the sound of an instrument played by that person, and the output unit outputs the tracking video together with the audio in which the components have been emphasized.

9. A content processing program that causes a computer to execute the following steps: a video acquisition process that acquires data of content including video; a generation process that detects each of a plurality of objects to be viewed from the video and generates position information indicating the position of each characteristic part of the plurality of objects to be viewed at a plurality of time points in the video and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation process that, based on the position information and the size information, cuts out each of the partial areas corresponding to the plurality of objects to be viewed from each frame corresponding to a plurality of time points in the video and generates a tracking video for each object to be viewed, the tracking video having the cut-out plurality of partial areas as frames; a result acquisition process that acquires selection information indicating which of the plurality of objects to be viewed has been selected as the tracking target; and an output process that, based on the selection information, outputs the tracking video in which the tracking target is shown from among the plurality of tracking videos.

10. A content processing method comprising: a generation step in which a computer detects each of a plurality of objects to be viewed from a video included in content, and generates position information indicating the position of each characteristic part of the plurality of objects to be viewed at a plurality of time points in the video, and size information indicating the size of a partial area including at least the characteristic part of the object to be viewed; a video generation step in which the computer cuts out each of the partial areas corresponding to the plurality of objects to be viewed from each frame corresponding to a plurality of time points in the video based on the position information and the size information, and generates a tracking video for each object to be viewed, the tracking video having the cut-out partial areas as frames; a result acquisition step in which the computer acquires selection information indicating an object selected as a tracking object from among the plurality of objects to be viewed; and an output step in which the computer outputs, from among the plurality of tracking videos, the tracking video in which the tracking object is shown based on the selection information.

Citation Information

Patent Citations

  • Communication device, communication system, communication method and computer program

    JP2017139628A

  • Data processing apparatus, data processing method and program

    JP2019220848A

  • Video distribution system and video distribution method

    JP2022086438A

  • Image processing device, image processing system, image processing method, program, and display device

    JP2023142787A

  • Video retargeting

    US20090251594A1