Information processing device and program
The information processing apparatus efficiently generates and distributes time-saving content by performing sound source separation using DNN models, addressing generation time and distribution restrictions in existing systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- TOSHIBA VISUAL SOLUTIONS CORPORATION
- Filing Date
- 2024-11-20
- Publication Date
- 2026-06-01
AI Technical Summary
Existing systems take time to generate short-time content and may face distribution restrictions due to content rights issues when creating content from multiple user views.
An information processing apparatus that performs sound source separation on audio data using a DNN model to generate time-saving content with a shortened playback time, allowing efficient generation and distribution of playlists based on content genre and service information.
Enables efficient generation and playback of time-saving content, reducing generation time and overcoming distribution limitations by separating audio data into sound sources for personalized and genre-specific playlists.
Smart Images

Figure 2026089140000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to an information processing apparatus and a program.
Background Art
[0002] Regarding distributed content such as videos, it is known that short-time content with a shortened playback time of the content can be created by editing the content based on information aggregated externally.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, when generating short-time content using information given by a plurality of users who have viewed the content after the distribution of the content, as in Patent Document 1, it may take time to generate the short-time content. In addition, due to the rights relationship of the content, there may be cases where the content cannot be distributed to a plurality of viewers.
[0005] The problem to be solved by the present invention is to provide an information processing apparatus and a program capable of efficiently generating short-time content.
Means for Solving the Problems
[0006] The information processing apparatus of this embodiment is an information processing apparatus that receives and processes content data, and comprises: an acquisition unit that stores audio data obtained from the content data in a storage area; an inference unit that performs sound source separation processing using the audio data stored in the storage area and outputs sound source data; and a generation unit that generates time-saving content with a shortened playback time of the content data based on the output sound source data. [Brief explanation of the drawing]
[0007] [Figure 1] Figure 1 shows an example of the overall configuration including the information processing device according to the embodiment. [Figure 2] Figure 2 is a diagram illustrating the sound source separation and playlist according to this embodiment. [Figure 3] Figure 3 is a schematic diagram showing the generated playlist according to the embodiment. [Figure 4] Figure 4 is a diagram illustrating the gain adjustment of sound source data according to this embodiment. [Figure 5] Figure 5 is a flowchart showing an example of the process for generating a playlist according to this embodiment. [Figure 6] Figure 6 is a flowchart showing an example of a gain adjustment process according to the embodiment. [Modes for carrying out the invention]
[0008] The information processing device and program of this embodiment will be described below with reference to the attached drawings.
[0009] Figure 1 shows an example of the overall configuration including the information processing device 1 according to the embodiment. As shown in Figure 1, the information processing device 1 comprises a video stream analysis unit 100, an AI (artificial intelligence) sound source separation processing unit 200, a playback control unit 300, and a UI (User Interface) setting unit 500. The information processing device 1 is also connected to a content DB 400.
[0010] In this embodiment, the information processing device 1 will be described as, for example, a device built into an edge device such as a television system.
[0011] The video stream analysis unit 100 decodes and analyzes video data, audio data, service information (SI), subtitle data, etc., contained in content data received using, for example, a receiver that receives broadcast waves (not shown) or a communication device (not shown). The content data may be broadcast waves.
[0012] Here, video data refers to video data that includes both video and audio data, such as sports or movies. Audio data refers to audio data that is synchronized with video data, such as cheers, dialogue, or instrument sounds. Service information refers to program scheduling information, such as the start time of the content or the genre of the content. Subtitle data refers to data that displays cheers, dialogue, etc., included in the audio data as text.
[0013] The video stream analysis unit 100 comprises a video decoder 101, an audio decoder 102, and a metadata analysis unit 103.
[0014] Specifically, the video decoder 101 decrypts the video data contained in the received content data, for example, if the received content data is encrypted.
[0015] Specifically, the audio decoder 102 decrypts the audio data contained in the received content data if the content data is encrypted. The audio decoder 102 sends the decrypted audio data to the DNN (Deep Neural Network) input buffer management unit 201. If the received content data is not encrypted, the audio data is sent to the DNN input buffer management unit 201 without decryption.
[0016] The metadata analysis unit 103 acquires, for example, service information included in content data, and identifies the genre of the content or the like using the acquired service information.
[0017] The metadata analysis unit 103 sends, for example, the identified genre of the content or the like to the sound source separation DNN inference unit 202. In addition, the metadata analysis unit 103 sends the identified genre of the content or the like to the sound source separation result analysis unit 204. The genre of the content or the like is, for example, sports, movies, dramas, music, and the like.
[0018] Next, the sound source separation AI processing unit 200 performs sound source separation processing on the audio data transmitted from the video stream analysis unit 100 and separates it into a plurality of sound source data. In addition, the sound source separation AI processing unit 200 uses information such as the sound pressure of the separated sound source data to generate a time-shortened content (hereinafter referred to as a playlist) that can shorten the playback time of the content data.
[0019] The sound source separation AI processing unit 200 includes a DNN input buffer management unit 201, a sound source separation DNN inference unit 202, a sound source separation result cache 203, a sound source separation result analysis unit 204, a playlist DB (Data Base) 205, and a DNN model storage unit 206.
[0020] The DNN input buffer management unit 201 acquires audio data from the audio decoder 102. In addition, the acquired audio data is stored in a storage area provided in a DNN input buffer management unit 201 (not shown). Note that the DNN input buffer management unit 201 is also referred to as an acquisition unit.
[0021] The sound source separation DNN inference unit 202 acquires audio data from a storage area (not shown) provided in the DNN input buffer management unit 201. In addition, it acquires the genre of the content or the like from the metadata analysis unit 103. The sound source separation DNN inference unit 202 selects and acquires an inference model corresponding to the genre of the content from the DNN model storage unit 206 based on the acquired genre of the content or the like. Note that the inference model is also referred to as a DNN model.
[0022] The sound source separation DNN inference unit 202 generates a plurality of sound source data obtained by separating the acquired audio data for each predetermined sound source, using the selected inference model and the acquired audio data.
[0023] Also, the sound source separation DNN inference unit 202 outputs the separated plurality of sound source data to the sound source separation result cache 203. Note that the sound source separation DNN inference unit 202 is also referred to as an inference unit.
[0024] In order to ensure sufficient separation performance of the DNN inference by the DNN inference model, since it is necessary to apply a certain amount of audio data to the DNN inference model collectively, the sound source separation DNN inference unit 202 separates the audio data acquired in the storage area (not shown) of the DNN input buffer management unit 201 into at least one sound source data using the DNN inference model after a predetermined amount or more of the acquired audio data has been accumulated. Then, the at least one separated sound source data is output to the sound source separation result cache 203.
[0025] The acquired audio data being a predetermined amount or more is, for example, the amount of audio data when a predetermined time has elapsed since the start of distribution of the content data. Specifically, the amount of audio data when 8 seconds or more of time has elapsed since the start of reception is taken as an example.
[0026] Also, the sound source separation DNN inference unit 202 is executed by a CPU (Central Processing Unit) that executes predetermined arithmetic processing and control processing using parameters and the like stored in the DNN model storage unit 206 and the like.
[0027] The sound source separation result cache 203 stores the sound source data separated by the sound source separation DNN inference unit 202. Note that the separated plurality of sound source data is also stored after the playlist is generated.
[0028] The sound source separation result analysis unit 204 retrieves the sound source data stored in the sound source separation result cache 203. It also retrieves the content genre from the metadata analysis unit 103. Although not shown in the diagram, the sound source separation result analysis unit 204 may also retrieve an inference model from the DNN model storage unit 206 based on the retrieved content genre.
[0029] Furthermore, the sound source separation result analysis unit 204 generates a playlist based on, for example, at least one acquired sound source data and service information. For example, it generates a playlist based on the genre of content included in the content's service information. It then outputs the generated playlist to the playlist DB 205. The sound source separation result analysis unit 204 is also referred to as the generation unit.
[0030] A playlist, for example, is a time-saving content that shortens the playback time of content data. A playlist is information that identifies important scenes from content data, aligning them with the content data's timeline. Alternatively, it could be a time-saving content that shortens the length of the content data.
[0031] Specifically, a playlist, for example, if the content genre is sports, is information that identifies exciting scenes using multiple audio data separated from the audio data. Exciting scenes can be determined, for example, from the amplitude of audio data separated as cheers.
[0032] Specifically, the sound source separation result analysis unit 204 generates a playlist using the separated sound source data. It also generates a playlist based on, for example, at least one sound source data and the genre of content included in the service information.
[0033] The sound source separation result analysis unit 204 generates a playlist using the sound source data output by the sound source separation processing performed by the sound source separation DNN inference unit 202. Therefore, compared to generating a playlist using audio data or video data, it is possible to generate a playlist more efficiently within the edge device.
[0034] Furthermore, the sound source separation result analysis unit 204 can generate playlists according to genre by, for example, applying different rules for each genre of content. Also, since it separates audio data into sound source data at a predetermined time, it can generate playlists while content data is being distributed.
[0035] Furthermore, the sound source separation result analysis unit 204 is executed by a CPU that performs predetermined calculation and control processing using parameters stored in the DNN model storage unit 206, etc.
[0036] The playlist DB 205 stores the playlists transmitted from the sound source separation result analysis unit 204. The playlist DB 205 may store multiple playlists for each piece of content data.
[0037] The DNN model memory unit 206 stores, for example, at least one inference model. The DNN model memory unit 206 also stores, for example, multiple inference models based on service information. Specifically, for example, it stores multiple inference models for each genre of content included in the service information.
[0038] The DNN model memory unit 206 stores, for example, inference models for sports, movies, music, etc. The DNN model memory unit 206 is a non-volatile memory that includes, for example, Flash Memory, Hard Disk, ROM (Read Only Memory), etc.
[0039] Next, the playback control unit 300 is exemplified by the chase playback control unit. For example, it can be used when performing chase playback of content data that is being streamed. Here, chase playback is exemplified by recording the content data being streamed and then rewinding to any point between the start of recording the content data being streamed and the current recording position of the content data.
[0040] The playback control unit 300, for example, when performing catch-up playback, can play content data in a shorter time than normal by using a playlist in which the content data has been shortened to specific scenes, making it possible to catch up with the content data being streamed. Furthermore, even when not performing catch-up playback, that is, when playing recorded content after it has been streamed, the playback can be performed in a shorter time than the playback time of the recorded content by using a playlist.
[0041] The playback control unit 300 includes a video decoder 301, an audio decoder 302, a trick play control unit 303, and an audio limiter 304.
[0042] The video decoder 301 acquires video data from, for example, recorded content in the content database 400. Specifically, for example, it acquires a playlist from the trick play control unit 303 and outputs video data according to the playlist. Also, for example, when playing a playlist, it outputs video data for each period specified by the playlist from the acquired recorded content.
[0043] The audio decoder 302 retrieves audio source data from, for example, the audio source separation result cache 203. It also retrieves a playlist from the trick play control unit 303 and outputs the separated audio source data according to the playlist.
[0044] The audio decoder 302 can also acquire and play audio data from the recorded content in the content DB 400. Furthermore, it plays back audio data whose gain has been adjusted by the audio limiter 304.
[0045] The audio decoder 302 can output only specific audio data because the audio source data is stored in the audio source separation result cache 203, even when performing normal playback based on the content DB 400, as well as when outputting based on a playlist. Furthermore, when playing a playlist, the audio decoder 302 outputs audio data for each period specified by the playlist from the acquired audio data.
[0046] The trick play control unit 303 reads a playlist from the playlist DB 205. Based on the playlist, it controls the video decoder 301 and audio decoder 302 to play only the important scenes.
[0047] Furthermore, the trick play control unit 303 specifically controls the playback of the playlist based on the sound source data output from the sound source separation DNN inference unit 202 and the video data of the content data. The trick play control unit 303 is also referred to as the control unit.
[0048] In this embodiment, the audio remixer 304 adjusts the gain of the audio data during playback by acquiring the audio source data separated from the audio source separation result cache 203.
[0049] Furthermore, the audio remixer 304, for example, when playing back content DB 400 or playlist DB 205, adjusts the gain of multiple audio source data included in the content data or playlist based on multiple audio source data output from the audio source separation DNN inference unit 202.
[0050] Furthermore, the audio limiter 304 specifically obtains information from the UI setting unit 500 regarding the genre of content selected by the user and the sound source data to be emphasized, and controls the audio decoder 302 to adjust the gain of multiple sound source data based on the obtained information.
[0051] The Content DB (Data Base) 400 stores various types of information, for example, using auxiliary storage devices that include non-volatile memory such as HDDs (Hard Disk Drives) and SSDs (Solid State Drives).
[0052] Furthermore, the content DB 400 is an external storage device separate from the information processing device 1, for example. It may also be included in the information processing device 1. The content DB 400 stores content data such as distributed recorded content.
[0053] The UI setting unit 500 is a unit that accepts user input. It may include input devices such as a keyboard or mouse. Furthermore, when playing content from the content DB 400, for example, the UI setting unit 500 sends information to the trick play control unit 303 regarding the genre of content selected by the user and the sound source data to be emphasized.
[0054] Next, we will explain sound source separation and playlist generation using Figure 2. Figure 2 is a diagram illustrating sound source separation and playlist generation according to this embodiment.
[0055] One example of sound source separation according to this embodiment is to output at least one sound source data from the audio data contained in the content data using an inference model.
[0056] As shown in Figure 2, for example, the sound source separation DNN inference unit 202 outputs sound source data using audio data acquired from content data and an inference model corresponding to the genre of content.
[0057] The sound source separation DNN inference unit 202 separates the audio data into categories such as "Dialogue," "Cheer," "Music," and "Effect" when the content genre is sports.
[0058] Furthermore, the audio source separation DNN inference unit 202 separates the audio data into categories such as "Dialogue" and "Other" when the content genre is movie or drama.
[0059] Furthermore, the sound source separation DNN inference unit 202 separates the content into sound source data such as "Drums," "Bass," "Other," and "Vocals" if the content genre is music. Note that the above separation example is just one example, and it is also possible to output at least one sound source data.
[0060] In this way, the sound source separation DNN inference unit 202 can output only one sound source data by separating it for each sound source data. Furthermore, gain adjustment can be performed for each sound source data.
[0061] Specifically, the sound source separation DNN inference unit 202 can, for example, output only "Vocals" when the content genre is music, or output everything except "Vocals". Furthermore, when the content genre is sports, it can reduce the amplitude of "Cheer" and increase the amplitude of "Dialogue". In other words, it can suppress the cheers and emphasize the announcer's voice.
[0062] Next, we will explain how to generate a playlist using audio data. The audio source separation result analysis unit 204 generates a playlist using audio data separated by content genre. For example, it generates a playlist using at least one audio data separated by content genre and an inference model for each content genre.
[0063] The sound source separation result analysis unit 204 may, for example, if the content genre is sports, consider sections of the sound source data where "Cheer" has a sound pressure level above a certain level as "exciting" scenes, and generate a playlist of extracted "exciting" scenes.
[0064] Furthermore, if the content genre is a movie or drama, the sound source separation result analysis unit 204 may, for example, consider sections of "Dialogue" in the sound source data with a certain sound pressure level or higher as "scenes with dialogue," and generate a playlist of extracted "scenes with dialogue."
[0065] Furthermore, if the content genre is music, the sound source separation result analysis unit 204 may, for example, consider sections of the sound source data other than "Vocals" that have a certain sound pressure level or higher as "performance" scenes and generate a playlist of extracted "performance" scenes.
[0066] The extraction conditions described above for each content genre are merely examples; extraction conditions may be set in advance, or they may be changed each time. Furthermore, an inference model may be used to set appropriate extraction conditions according to the content genre.
[0067] Next, we will explain an example of content data and playlist playback using Figure 3. Figure 3 is a schematic diagram showing a playlist generated according to the embodiment. Figure 3 shows an example of content data to be distributed and a playlist generated from the content data.
[0068] The content data shown in Figure 3, for example, indicates the playback time of a single recorded piece of content.
[0069] Next, the playlist shown below the content data in Figure 3 is an example of data where specific periods (important scenes, performance scenes, etc.) are identified from the time axis of the content data. Specifically, as shown in the playlist in Figure 3, for example, specific periods (a), (b), (c), and (d) are identified from the time axis of the content data.
[0070] The playlist playback shown below the playlist in Figure 3 indicates the playback time for specific periods (a), (b), (c), and (d) of the playlist generated from the content.
[0071] As shown in Figure 3, playing a playlist allows for playback in a shorter time compared to the playback time of the content data. Furthermore, as shown in the playlist playback in Figure 3, the playlist may be pre-edited data containing only specific periods (such as important scenes or performance scenes).
[0072] Next, Figure 4 is a diagram illustrating the gain adjustment of sound source data according to the embodiment. The upper part of Figure 4 shows an example of playing back sound source data separated for each audio stored in the sound source separation result cache 203. The lower part of Figure 4 shows an example of playing back the separated sound source data after gain adjustment.
[0073] Here, the audio limiter 304 of the playback control unit 300 can, for example, adjust the gain for each separated audio source data when playing back audio source data stored in the audio source separation result cache 203. Furthermore, mixing using the separated audio source data becomes possible.
[0074] Here, we will explain using the example of the case where the sound source data is arranged from top to bottom as "Drums," "Bass," "Other," and "Vocals," as shown in the lower part of Figure 4.
[0075] The audio limiter 304 of the playback control unit 300 represents the case of original sound reproduction as an example from time 0 to t1. Specifically, it represents the case where the sound source data for "Drums", "Bass", "Other", and "Vocals" are all multiplied by a coefficient of 1 and added together.
[0076] Next, the audio limiter 304 of the playback control unit 300 will show the case of karaoke playback as an example from time t1 to t2. Specifically, it will show the case where all the sound source data for "Drums", "Bass", and "Other" (excluding "Vocals") are multiplied by a coefficient of 1 and added together.
[0077] Furthermore, the audio limiter 304 of the playback control unit 300 represents the time t2 onward as an example of playback with emphasis on vocals. Specifically, it represents the case where the "Vocals" sound source data is multiplied by a coefficient of 1, and the "Drums," "Bass," and "Other" sound source data are multiplied by a coefficient of 0.1 and then added together.
[0078] Furthermore, even during playback, if the playback mode is changed by, for example, the UI setting unit 500, the audio remixer 304 can perform gain adjustment. For example, gain adjustment can be performed at any timing, such as the timings t1 and t2 during playback.
[0079] Thus, since the separated audio data is stored in the audio source separation result cache 203, the audio remixer 304 can adjust the gain during playlist playback. Furthermore, even after the playlist has been generated, for example, since the audio data is stored in the audio source separation result cache 203, the gain can be adjusted even when the content is being played.
[0080] Next, Figure 5 is a flowchart showing an example of the process for generating a playlist according to the embodiment. As shown in Figure 5, the audio decoder 102 of the video stream analysis unit 100 receives content data (step S100).
[0081] Furthermore, the audio decoder 102 sends the audio data contained in the content data to the DNN input buffer management unit 201 of the sound source separation AI processing unit 200 (step S101).
[0082] Next, the DNN input buffer management unit 201 stores the acquired audio data in a memory area (step S102).
[0083] Next, the sound source separation DNN inference unit 202 acquires the audio data exceeding a predetermined amount when the audio data stored in the DNN input buffer management unit 201 exceeds a predetermined amount (step S103).
[0084] Here, the metadata analysis unit 103 of the video stream analysis unit 100 acquires service information contained in the received content data. Then, it identifies the content genre, etc., from the service information. The content genre, etc., is sent to the sound source separation DNN inference unit 202 (step S104).
[0085] The sound source separation DNN inference unit 202 selects and retrieves an inference model from the DNN model storage unit 206 based on the genre of the acquired content (step S105).
[0086] Furthermore, the sound source separation DNN inference unit 202 separates the acquired audio data into multiple sound source data using the selected inference model. Note that the sound source separation process may output at least one sound source data (step S106).
[0087] Furthermore, the sound source separation DNN inference unit 202 outputs the separated sound source data to the sound source separation result cache 203. The sound source separation result cache 203 then stores the separated sound source data (step S107).
[0088] Next, the sound source separation result analysis unit 204 acquires that the sound source data has been stored in the sound source separation result cache 203 (step S108).
[0089] Furthermore, the sound source separation result analysis unit 204 obtains service information from the metadata analysis unit 103. It also obtains the genre of content included in the service information (step S109).
[0090] Furthermore, the sound source separation result analysis unit 204 obtains the separated sound source data from the sound source separation result cache 203 (step S110).
[0091] Furthermore, the sound source separation result analysis unit 204 generates a playlist from the separated sound source data using an inference model based on the content genre (step S111).
[0092] The sound source separation result analysis unit 204 may, for example, select an inference model based on the content genre from the DNN model storage unit 206. Alternatively, the extraction conditions for playlists based on the content genre may be set in advance, and a playlist may be generated using the extraction conditions.
[0093] Furthermore, the sound source separation result analysis unit 204 sends the generated playlist to the playlist DB 205 (step S112).
[0094] The audio source separation AI processing unit 200 repeats the processes of separating audio data into individual audio source data and generating playlists until the content distribution is completed.
[0095] Next, the processing during playlist playback will be explained using Figure 6. Figure 6 is a flowchart showing an example of the gain adjustment process according to the embodiment. As shown in Figure 6, the trick play control unit 303 of the playback control unit 300 obtains information selected by the user regarding the playlist of content data from the UI setting unit 500 (step S200).
[0096] Furthermore, the trick play control unit 303 reads a playlist from the playlist DB 205 (step S201).
[0097] Furthermore, the trick play control unit 303 sets the input source based on the playlist to the video decoder 301 (step S202).
[0098] Furthermore, the trick play control unit 303 sets the input source based on the playlist to the audio decoder 302 (step S203).
[0099] Next, the video decoder 301 retrieves content data corresponding to the playlist from the content DB 400 (step S204).
[0100] Next, the audio decoder 302 retrieves the separated audio data corresponding to the playlist from the audio source separation result cache 203 (step S205).
[0101] Furthermore, the audio limiter 304 adjusts the gain of the audio data included in the playlist according to the information set by the user from the UI setting unit 500. Specifically, the trick play control unit 303 acquires the information from the UI setting unit 500 and adjusts the gain of the audio decoder 302 (step S206).
[0102] The video decoder 301 then outputs video data based on the playlist. The audio decoder 302 also outputs sound source data based on the playlist. (Step S207).
[0103] Furthermore, the audio remixer 304 may perform gain adjustments as appropriate while outputting data based on the playlist in step S207.
[0104] The trick play control unit 303 obtains the playback position of the playlist (step S208).
[0105] In step S208, for example, using chase playback as an example, if the playback position of the playlist catches up with the content data being streamed, the trick play control unit 303 terminates the playback of the playlist. Also, using content data playback as an example, if the playback position of the content data reaches its end, the playback of the playlist is terminated.
[0106] Furthermore, in step S208, for example, if the playback position of the playlist has not caught up with the content being streamed, the trick play control unit 303 may, for example, return to step S206 and output the playlist using the stored content data and sound source data via the video decoder 301 and audio decoder 302.
[0107] Furthermore, in step S208, for example, if the playback position of the playlist has not caught up with the content data being streamed, the trick play control unit 303 reads the playlist again, returns to step S204 to acquire the content data, and outputs the playlist using the video decoder 301 and audio decoder 302.
[0108] As described above, the information processing device of this embodiment is an information processing device that receives and processes content, and comprises: an acquisition unit that stores audio data acquired from content data in a storage area; an inference unit that performs sound source separation processing using the audio data stored in the storage area and outputs sound source data; and a generation unit that generates time-saving content with a shortened playback time of content data based on the output sound source data.
[0109] This allows for the efficient generation of time-saving content based on audio data. Furthermore, it reduces generation time compared to generating time-saving content using video or audio data. Additionally, generating time-saving content allows users to efficiently play content during catch-up playback.
[0110] Furthermore, the inference unit in the information processing device of this embodiment outputs sound source data according to service information related to content data. This makes it possible to generate different sound source data depending on the service information related to content data. It is also possible to output only specific sound source data.
[0111] Furthermore, the inference unit in this information processing device outputs sound source data using an inference model corresponding to the service information related to the content data. This allows for the generation of sound source data for each service information by using an inference model corresponding to the service information. In other words, multiple sound source data can be appropriately generated from audio data according to the genre of content.
[0112] Furthermore, in the information processing device of this embodiment, the generation unit generates time-saving content based on sound source data and service information. This allows for the generation of appropriate time-saving content based on the service information. In other words, if the content genre is sports, it can generate time-saving content that compiles exciting scenes with lots of cheering.
[0113] Furthermore, the information processing device of this embodiment further includes a control unit that plays time-saving content based on sound source data and content data. This allows the information processing device, for example, to play time-saving content using only specific sound source data from the sound source data via the control unit.
[0114] Furthermore, the information processing device of this embodiment is further equipped with an audio remixer that performs gain adjustment of the sound source data included in the content data or time-saving content based on the sound source data when playing back content data or time-saving content. This allows the user to play back the sound source data they want to hear by adjusting the gain of specific sound source data from the sound source data.
[0115] Furthermore, the information processing device of this embodiment includes an audio decoder that, when playing back time-saving content, outputs sound source data for specific periods from the sound source data acquired based on the time-saving content. As a result, since the time-saving content is played back using the sound source data output by the inference unit, gain adjustment can be performed during playback.
[0116] Furthermore, the information processing device of this embodiment includes a video decoder that, when playing back time-saving content, outputs video data for specific periods from the acquired content data based on the time-saving content. This allows content data stored in the content DB to be played back, eliminating the need to store large video data in the sound source separation AI processing unit. In addition, the video decoder can play back the video data.
[0117] The information processing device described in this embodiment is, for example, built into an edge device that does not transmit acquired content externally. The edge device is, for example, a television system comprising at least a video stream analysis unit 100, a sound source separation AI processing unit 200, and a playback control unit 300. The edge device may also be a Blu-ray recorder, a DVD (Digital Versatile Disk) recorder, etc. Furthermore, since the information processing device can generate time-saving content within the edge device, there is no need to send the content externally.
[0118] In this embodiment, we have described the process of separating audio sources and generating a playlist using received content data. However, the method is not limited to this; for example, the audio source data that has undergone source separation may be included in the broadcast wave to be distributed. Alternatively, the generated playlist may be distributed as a broadcast wave.
[0119] The program for realizing the functions of the information processing device 1 as described above may be provided as a file in a format installable or executable on a computer, recorded on a computer-readable recording medium such as a CD-ROM, flexible disk (FD), CD-R, or DVD (Digital Versatile Disk). Alternatively, the program may be stored on a computer connected to a network such as the Internet and provided by downloading it via the network. Furthermore, the program may be provided or distributed via a network such as the Internet.
[0120] While embodiments of the present invention have been described above, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. [Explanation of Symbols]
[0121] 1…Information Processing Device 100...Video Stream Analysis Department 103…Metadata Analysis Department 200...Sound source separation AI processing unit 201...DNN Input Buffer Management Unit (Acquisition Unit) 202…Sound source separation DNN inference unit (inference unit) 203... Audio source separation result cache 204...Sound source separation result analysis section (generation section) 205... Playlist DB 206...DNN Model Memory Unit 300...Regeneration Control Unit 301...Video Decoder 302…Audio Decoder 303... Trick Play Control Unit (Control Unit) 304... Audio Remixer 400... Content DB 500...UI settings section
Claims
1. An information processing device that receives and processes content data, An acquisition unit that stores audio data obtained from the aforementioned content data in a storage area, An inference unit that performs sound source separation processing using the audio data stored in the memory area and outputs sound source data, A generation unit generates time-saving content that shortens the playback time of the content data based on the output sound source data, An information processing device equipped with the following features.
2. The inference unit outputs the sound source data according to the service information relating to the content data. The information processing apparatus according to claim 1.
3. The inference unit outputs the sound source data using an inference model corresponding to the service information relating to the content data. The information processing apparatus according to claim 1 or 2.
4. The generation unit generates the time-saving content based on the sound source data and the service information. The information processing apparatus according to claim 2.
5. The system further includes a control unit that controls playback of the time-saving content based on the sound source data and the content data. The information processing apparatus according to claim 1.
6. The audio further comprises an audio remixer that performs gain adjustment of the audio data included in the content data or time-saving content based on the audio data when the content data or time-saving content is played back. The information processing apparatus according to claim 1.
7. The audio decoder further comprises, when playing the aforementioned time-saving content, an audio decoder that outputs audio data for specific periods from the acquired audio data based on the aforementioned time-saving content. The information processing apparatus according to claim 1.
8. The system further includes a video decoder that, when playing the aforementioned time-saving content, outputs video data for specific periods from the content data acquired based on the aforementioned time-saving content. The information processing apparatus according to claim 1 or 7.
9. A program for controlling an information processing device using a computer, The aforementioned computer An acquisition means for storing audio data obtained from content data in a memory area, An inference means that performs sound source separation processing using the audio data stored in the memory area and outputs sound source data, A generation means that generates time-saving content with a reduced playback time based on the output sound source data, A program that makes something work.