Method and system for outputting generated media
The method and system address limitations in generating matching audio tracks and creating new audio-video content by using AI to process and combine network audio-video information, enabling efficient and versatile media generation for various applications.
Patent Information
- Application Number
- JP2025077982
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-19
- Filing Date
- 2025-05-08
- Publication Date
- 2026-01-07
AI Technical Summary
Existing technologies are limited in their ability to automatically generate matching audio tracks, perform video object identification, and create new audio-video content based on target drum audio tracks, lacking the capability to autonomously generate new matching footage or videos.
A method and system that utilize network audio-video information, perform track division processing, and combine artificial intelligence to generate and output media, involving an apparatus server, generation AI server, and sampling module to separate, align, and mix audio-video information on a time axis, using APIs to connect systems and enable user-friendly operations.
Enables automatic generation of media that matches input tracks, supports seamless integration of generated content with existing audio-video sources, and facilitates applications in entertainment, education, and creativity through AI-generated media output.
Smart Images

Figure 2026001696000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method and system for performing track splitting processing on network audio and video information, and combining it with artificial intelligence to generate generated information, and further output generated media. Hereinafter, in this specification, "audio and video" will be abbreviated to "audio-video". [Background technology]
[0002] In prior patent documents, for example, Patent Document 1 below discloses a music matching method. The device, electronic device, and computer are capable of reading storage media and mainly perform audio track separation for music awaiting matching to obtain target audio tracks for the music awaiting matching. The target audio tracks include a target person's voice audio track, a target accompaniment audio track, a target bass audio track, and a target drum audio track. Based on the target drum audio track, a first target drum order and a second target drum order are obtained. The terminal selects at least one target piece of music from the candidate music, and defines a video awaiting matching that includes the target music as the target video. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Chinese Patent Application Publication No. 116939323A Summary of the Invention [Problem to be solved by the invention]
[0004] However, the above-mentioned Patent Document 1 has the following limitations. 1. Target music and target video matching is performed based on the target drum audio track. 2. It is not possible to automatically generate a generated audio track that matches the target audio track. 3. It can only separate multiple target audio tracks for music waiting to be matched, and cannot perform operations such as identifying objects in the video, detecting edges, and capturing the main character for audio-video sources that include video. 4. It is not possible to autonomously generate new matching footage, matching movies, or matching videos based on the selected target audio track.
[0005] Therefore, the inventors believed that the above drawbacks could be improved, and after extensive research, they came up with a method and system for acquiring network video information, performing track division processing, and combining it with artificial intelligence to output generated media, which effectively improves the above issues through rational design. [Means for solving the problem]
[0006] In order to achieve the above object, one aspect of the present invention is a method for acquiring network video information, performing track division processing, combining artificial intelligence, and outputting generated media, the method comprising the steps of: an apparatus server of an audio-video playback device, or a generation artificial intelligence server connected to the apparatus server by signal, acquiring network audio-video information from a streaming platform; the apparatus server or the generation artificial intelligence server separating multiple division track information from the network audio-video information using a sampling module; the generation artificial intelligence server selecting some or all of the division track information and positioning the selected division track information to align it on a time axis; the generation artificial intelligence server generating at least one generation information based on the selected division track information and positioning the generation information to align it with the division track information on the time axis; the generation artificial intelligence server mixing the selected division track information and the generation information, outputting it as generated media, and feeding it back to the apparatus server; and the apparatus server playing the generated media in the playback interface of the audio-video playback device.
[0007] To achieve the above-mentioned object, another aspect of the present invention provides a system for acquiring network video information, performing track segmentation processing, and outputting generated media by combining artificial intelligence, the system being connected to a streaming platform. The system includes an audio / video playback device having an apparatus server and a playback interface connected to the apparatus server, a sampling module connected to the apparatus server, and a generation artificial intelligence server connected to the apparatus server and the sampling module. The apparatus server or the generation artificial intelligence server acquires network audio / video information from the streaming platform, and then separates multiple segment track information from the network audio / video information using the sampling module. The generation artificial intelligence server selects some or all of the segment track information and positions the selected segment track information on a time axis. The generation artificial intelligence server generates at least one generated information based on the selected segment track information and positions the generated information on the time axis to align it with the segment track information. The artificial intelligence server mixes the selected split track information and the generated information, outputs the mixed information as generated media, and feeds it back to the device server, which plays the generated media in the playback interface. [Effects of the Invention]
[0008] The present invention is configured as described above and therefore provides the following effects. 1. Application Programming Interface (API) technology is used to connect different systems or application programs, and the API can be modularized and set up according to different capture targets, making operation smoother and maintenance easier. 2. The artificial intelligence generating server uses artificial intelligence to automatically generate one or more pieces of generated information that match on the timeline based on the input segment track information, and all generated information generated by the artificial intelligence generating server is a selectable independent item. The artificial intelligence generating server superimposes the selected generated information onto the input segment track information to quickly generate a completely new generated media, making it widely applicable for entertainment, education, creativity, etc.
[0009] At least the following points become clear from the description and drawings described below. Note that "YouTube," "Netflix," "Spotify," and "KKBOX" in Figures 3 and 4 are all registered trademarks. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram illustrating a system for acquiring network video information, performing track division processing, and outputting generated media in combination with artificial intelligence according to an embodiment of the present invention. [Figure 2] 1 is a flowchart illustrating a method and system for acquiring network video information, performing track splitting processing, and combining artificial intelligence to output generated media according to an embodiment of the present invention. [Figure 3] 1 is a schematic diagram (1) of a search and sort interface according to an embodiment of the present invention, in which a single search mode and an integrated search mode are provided; [Figure 4] FIG. 2 is a second schematic diagram of a search sort interface displaying search results according to an embodiment of the present invention; [Figure 5]FIG. 3 is a schematic diagram (III) of an embodiment of the present invention in which audio-only network audio-video information is divided into split tracks. [Figure 6] FIG. 4 is a schematic diagram (4) of a case where generated information is mixed based on audio-only network audio-video information according to an embodiment of the present invention and output as generated media. [Figure 7] 5 is a schematic diagram (V) of a case where network audio-video information including audio and video is divided into tracks according to an embodiment of the present invention; [Figure 8] 6 is a schematic diagram (VI) of a case where generated information is mixed based on network audio-video information including audio and video according to an embodiment of the present invention and output as generated media. [Figure 9] 7 is a schematic diagram (7) of a case where generated information is mixed based on network audio-video information including audio and video according to an embodiment of the present invention and output as other generated media. DETAILED DESCRIPTION OF THE INVENTION
[0011] The following describes in detail the embodiments of the present invention, but the present invention is not limited to these, and various modifications are possible within the scope of the description, and embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.
[0012] FIG. 1 shows a system (hereinafter referred to as the present system) for acquiring network video information according to the present invention, performing track division processing, and combining it with artificial intelligence to output generated media, which is used to execute a method (hereinafter referred to as the present method) for acquiring network video information according to the present invention, performing track division processing, and combining it with artificial intelligence to output generated media.
[0013] This system is connected to multiple streaming platforms A, such as YouTube (registered trademark) and Netflix (registered trademark) for playing audio and video, and Spotify (registered trademark) and KKBOX (registered trademark) for playing music. Each streaming platform A has an API server A1.
[0014] This system includes the following components, each of which will be explained below.
[0015] <Audio / Video Playback Device 1> The system includes a device server 11, a playback interface 12 whose signals are connected to the device server 11, a search and sort interface 13, an open interface 14, and a database 15.
[0016] The device server 11 is connected to all API servers A1, and may have multiple open interfaces 14 corresponding to the streaming platforms A according to the number of streaming platforms A. The database 15 may be a physical hard disk, a storage area network, a cloud hard disk, etc.
[0017] <Sampling Module 2> The sampling module 2 is connected to a device server 11. Preferably, the sampling module 2 includes a forward propagation neural network, an item identification model, and an edge detection algorithm. The forward propagation neural network is one of a multi-source separation model (UMX), a voice frequency separation network (TasNet), a speech separation model (Looking to Listen), a voice separation model (Demucs), and a dual path transform network (DPTNet). The item identification model is one of YOLO8, Vision Master, Vision AI, and Amazon Rekognition®. The edge detection algorithm is one of a Canny edge detection algorithm, a Laplacian edge detection algorithm, and a Sobel edge detection algorithm.
[0018] <Generated AI Server 3> The signal is connected to the device server 11 via the sampling module 2 .
[0019] With reference to Figures 1 and 2, the method comprises the following steps: First, the system is installed. For example, an audio-video playback device 1 is purchased and connected to the cloud sampling module 2 and the generation AI server 3. Alternatively, the user can load the application program deployed by the system onto a playback device or PC they already own, and use the computing unit of the playback device or PC itself as the device server 11 to execute the method.
[0020] Please refer to Figures 3 and 4. Please also refer to Figures 1 and 2. The user searches for network audio-video information B, such as audio-video, music, etc., to be played through the single search mode or integrated search mode of the search and sort interface 13 (see Figure 5 for the network audio-video information B).
[0021] In the single search mode, the user selects search options supported by different streaming platforms A shown in FIG. 3 based on the music without video or audio video with video that the user wants to play.
[0022] After selecting a single search mode of any one of the streaming platforms A, the user then inputs the keywords he / she wants to search, and the device server 11 sends an audio / video search request to the selected API server A1 through application programming interface technology (API) based on the keywords.
[0023] The API server A1 executes a search based on the audio / video search request and obtains search results corresponding to the keywords. The device server 11 obtains multiple link codes of the search results from the API server A1 through the API, each link code corresponding to a different network audio / video information B. The search sorting interface 13 first excludes link codes corresponding to network audio / video information B that have invalid status, and then arranges the link codes in the search sorting interface 13 to display the search results.
[0024] Taking YouTube (registered trademark) as an example, after YouTube (registered trademark) receives an audio / video search request, YouTube (registered trademark)'s API server A1 performs a search using YouTube (registered trademark)'s own search engine, and the obtained search results include multiple audio / videos.
[0025] The present invention links information such as the address, name, creator, etc. of these audio / videos as link codes, and eliminates link codes that are invalid due to copyright issues, playback rights issues, or violations of streaming platform A's policies, etc., and displays complete YouTube® search results using the search and sort interface 13 without opening the YouTube® web page or YouTube® application program screen. Playback is not interrupted due to invalid links, and the user's audio / video viewing experience is not affected.
[0026] In the integrated search mode, for example, as shown in FIG. 3, in the integrated search option, when the user inputs the keyword he / she wants to search, the device server 11 sends an audio / video search request to the API servers A1 of all streaming platforms A, and each API server A1 independently searches and obtains its own search results.
[0027] The device server 11 integrates multiple network audio-video information B with the same name or creator from different streaming platforms A into one data, or integrates multiple network audio-video information B with the same creator but not exactly the same name but the same keyword into one data, and indicates in the search sorting interface 13 that the data comes from different streaming platforms A, and displays the search results of all streaming platforms A in the single search sorting interface 13 and allows the user to select (see Figure 4).
[0028] Users can add the link code of a specific source to the playlist by clicking the corresponding source symbol. Users can also delete or rearrange the link codes through the playlist options in Figure 3. This is an intuitive operation method that users can easily get used to.
[0029] The device server 11 or the generation artificial intelligence server 3 sequentially obtains link codes from the playlist, analyzes the link codes through API, and then obtains the network audio-video information B from the API server A1. Similarly, there is no need to jump to the web page of another streaming platform A, use an application program, or download the network audio-video information B.
[0030] Directly downloading network audio-video information B from streaming platform A without adopting API is also a possible embodiment of the present invention.
[0031] Preferably, the device server 11 or the generating artificial intelligence server 3 first verifies whether the streaming platform A is open, for example, whether the user is logged in, whether the paid mode is open, etc.
[0032] If the streaming platform A is not open, the device server 11 displays an open interface 14 corresponding to the streaming platform A in the audio-video playback device 1, opens it to the user, and then obtains the network audio-video information B, and makes the source of the network audio-video information B normal.
[0033] Next, the forward propagation neural network of the sampling module 2 separates a plurality of divided track information C1, C2, and C3 from the network audio / video information B based on frequency characteristics for sampling (see FIG. 5). The divided track information C1, C2, and C3 includes audio data of the divided track, video of the divided track, video of the main character, lyrics, white text, narration, etc. The video of the main character is preferably obtained by capturing using an item identification model and an edge detection algorithm.
[0034] The generation artificial intelligence server 3 randomly or according to a user's command selects some or all of the segment track information C1, C2, C3, and aligns the selected segment track information C1, C2, C3 on the time axis. For example, all segment track information C1, C2, C3 start from the same playback time of 0 seconds during the sampling process, and mark the playback time start point as the start point of each segment track information C1, C2, C3 on the time axis. The start points of the selected segment track information C1, C2, C3 are aligned to complete the positioning of the segment track information C1, C2, C3.
[0035] The artificial intelligence server 3 generates at least one piece of generated information D1, D2, D3 based on the selected divided track information C1, C2, C3 (see FIG. 6), and positions the generated information D1, D2, D3 on the time axis so as to align it with the divided track information C1, C2, C3. The generated information D1, D2, D3 includes generated audio data, static generated video, dynamically continuous generated video, generated text, etc.
[0036] The artificial intelligence generating server 3 automatically or based on the user's command performs additional changes such as changes in style, rhythm, melody, volume, and tone on the selected partial divided track information C1, C2, C3, to meet the different needs of each user for a wide range of purposes such as entertainment, education, and creation.
[0037] Finally, the generation artificial intelligence server 3 mixes the selected split track information C1, C2, C3 and the generation information D1, D2, D3, outputs it as generated media E, and also feeds it back to the device server 11 (see Figure 6), so that the device server 11 plays the generated media E in the playback interface 12 for the user to view.
[0038] Please refer to Figures 5 and 6. Please also refer to Figures 1 and 2. In one embodiment, network audio video information B contains only voice. After track division, divided track information C1, C2, and C3 contain only audio data of the divided tracks, such as drum sounds, violin sounds, and piano sounds. The artificial intelligence generation server 3 can generate generated information D1, D2, and D3, such as generated images, videos, and text, such as static seaside buildings and continuously changing sailing ships. The artificial intelligence generation server 3 mixes the divided track information C1, C2, and C3 and the generated information D1, D2, and D3, and outputs them as generated media E.
[0039] Please refer to Figures 7 to 9. Please also refer to Figures 1 and 2. In another embodiment, the network audio-video information B' simultaneously contains audio and video, and the divided track information C1', C2', C3', and C4' after track division further contains, in addition to the audio data of the divided tracks, for example, the video of the divided track of the main video. The generation artificial intelligence server 3 can generate the generated information D1' and D2', and mix the divided track information C1', C2', C3', and C4' with the generated information D1' and D2' to output the generated media E1'.
[0040] Only the split track information C1', C2', and C3' that constitute the audio data of the split tracks is selected, and mixed with generated information D3', D4', and D5' such as generated video, generated animation, and generated text, and output as different generated media E2' (see Figure 9).
[0041] Alternatively, only the audio data of a partial division track and the video of a partial division track may be selected and combined with the generated text, or only the audio data of a partial division track may be selected and combined with the generated audio data and generated video, or the sailing ship in the main video may be removed and the waves in the video of the division track may be selected and combined with the generated audio data, etc. The selection and combination methods are not limited.
[0042] Returning to Figure 6, please also refer to Figures 1 and 2. The present invention has many more applications. For example, suppose a user is a violinist and splits a music band ensemble on Spotify (registered trademark) into an audio track for violin, an audio track for piano, an audio track for trumpet, etc., then removes the audio track for violin, and combines the remaining split track information C1, C2, and C3 with sheet music or a favorite generated video, which is output as generated media E. The user can play generated media E and practice the violin at the same time to further master the technique of the song.
[0043] Alternatively, a user can upload a recording of a company meeting to YouTube (registered trademark), separate the voices of different colleagues into separate tracks, and amplify the voices of colleagues with quieter voices. The artificial intelligence generation server 3 then combines the speech-identified dialogue and translation, and outputs the resulting media E as a multilingual meeting transcript.
[0044] For link codes whose number of acquisitions reaches a threshold, the device server 11 further stores these link codes in the database 15, so that the next time the user plays them, the device server 11 or the generation artificial intelligence server 3 can quickly acquire them. Furthermore, the device server 11 collects statistics on link codes frequently used by users of different regions, ages, and genders, and makes customized recommendations based on the user's region, age, gender, etc.
[0045] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. [Explanation of symbols]
[0046] 1. Audio-video playback device 11 Device Server 12 Playback Interface 13 Search and sort interface 14 Open Interface 15 Databases 2 Sampling Module 3. Generating AI Server A streaming platform A1 API Server B Network Audio Video Information B' Network Audio Video Information C1 Split track information C2 Split Track Information C3 Split Track Information C1' Split track information C2' Split track information C3' Split Track Information C4' Split Track Information D1 generation information D2 generation information D3 Generation Information D1' Generation information D2' generation information D3' Generation information D4' Generation information D5' Generation information E-Generated Media E1' Generative Media E2' Generation Media
Claims
1. A method for acquiring network video information, performing track splitting processing, and combining artificial intelligence to output generated media, comprising: An apparatus server of an audio-video playback apparatus or a generating artificial intelligence server connected to the apparatus server receives network audio-video information from a streaming platform, and the apparatus server or the generating artificial intelligence server separates a plurality of split track information from the network audio-video information through a sampling module; The generation artificial intelligence server selects a part or all of the division track information and positions the selected division track information so as to align it on a time axis, and the generation artificial intelligence server generates at least one generation information based on the selected division track information and positions the generation information so as to align it with the division track information on the time axis; The generating artificial intelligence server mixes the selected divided track information and the generated information, outputs the mixed information as the generated media, and feeds it back to the device server; and the device server plays the generated media in a playback interface of the audio-video playback device.
2. 2. The method of claim 1, wherein the audio-video playback device has a search and sort interface connected to the device server, the device server sending an audio-video search request to an API server of the streaming platform based on a keyword through an application programming interface (API). The API server performs a search based on the audio-video search request and obtains search results corresponding to the keyword. Then, the device server obtains link codes of the search results from the API server through the application programming interface (API). Each link code corresponds to a different piece of network audio-video information. The device server obtains one of the link codes. The device server or the generation artificial intelligence server analyzes the link code to obtain the corresponding piece of network audio-video information from the API server. The search and sort interface arranges the link codes from the audio-video playback device to display the search results.
3. 3. The method for outputting generated media according to claim 2, wherein the audio / video playback device has a database whose signal is connected to the device server, and after the search results are displayed on the search and sort interface, the user selects at least one of the link codes, and adds the selected link code to a playlist through the search and sort interface, and the device server or the generation artificial intelligence server sequentially acquires the link codes from the playlist, and when the number of times one of the link codes has been acquired reaches a threshold, the device server stores the one of the link codes in the database.
4. 3. The method of claim 2, wherein the streaming platform includes a plurality of API servers, each of which is connected to the device server via a respective signal; the search and sort interface includes a single search mode and an integrated search mode; when the search and sort interface is operated in the single search mode, the device server sends the audio-video search request to the API server of one of the streaming platforms and displays the search results of the one of the streaming platforms in the audio-video playback device; and when the search and sort interface is operated in the integrated search mode, the device server sends the audio-video search request to the API servers of all of the streaming platforms and displays the search results of all of the streaming platforms in the audio-video playback device; and the device server comprehensively displays the link codes corresponding to the same network audio-video information from different streaming platforms.
5. 3. The generated media output method of claim 2, wherein before displaying the search results on the search sorting interface, the device server first excludes the link code corresponding to the network audio-video information having an invalid status, and before the device server or the generating artificial intelligence server obtains the network audio-video information from the streaming platform, the device server first verifies whether the streaming platform is open.
6. The generated media output method of claim 1, characterized in that before the generating artificial intelligence server outputs the generated media, it first performs additional changes to at least a selected portion of the divided track information, including changes to one of musical style, rhythm, melody, volume, and tone.
7. The generated media output method of claim 1, wherein the sampling module includes a forward propagation neural network, samples the divided track information from the network audio-video information based on frequency characteristics, and captures the main image from the network audio-video information using an item identification model and an edge detection algorithm.
8. A system that acquires network video information whose signal is connected to a streaming platform, performs track division processing, and combines artificial intelligence to output generated media, an audio-video playback device having a device server and a playback interface signal-connected to the device server; a sampling module signally connected to the device server; a generating artificial intelligence server signal-connected to the device server and the sampling module; The device server or the generating artificial intelligence server obtains network audio-video information from the streaming platform, and then separates a plurality of split track information from the network audio-video information through the sampling module; The artificial intelligence generation server selects a part or all of the divided track information and positions the selected divided track information on the time axis to align it with the divided track information. The artificial intelligence generation server generates at least one piece of generated information based on the selected divided track information and positions the generated information on the time axis to align it with the divided track information. The artificial intelligence server for generating the music data mixes the selected divided track information and the generated information, outputs the mixed information as the generated media, and also feeds the mixed information back to the device server. The device server plays the generated media in the playback interface.
9. 10. The generated media output system of claim 8, wherein the streaming platform includes an API server, the device server is connected to the API server by a signal, and the audio-video playback device includes a search and sort interface connected to the device server by a signal, the device server sends an audio-video search request to the API server based on a keyword through an application programming interface technology, the API server performs a search and obtains search results corresponding to the keyword, and then the device server obtains from the API server by the application programming interface technology a plurality of link codes corresponding to the search results, each link code corresponding to a different piece of network audio-video information, the device server obtains one of the link codes, the device server or the generation artificial intelligence server analyzes the one of the link codes through the application programming interface technology, and then obtains the corresponding network audio-video information from the API server based on the one of the link codes through the application programming interface technology, and the search and sort interface arranges the link codes in the audio-video playback device and displays the search results.
10. 10. The generated media output system of claim 9, wherein the audio / video playback device has a database signal-connected to the device server, and after displaying the search results on the search / sort interface, a user selects at least one of the link codes, and adds the selected link code to a playlist through the search / sort interface; the device server or the generating artificial intelligence server sequentially acquires the link codes from the playlist; and when the number of times one of the link codes has been acquired reaches a threshold, the device server stores the one of the link codes in the database.
11. 10. The generated media output system of claim 9, further comprising: a plurality of streaming platforms, each of which has its own API server connected to the device server; the search and sort interface having a single search mode and an integrated search mode, wherein when the search and sort interface is operated in the single search mode, the device server sends the audio-video search request to the API server of one of the streaming platforms and displays the search results of the one of the streaming platforms in the audio-video playback device; and when the search and sort interface is operated in the integrated search mode, the device server sends the audio-video search request to the API servers of all of the streaming platforms and displays the search results of all of the streaming platforms in the audio-video playback device, and the device server comprehensively displays the link codes corresponding to the same network audio-video information from different streaming platforms.
12. 10. The generated media output system of claim 9, wherein before displaying the search results on the search sorting interface, the device server first eliminates the link code corresponding to the network audio-video information having an invalid status, and before the device server or the generating artificial intelligence server obtains the network audio-video information from the streaming platform, the device server first verifies whether the streaming platform is open.
13. The generated media output system of claim 8, characterized in that before the generating artificial intelligence server outputs the generated media, it first performs additional changes to at least a selected portion of the divided track information, including changes to one of style, rhythm, melody, volume, and tone.
14. The generated media output system of claim 8, wherein the sampling module includes a forward propagation neural network, samples the segmented track information from the network audio-video information based on frequency characteristics, and captures the main image from the network audio-video information using an item identification model and an edge detection algorithm.
Citation Information
Patent Citations
Receiving device, query generation method and program
JP2014160503A
Systems and methods for immersive audio experiences
JP2024522251A
Unified playlist
US20160241922A1
Music streaming, playlist creation and streaming architecture
US20230185846A1
Music matching method and device, electronic equipment and computer readable storage medium
CN116939323A