Broadcast television cataloguing method based on artificial intelligence large model and related products
By automatically identifying and splitting the timing, type and title of radio and television program content based on a method based on a large artificial intelligence model, the problem of high labor costs in the existing technology is solved, and a low-cost and efficient cataloging process is achieved.
Patent Information
- Application Number
- CN202510953285.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-03
AI Technical Summary
Existing intelligent cataloging systems require a large amount of manpower for template labeling and maintenance, resulting in high labor costs and low cataloging efficiency.
By adopting a method based on artificial intelligence big models, we obtain multimodal time series data, call artificial intelligence big models for cognition, automatically identify and split the timing, type and title of radio and television program content, and generate a catalog playlist.
There is no need for manual labeling templates, which improves cataloging efficiency, reduces labor costs, and realizes low-cost cataloging of radio and television.
Smart Images

Figure CN120751183A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of television cataloging, and in particular to a broadcast television cataloging method based on an artificial intelligence large model and related products. Background Art
[0002] Television cataloging is a crucial component of broadcasting and television operations. Existing intelligent cataloging systems primarily utilize audio and video template matching technology, which labels advertisements and programs on broadcast television with program names and types. After template labeling, the system automatically annotates the advertisements and program information through content consistency comparison, ultimately creating a playlist of broadcast and television programs. While intelligent cataloging has been achieved to a certain extent, it still requires significant manpower for template labeling and the daily maintenance of a large cataloging template library. Failure to maintain these templates in a timely manner significantly increases the workload of manual cataloging and annotation. Summary of the Invention
[0003] The present invention provides a radio and television cataloging method and related products based on an artificial intelligence large model, which are used to solve the defect of large manual input in the existing technology and realize low-cost cataloging of radio and television.
[0004] The present invention provides a radio and television cataloging method based on an artificial intelligence large model, comprising the following steps.
[0005] Obtain multimodal time series data of target broadcasting and television;
[0006] Invoke an artificial intelligence model to recognize the multimodal time series data and obtain the time series, type, and title of the program content contained in the target broadcast and television;
[0007] A cataloged playlist of the target broadcasting television is generated based on the timing, type and title of the program content contained in the target broadcasting television.
[0008] According to the radio and television cataloging method based on the artificial intelligence large model provided by the present invention, multimodal time series data of target radio and television are obtained, including:
[0009] Obtain audio and video timing data of the target broadcasting and television;
[0010] Convert the audio time series data into text time series data to obtain first text time series data;
[0011] Perform text recognition on the subtitles in the video timing data to obtain second text timing data; the first text timing data, the second text timing data and the video timing data constitute the multimodal timing data.
[0012] According to the radio and television cataloging method based on the artificial intelligence big model provided by the present invention, the artificial intelligence big model is called to recognize the multimodal time series data to obtain the time series, type and title of the program content contained in the target radio and television, including:
[0013] Invoking the artificial intelligence model to obtain time periods corresponding to different program contents in the target broadcast television program based on the multimodal time series data and first prior knowledge, wherein the first prior knowledge is used to describe examples of different program contents;
[0014] Splitting the multimodal time series data according to time periods corresponding to different program contents to obtain multimodal time series sub-data corresponding to each program content;
[0015] Call the artificial intelligence big model to obtain the type and title corresponding to each program content based on the multimodal sub-time series sub-data and the second prior knowledge; the second prior knowledge is used to describe examples of extracting program content titles and classification examples of different program content types.
[0016] According to the broadcast and television cataloging method based on the artificial intelligence large model provided by the present invention, the multimodal time series data includes video time series data;
[0017] Calling the artificial intelligence big model to obtain time periods corresponding to different program contents in the target broadcast television according to the multimodal time series data and the first prior knowledge, including:
[0018] Invoking an artificial intelligence big model, the artificial intelligence big model outputs initial time periods corresponding to different program contents based on the multimodal time series data and the first prior knowledge;
[0019] Determine a first blank time period according to an initial time period corresponding to each program content;
[0020] Obtaining video timing data corresponding to the first blank time period, and determining a shot splitting time point within the first blank time period based on the video timing data corresponding to the first blank time period;
[0021] Extract multiple frames of images near each shot split time point;
[0022] The artificial intelligence big model is called, and the artificial intelligence big model outputs a shot split time point as the actual dividing point of different program contents based on the multiple frames of images corresponding to each shot split time point;
[0023] The initial time periods corresponding to the different program contents are adjusted according to the actual dividing points of the different program contents to obtain the time periods corresponding to the different program contents.
[0024] According to the radio and television cataloging method based on the artificial intelligence large model provided by the present invention, the second prior knowledge includes: third prior knowledge and fourth prior knowledge, the third prior knowledge is used to describe examples of different program content types, and the fourth prior knowledge is used to describe examples of title extraction of program content under different types;
[0025] The AI model is called to obtain the type and title of each program content based on the multimodal sub-sequence sub-data and the second prior knowledge, including:
[0026] Calling the artificial intelligence big model, the artificial intelligence big model outputs the type of each program content based on the multimodal sub-time series sub-data and third prior knowledge;
[0027] The artificial intelligence big model is called, and the artificial intelligence big model outputs the title of each program content according to the type of program content, multimodal sub-time series sub-data and fourth prior knowledge.
[0028] According to the broadcast and television cataloging method based on the artificial intelligence big model provided by the present invention, after the artificial intelligence big model outputs the type of each program content based on the multimodal sub-time-sequence sub-data and the third prior knowledge, the method further includes:
[0029] If the program content is of the news type, the artificial intelligence model is invoked to obtain the time periods corresponding to the different news items contained in the program content based on the multimodal sub-time series sub-data of the program content and fifth prior knowledge; the fifth prior knowledge is used to describe examples of different news items;
[0030] Splitting the multimodal time series sub-data according to time periods corresponding to different news items to obtain multimodal time series grandchild data corresponding to each news item;
[0031] The artificial intelligence big model is called, and the artificial intelligence big model outputs the title of each news item based on the multimodal sub-time series data and the sixth prior knowledge; the sixth prior knowledge is used to describe an example of extracting the news item title; the catalog playlist includes the time sequence and title of the news items contained in the news program content.
[0032] According to the broadcast and television cataloging method based on the artificial intelligence large model provided by the present invention, the multimodal time series data includes video time series data;
[0033] The artificial intelligence model is called to obtain time periods corresponding to different news items included in the program content based on the multimodal sub-time series sub-data of the program content and the fifth prior knowledge, including:
[0034] Invoke the artificial intelligence model to output initial time periods corresponding to different news items based on the multimodal time series sub-data and the fifth prior knowledge;
[0035] Determining a second blank time period according to the initial time period corresponding to each news item;
[0036] Obtaining video timing data corresponding to the second blank time period, and determining a shot splitting time point within the second blank time period based on the video timing data corresponding to the second blank time period;
[0037] Acquire multiple frames of images near each shot cutting time point during the second blank time period;
[0038] The AI model is called to output the time point of each shot segmentation as the actual dividing point between different news items based on the multiple frames of images corresponding to each shot segmentation time point.
[0039] Adjust the initial time periods of different news items according to the actual dividing points of different news items to obtain the time periods corresponding to different news items
[0040] The present invention also provides a radio and television cataloging device based on an artificial intelligence large model, comprising the following modules:
[0041] The data acquisition module is used to acquire multimodal time series data of the target broadcast and television.
[0042] The program recognition module is used to call the artificial intelligence large model to recognize the multimodal time series data and obtain the timing, type and title of the program content contained in the target radio and television.
[0043] The catalog playlist generation module is used to generate the catalog playlist of the target radio and television according to the timing, type and title of the program content contained in the target radio and television.
[0044] The present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the radio and television cataloging method based on the artificial intelligence large model as described above is implemented.
[0045] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described radio and television cataloging methods based on an artificial intelligence large model.
[0046] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described radio and television cataloging methods based on an artificial intelligence large model.
[0047] The radio and television cataloging method and related products based on the artificial intelligence big model provided by the present invention use the artificial intelligence big model to recognize the multimodal time series data of radio and television, obtain the time series, type and title of the program content contained in the target radio and television, and generate a cataloging and playlist of the target radio and television based on the time series, type and title of the program content contained in the target radio and television. This eliminates the need to invest a large amount of manpower in template labeling, thereby realizing low-cost cataloging of radio and television. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 This is one of the flow charts of the radio and television cataloging method based on the artificial intelligence large model provided by an embodiment of the present invention.
[0050] Figure 2 This is the second flow chart of the radio and television cataloging method based on the artificial intelligence large model provided by an embodiment of the present invention.
[0051] Figure 3 3 is a flow chart of a method for acquiring multimodal time series data in an embodiment of the present invention.
[0052] Figure 4 4 is a flow chart of a method for obtaining broadcast and television cataloging parameters in an embodiment of the present invention.
[0053] Figure 5 4 is a flow chart of a method for determining initial time periods corresponding to different program contents in an embodiment of the present invention.
[0054] Figure 6 4 is a flow chart of a method for determining a program content type title in an embodiment of the present invention.
[0055] Figure 7 4 is a flow chart of a method for cataloging news program content in an embodiment of the present invention.
[0056] Figure 8 4 is a flow chart of a method for determining initial time periods corresponding to different news items in an embodiment of the present invention.
[0057] Figure 9 It is a structural diagram of a radio and television editing device based on an artificial intelligence large model provided by an embodiment of the present invention.
[0058] Figure 10It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0059] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0060] Existing intelligent cataloging systems for broadcast and television require extensive manpower to label cataloging templates and daily maintenance, resulting in complex and costly labor. To address these issues, the present invention introduces an artificial intelligence (AI) big model. Based on its cognitive understanding, it automatically catalogs broadcast and television programs. This eliminates the need for template labeling, improving efficiency and reducing labor costs.
[0061] It should be noted that the serial numbers assigned to the objects described in the present invention, such as "first", "second", etc., are only used to distinguish the objects described and do not have any order or technical meaning.
[0062] The following combination Figure 1-Figure 3 The present invention describes a radio and television cataloging method based on a large artificial intelligence model. This method can be applied to electronic devices such as terminal devices or servers. Terminal devices may include mobile phones, computers, tablets, and smart terminals; servers may include standalone servers, clustered servers, or cloud servers. This method can also be applied to a radio and television cataloging device installed in an electronic device such as a terminal device or server. This device can be implemented using software, hardware, or a combination of both.
[0063] Figure 1 and Figure 2 The flowchart of the radio and television cataloging method based on the artificial intelligence large model provided by the embodiment of the present invention is exemplified as follows: Figure 1 and Figure 2 As shown, the method includes the following:
[0064] Step 101: Acquire multimodal time series data of a target broadcast television.
[0065] The "modality" in multimodal timing data refers to data forms in different forms. In an embodiment of the present invention, the multimodal timing data of the target broadcasting television includes at least two of the video timing data, audio timing data and text timing data of the target broadcasting television.
[0066] Step 102: Call the artificial intelligence big model to recognize the multimodal time series data to obtain the timing, type and title of the program content contained in the target radio and television.
[0067] It should be noted that the program content in the embodiment of the present invention includes advertisements and programs, wherein the programs refer to news programs, variety shows, and film and television programs.
[0068] Step 103: Generate a catalog playlist of the target broadcasting television according to the timing, type and title of the program contents contained in the target broadcasting television.
[0069] The embodiments of the present invention can effectively make up for the shortcomings of traditional audio and video template comparison technology by utilizing the multimodal semantic recognition capabilities of large models, thereby further improving the intelligence level of cataloging, liberating the cataloging work of manually annotated templates, and significantly reducing the workload of manual labor.
[0070] In an example embodiment, Figure 3 The flowchart of the multimodal time series data acquisition method according to the embodiment of the present invention is shown as follows: Figure 3 As shown, step 101 can be specifically implemented through the following steps 201 to 203.
[0071] Step 201: Acquire audio and video timing data of a target broadcast television program.
[0072] Step 202: Convert the audio time series data into text time series data to obtain first text time series data.
[0073] Step 203: Perform text recognition on the subtitles in the video time series data to obtain second text time series data. The first text time series data, the second text time series data, and the video time series data constitute the multimodal time series data.
[0074] The above-mentioned conversion of audio time series data into text time series data and subtitle recognition can be achieved through existing dedicated models, which will not be described in detail in the present invention.
[0075] In an example embodiment, Figure 4 The flowchart of the method for obtaining broadcast and television catalog parameters in an embodiment of the present invention is shown as an example. Figure 4 As shown, step 102 can be specifically implemented through the following steps 301 to 303.
[0076] Step 301: Call the artificial intelligence model to obtain the time periods corresponding to different program contents in the target radio and television based on the multimodal time series data and the first prior knowledge.
[0077] The first prior knowledge is used to describe different program content examples, such as several advertisement examples, as well as several film and television programs, variety shows, and news programs. For advertisements, it's preferable to describe several advertisements with different styles. Similarly, for other program content, it's also possible to describe several program content with different styles. In this way, the AI model can segment different program content not only by understanding cognitive content but also by changes in language style.
[0078] When invoking the AI model, the prompt is to output the time period information corresponding to each program content. The AI model will analyze and understand the time period of different program content based on the second prior knowledge, and then provide the time period information corresponding to each program content.
[0079] This embodiment uses semantic understanding and time series analysis to determine program content boundaries and automatically segment program content. Specifically for continuously broadcast news, short advertisements, and columns, the large model can accurately segment content based on features such as language style changes and scene transitions.
[0080] Step 302: Split the multimodal time series data according to time periods corresponding to different program contents to obtain multimodal time series sub-data corresponding to each program content.
[0081] Step 303: Call the artificial intelligence model to obtain the type and title corresponding to each program content based on the multimodal sub-time series sub-data and the second prior knowledge. The second prior knowledge is used to describe examples of extracting program content titles and classifying different program content types.
[0082] When calling the AI model in step 303, the prompt is to output the type and title of each program content. The AI model will analyze and understand the multimodal sub-sequence sub-data corresponding to each program content based on the second prior knowledge, and then output the type and title of each program content.
[0083] Specifically, after splitting out independent advertisements and programs, the large model further identifies their specific categories. For example, the model can distinguish between TV dramas, variety shows, news columns, public service announcements, commercials, etc., and classify them based on a predefined label system. In order to further improve the content cataloging, the name of the advertisement or program can be automatically extracted from the text information of the advertisement or program. NER (named entity recognition) technology can be combined to identify key information such as title, brand, broadcast time, and guests. For example, in a variety show, the system can automatically extract the program name and distinguish different episodes. In advertisements, the model can extract brand names or product names from screen subtitles and voice content.
[0084] In an example embodiment, Figure 5 The flowchart of the method for determining the initial time period corresponding to different program contents in an embodiment of the present invention is shown as an example. Figure 5 As shown, step 301 can be specifically implemented through the following steps 401 to 406.
[0085] Step 401: Call the artificial intelligence big model, and the artificial intelligence big model outputs the initial time period corresponding to different program contents based on the multimodal time series data and the first prior knowledge.
[0086] Step 402: Determine a first blank time period according to the initial time period corresponding to each program content.
[0087] Step 403: Obtain video timing data corresponding to the first blank time period, and determine the shot splitting time point within the first blank time period based on the video timing data corresponding to the first blank time period.
[0088] The video timing data may be a multi-frame timing image. Based on the multi-frame timing image, an existing dedicated model may be used to identify the time point of shot switching. The present invention will not elaborate on the existing dedicated model.
[0089] Step 404: extract multiple frames of images near each shot segmentation time point.
[0090] Step 405: The AI model is called. Based on the multiple frames corresponding to each shot split time point, the AI model outputs a shot split time point as the actual demarcation point between different program contents. The prior knowledge provided to the AI model here is examples of temporally adjacent program contents and multiple image frames near the actual demarcation time points of these temporally adjacent program contents. The prompt is to output the actual demarcation point between different program contents.
[0091] Step 406: Adjust the initial time periods corresponding to the different program contents according to the actual demarcation points of the different program contents to obtain the time periods corresponding to the different program contents.
[0092] Because in step 401, after the artificial intelligence model splits different program contents based on understanding and analysis, a blank time period sometimes appears between two different program contents, that is, a time period that is not determined to be a certain program content. For example, the first program content is a news program, and the corresponding time period is 7:00:00-7:31:26. The second program content is an advertisement, and the corresponding time period is 7:31:31-7:32:00. The blank time period is 7:31:26-7:31:31. This blank time period may be some transitions. In order to more accurately divide the corresponding time periods of the program content, this embodiment specifically analyzes which parts of the blank time period belong to the previous program content and which parts belong to the next program content, which specifically corresponds to steps 402 to 406. This embodiment achieves accurate division of the corresponding time periods of different program contents by combining the transition distinction of video timing data.
[0093] In one exemplary embodiment, step 303 involves invoking the AI model to obtain the genre and title of each program content. Furthermore, this step can be performed by invoking the AI model once to obtain both the genre and title of each program content, or by invoking the AI model twice to obtain both the genre and title of each program content.
[0094] If the artificial intelligence large model is called twice to obtain the type and title of the program content respectively, then the second prior knowledge includes the third prior knowledge and the fourth prior knowledge. The third prior knowledge is used to describe examples of different program content types, and the fourth prior knowledge is used to describe examples of title extraction of program content under different types.
[0095] Figure 6 The flowchart of the method for determining the program content type title in the embodiment of the present invention is shown as an example. Figure 6 As shown, the determination of the program content type and title can be specifically achieved through the following steps 501 to 502.
[0096] Step 501: Invoke the AI model. The AI model outputs the type of each program content based on the multimodal sub-time series sub-data and the third prior knowledge. It is understood that the prompt word input for this invocation of the AI model is "output program content type."
[0097] Step 502: The AI model is invoked. Based on the program type, the multimodal sub-time-series sub-data, and the fourth prior knowledge, the AI model outputs the title of each program. It is understood that the prompt input for this invocation of the AI model is "output program title."
[0098] This embodiment first calls the AI model to determine the type of program content, and then calls it again to determine the program title. This allows the AI model to be fed with experience in extracting titles specific to specific genres, enabling more accurate title extraction. For example, after determining that a program is a variety show, when obtaining its title, the AI model will describe the unique patterns in variety show title acquisition within the fourth prior knowledge. The AI model then extracts the variety show title based on this prior knowledge.
[0099] In an exemplary embodiment, if the type of the program content is news, the news items contained therein are further split and the title of each news item is extracted. Figure 7 , which is specifically achieved through the following steps 601 to 603.
[0100] In step 601, if the program content is of the news type, the AI master model is invoked to obtain the time periods corresponding to the different news items contained in the news program content based on the multimodal sub-time series sub-data of the program content and the fifth prior knowledge. It is understood that the fifth prior knowledge is used to describe examples of different news items. The prompt input for invoking the AI master model this time is "output the time periods corresponding to the different news items."
[0101] Step 602: Split the multimodal time series sub-data according to the time periods corresponding to different news items to obtain multimodal time series sub-data corresponding to each news item.
[0102] Step 603: Invoke the AI model. Based on the multimodal sub-time series data and the sixth prior knowledge, the AI model outputs the title of each news item. It is understood that the prompt input for invoking the AI model this time is "output the title of each news item." The sixth prior knowledge describes an example of extracting news item titles, such as identifying the "headline" of a news item and, combined with the semantic information of the corresponding news item, summarizing and generating a news item title that best matches the content of the news item.
[0103] The embodiment of the present invention combines the large model with the OCR recognition results of the "title board" of the corresponding news item and the semantic information of the corresponding news item to summarize and generate the news item name that best matches the content of the news item, thereby achieving further (subdivided) item-level splitting of news programs and automatically generating a news item playlist.
[0104] It should be noted that the catalog playlist includes not only the sequence, type and title of each program content included in radio and television, but also the sequence and title of news items included in news programs.
[0105] In an example embodiment, after the artificial intelligence model splits the news items of a news program based on understanding and analysis, a blank time period may sometimes appear between two different news items. In order to more accurately determine which part of the blank time period belongs to the previous news item and which part belongs to the next news item, this embodiment provides a similar processing method to the first blank time period described above. Figure 8 , specifically as steps 701 to 706.
[0106] Step 701: Call the artificial intelligence model to output the initial time periods corresponding to different news items based on the multimodal time series sub-data and the fifth prior knowledge.
[0107] Step 702: Determine a second blank time period according to the initial time period corresponding to each news item.
[0108] Step 703: Obtain video timing data corresponding to the second blank time period, and determine the shot splitting time point within the second blank time period based on the video timing data corresponding to the second blank time period.
[0109] Step 704: Acquire multiple frames of images near each shot segmentation time point within the second blank period.
[0110] Step 705: Call the artificial intelligence model to output a shot-segmentation time point as the actual dividing point between different news items based on the multiple frames of images corresponding to each shot-segmentation time point.
[0111] Step 706: Adjust the initial time periods of different news items according to the actual dividing points of different news items to obtain time periods corresponding to the different news items.
[0112] The embodiments of the present invention catalog radio and television by utilizing the multimodal semantic recognition capabilities of large models, thereby relieving the cataloging work of manually annotating templates, significantly reducing the workload of manual labor, and further improving the intelligence level of cataloging.
[0113] The following describes the radio and television cataloging device based on the artificial intelligence big model provided by the present invention. The radio and television cataloging device based on the artificial intelligence big model described below and the radio and television cataloging method based on the artificial intelligence big model described above can be referenced to each other.
[0114] See also Figure 9 The radio and television cataloging device based on the artificial intelligence large model includes the following modules:
[0115] The data acquisition module 801 is used to acquire multimodal time series data of the target broadcast television.
[0116] The program recognition module 802 is used to call the artificial intelligence big model to recognize the multimodal time series data and obtain the time series, type and title of the program content contained in the target broadcasting and television.
[0117] The catalog playlist generating module 803 is used to generate a catalog playlist of the target broadcast television according to the timing, type and title of the program content contained in the target broadcast television.
[0118] Figure 10 An example of a physical structure diagram of an electronic device is shown below. Figure 10 As shown, the electronic device may include: a processor 1010, a communication interface 1020, a memory 1030, and a communication bus 1040, wherein the processor 1010, the communication interface 1020, and the memory 1030 communicate with each other via the communication bus 1040. The processor 1010 may call the logic instructions in the memory 1030 to execute a radio and television cataloging method based on an artificial intelligence large model, the method comprising: obtaining multimodal time series data of a target radio and television; calling the artificial intelligence large model to recognize the multimodal time series data to obtain the time series, type, and title of the program content contained in the target radio and television; and generating a catalog playlist of the target radio and television according to the time series, type, and title of the program content contained in the target radio and television.
[0119] In addition, the logic instructions in the above-mentioned memory 1030 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0120] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the radio and television cataloging method based on the artificial intelligence big model provided by the above methods, and the method includes: obtaining multimodal time series data of the target radio and television; calling the artificial intelligence big model to recognize the multimodal time series data, and obtaining the timing, type and title of the program content contained in the target radio and television; generating a cataloging playlist of the target radio and television based on the timing, type and title of the program content contained in the target radio and television.
[0121] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the radio and television cataloging method based on the artificial intelligence big model provided by the above-mentioned methods, the method comprising: obtaining multimodal time series data of the target radio and television; calling the artificial intelligence big model to recognize the multimodal time series data, and obtaining the timing, type and title of the program content contained in the target radio and television; generating a cataloging playlist of the target radio and television based on the timing, type and title of the program content contained in the target radio and television.
[0122] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0123] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A radio and television cataloging method based on an artificial intelligence large model, characterized in that: include: Obtain multimodal time series data of target broadcasting and television; Invoke an artificial intelligence model to recognize the multimodal time series data and obtain the time series, type, and title of the program content contained in the target broadcast and television; A catalog playlist of the target broadcasting television is generated based on the timing, type and title of the program content contained in the target broadcasting television.
2. The radio and television cataloging method based on artificial intelligence large model according to claim 1 is characterized in that: Obtain multimodal time series data of the target broadcast television, including: Obtain audio and video timing data of the target broadcasting and television; Convert the audio time series data into text time series data to obtain first text time series data; Perform text recognition on the subtitles in the video timing data to obtain second text timing data; the first text timing data, the second text timing data and the video timing data constitute the multimodal timing data.
3. The radio and television cataloging method based on artificial intelligence large model according to claim 1 or 2 is characterized in that: Calling the artificial intelligence model to recognize the multimodal time series data to obtain the time series, type and title of the program content contained in the target broadcast and television, including: Invoking the artificial intelligence model to obtain time periods corresponding to different program contents in the target broadcast television program based on the multimodal time series data and first prior knowledge, wherein the first prior knowledge is used to describe examples of different program contents; Splitting the multimodal time series data according to time periods corresponding to different program contents to obtain multimodal time series sub-data corresponding to each program content; Call the artificial intelligence big model to obtain the type and title corresponding to each program content based on the multimodal sub-time series sub-data and the second prior knowledge; the second prior knowledge is used to describe examples of extracting program content titles and classification examples of different program content types.
4. The radio and television cataloging method based on artificial intelligence large model according to claim 3 is characterized in that: The multimodal time series data includes video time series data; Calling the artificial intelligence big model to obtain time periods corresponding to different program contents in the target broadcast television according to the multimodal time series data and the first prior knowledge, including: Invoking an artificial intelligence big model, the artificial intelligence big model outputs initial time periods corresponding to different program contents based on the multimodal time series data and the first prior knowledge; Determine a first blank time period according to an initial time period corresponding to each program content; Obtaining video timing data corresponding to the first blank time period, and determining a shot splitting time point within the first blank time period based on the video timing data corresponding to the first blank time period; Extract multiple frames of images near each shot split time point; The artificial intelligence big model is called, and the artificial intelligence big model outputs a shot split time point as the actual dividing point of different program contents based on the multiple frames of images corresponding to each shot split time point; The initial time periods corresponding to the different program contents are adjusted according to the actual dividing points of the different program contents to obtain the time periods corresponding to the different program contents.
5. The radio and television cataloging method based on artificial intelligence large model according to claim 3 is characterized in that: The second priori knowledge includes: third priori knowledge and fourth priori knowledge, wherein the third priori knowledge is used to describe examples of different program content types, and the fourth priori knowledge is used to describe examples of extracting titles of program content under different types; The AI model is called to obtain the type and title of each program content based on the multimodal sub-sequence sub-data and the second prior knowledge, including: Calling the artificial intelligence big model, the artificial intelligence big model outputs the type of each program content based on the multimodal sub-time series sub-data and third prior knowledge; The artificial intelligence big model is called, and the artificial intelligence big model outputs the title of each program content according to the type of program content, multimodal sub-time series sub-data and fourth prior knowledge.
6. The method for cataloguing broadcast and television based on artificial intelligence large model according to claim 5, characterized in that: After the artificial intelligence model outputs the type of each program content based on the multimodal sub-time series sub-data and the third prior knowledge, it also includes: If the program content is of the news type, the artificial intelligence model is invoked to obtain the time periods corresponding to the different news items contained in the program content based on the multimodal sub-time series sub-data of the program content and fifth prior knowledge; the fifth prior knowledge is used to describe examples of different news items; Splitting the multimodal time series sub-data according to time periods corresponding to different news items to obtain multimodal time series grandchild data corresponding to each news item; The artificial intelligence big model is called, and the artificial intelligence big model outputs the title of each news item based on the multimodal sub-time series data and the sixth prior knowledge; the sixth prior knowledge is used to describe an example of extracting the news item title; the catalog playlist includes the time sequence and title of the news items contained in the news program content.
7. The radio and television cataloging method based on artificial intelligence large model according to claim 6 is characterized in that: The multimodal time series data includes video time series data; The artificial intelligence model is called to obtain time periods corresponding to different news items included in the program content based on the multimodal sub-time series sub-data of the program content and the fifth prior knowledge, including: Invoke the artificial intelligence model to output initial time periods corresponding to different news items based on the multimodal time series sub-data and the fifth prior knowledge; Determining a second blank time period according to the initial time period corresponding to each news item; Obtaining video timing data corresponding to the second blank time period, and determining a shot splitting time point within the second blank time period based on the video timing data corresponding to the second blank time period; Acquire multiple frames of images near each shot cutting time point during the second blank time period; The AI model is called to output the time point of each shot segmentation as the actual dividing point between different news items based on the multiple frames of images corresponding to each shot segmentation time point. The initial time periods of different news items are adjusted according to actual dividing points of different news items to obtain time periods corresponding to the different news items.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, it implements the radio and television cataloging method based on the artificial intelligence large model as described in any one of claims 1 to 7.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the radio and television cataloging method based on the artificial intelligence large model as described in any one of claims 1 to 7 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the radio and television cataloging method based on the artificial intelligence large model as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and system for cataloging news video
CN101616264A
New generation intelligent cataloging system and method facing large amount of broadcast television programs
CN102075695A
A method and system for automatic classification of video content recognition
CN119741636A