Data synthesis method for model training of media asset query task in multimedia field
Through the LLM large language model, query corpus templates matching the multimedia search scenario are generated, and combined with the expansion of the multimedia media library data, the problem of media query task model training relies on a large amount of labeled data, and efficient and low-cost data processing and model training effects are achieved.
Patent Information
- Application Number
- CN202510137228.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-27
AI Technical Summary
In the multimedia field, model training of media resource query tasks relies on a large amount of labeled data, which makes data labeling time-consuming and labor-intensive, cost-effective and ineffective.
The LLM large language model is adopted, based on semantic analysis and scene adaptation, a query corpus template matching the multimedia search scenario is generated, and the data from the multimedia media library is expanded to generate new query corpus data for model training.
Through automated and intelligent data synthesis technology, the cost and time of model training is significantly reduced, the efficiency of data preparation and the effect of model training is improved, and it is suitable for media query tasks in different fields and scenarios.
Smart Images

Figure CN120046756A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data processing technologies in the multimedia field, and particularly to a data synthesis method for model training of media resource query tasks in the multimedia field. Background Art
[0002] In the multimedia field, model training for media resource query tasks usually relies on a large amount of labeled data. However, in traditional methods, the data annotation process is time-consuming, laborious, and costly. In addition, the shortage of high-quality corpus in the original data also limits the effect of model training. Therefore, there is an urgent need for an efficient, low-cost, and high-quality data processing method. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a data synthesis method for model training of media resource query tasks in the multimedia field, so as to solve the technical problems that current data processing relies on a large amount of labeled data, thus being time-consuming, laborious, costly, and having poor effects.
[0004] Technical Solution
[0005] A data synthesis method for model training of media resource query tasks in the multimedia field, characterized in that:
[0006] First, obtain the data of the basic query corpus, and use the LLM large language model to generate a query corpus template that matches the multimedia retrieval scenario based on semantic analysis and scenario adaptation, including: performing semantic parsing on the query corpus through the LLM, extracting key features, and generating diverse query abstract templates according to the multimedia retrieval scenario; based on the query abstract template, combine and expand the media resource data in the multimedia media resource library as the entity objects of the query abstract template to generate new query corpus data, and finally use the new query corpus data as training data for model training of media resource query tasks.
[0007] Further, the data of the basic query corpus includes the data of keywords recognized by voice.
[0008] Further, the query abstract template includes statement templates of various feature combinations formed by the data of the extracted keywords, and query statement templates generated based on the multimedia retrieval scenario; this template includes a task part template and an entity object part template. The generation method of the task part template is to generate diversely through the generation ability of the large language model according to the requirements of the multimedia retrieval scenario, or to generate customized and scenario-based expansions through the large language model, and perform type matching according to the media types in the multimedia media resource library; the generation method of the entity object part template is to perform feature replacement and expansion using the data of the media types that match the corresponding multimedia media resource library.
[0009] Furthermore, the generation method of the entity object part template is as follows: using the abstract template for query to replace and expand all or some common data of the appropriate media type in the corresponding multimedia media library, so as to synthesize and generate a complete new query corpus, and the quantity and type of selected data are determined according to the training task.
[0010] Furthermore, the data in the multimedia media library needs to be cleaned, sorted out and prepared first. The data is classified according to the media type, and it is ensured that the data of each field defined in each category is accurate, not empty and not repeated.
[0011] The data in the multimedia media library includes video media data, and the defined data fields included in this data are: media_name, region, actor, director, year, type, tag; the data in the multimedia media library also includes channel media data, and the defined data fields included in this data are: tv_channel; the data in the multimedia media library also includes music media data, and the defined data fields included in this data are: media_name, region, singer, producer, year, type, tag.
[0012] The data of the media type adopted by the entity object part template uses the data fields defined by the video media data, the channel media data or the music media data, and corresponding data field replacement and expansion are performed during expansion.
[0013] After the expansion is completed, the relevance of the values of each data field of the replaced object data is automatically detected.
[0014] Furthermore, the trained model for the media query task is used for subsequent inference and testing.
[0015] A storage medium stores a program, characterized in that the program, when executed by a processor, implements the method described in any one of the above.
[0016] Beneficial effects
[0017] The method of the present invention uses a large language model (LLM) to generate multiple abstract templates for queries based on the obtained query corpus data and in the context of a multimedia retrieval scenario. On the basis of the generated abstract templates for queries, a method of generating new query corpus by combining and expanding with the data in the multimedia asset library utilizes automated and intelligent data synthesis technology, effectively reducing the cost of model training. The adopted method of generating templates has extremely high scalability and can adapt to media asset query tasks in different fields and scenarios. By flexibly adjusting the parameters and inputs of the large language model, corpus templates applicable to new fields and new tasks can be quickly generated. The method of the present invention realizes the engineering expansion of data synthesis, provides the ability of precise control and high customization, can meet different requirements of various trainings, expands the application while reducing the cost, and has high application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The following further elaborates the present invention in conjunction with specific embodiments and the accompanying drawings.
[0020] The present invention proposes a data synthesis method for model training in media asset query tasks in the multimedia field. First, it obtains the data of the query corpus, uses a large language model (LLM), and based on semantic analysis and scenario adaptation, generates query corpus templates that match the multimedia retrieval scenario. Specifically, it includes: performing semantic parsing on the query corpus through the LLM, extracting key features (such as media type, name, tag, etc.), and generating diverse abstract templates for queries according to the multimedia retrieval scenario; based on the abstract templates for queries, combining and expanding the media asset data in the multimedia asset library as the entity objects of the abstract templates for queries to generate new query corpus data. Finally, the new query corpus data is used as training data for model training in media asset query tasks. The specific implementation manners include the following steps:
[0021] 1. Collect and clean the data in the multimedia asset library to ensure the accuracy and integrity of the data; including cleaning, sorting, and preparing the data in the multimedia asset library to ensure that the data in each field is accurate, not empty, and not repeated, specifically:
[0022] - The video media asset data includes the following fields: media_name, region, actor, director, year, type, tag;
[0023] - The channel media asset data includes the following field: tv_channel;
[0024] - The music media asset data contains the following fields: media_name, region, singer, producer, year, type, tag.
[0025] 2. Generate an abstract template for query using a large language model; first, obtain some query basic corpus, and then use the LLM to generate some abstract templates for query based on the multimedia retrieval scenario and the query basic corpus, such as:
[0026] - The basic corpus is: I want to watch the movie "Shanghai Beach".
[0027] - Annotation: {"intent": "play_video", "entities": {"media_name": "Shanghai Beach", "type": "movie"}}
[0028] Generate the abstract template: Want to watch XY, corresponding to {"intent": "play_video", "entities": {"media_name": "Y", "type": "X"}}.
[0029] The generated template includes a task part template and an entity object part template. The previous "want to watch" task part template can be diversified or customized or scene-expanded using the text generation ability of the LLM and matched according to the media types in the multimedia media asset library. The latter entity object part template needs to be expanded according to the data in the media asset library established in step 1.
[0030] 3. Based on the multimedia media asset library data, expand the abstract template for query according to different entity objects. For example: For the above-generated abstract template: Want to watch XY, use the multimedia media asset library data to replace and expand the entity object part XY to generate new query corpus. The specific examples are as follows:
[0031] Expansion generation example 1: (Data replacement and combination of the same template and the same video type)
[0032] - The newly generated query corpus after expansion: I want to watch the movie "Shaolin Temple".
[0033] - Annotation: {"intent": "play_video", "entities": {"media_name": "Shaolin Temple", "type": "movie"}}
[0034] The task part template of the abstract template part is the same, and the entity object part template for matching uses data replacement and combination of the same video type.
[0035] Expansion Generation Example 2: (Same template, but data replacement and combination for different video types)
[0036] - New query corpus generated by expansion: I want to watch the TV drama The Knockout
[0037] - Labeled as: {"intent":"play_video","entities":{"media_name":"The Knockout","type":"TV drama"}}
[0038] The task part template of the abstract template is the same, and the entity object part template for matching uses data replacement and combination of different video types.
[0039] Expansion Generation Example 3: (Data replacement and expansion for different media asset types)
[0040] For the abstract template generated above: Want to watch XY, according to the scenario and LLM expansion, the expanded abstract template is: I want to listen to X, and then combined with the data of the media asset type corresponding to the multimedia media asset library data, the new query corpus generated by expansion: I want to listen to Qi Li Xiang
[0041] - Labeled: {"intent":"play_music","entities":{"media_name":"Qi Li Xiang"}}
[0042] Expansion Generation Example 4: (Corpus templates for different media asset types)
[0043] Still for the abstract template generated above: Want to watch XY, according to the scenario and LLM expansion, the expanded abstract template is: I want to watch Z, and then combined with the data of the media asset type corresponding to the multimedia media asset library data, the new query corpus generated by expansion: I want to watch CCTV-1
[0044] - Labeled: {"intent":"play_tv_live","entities":{"tv_channel":"CCTV-1"}}
[0045] The above four extended generation examples are to replace and combine the corresponding data types of data in the multimedia asset library based on the abstract template in step 2. For data of different data types, corresponding abstract templates will be generated, and replacement or extension will be performed according to the situation of the data of this type to form new query corpora respectively. The scope of data replacement is limited to various defined feature fields and entities in the asset library. For specific fields, please refer to the previous field description section. Each time a replacement is made, the values of these fields are read from the database. Moreover, it is necessary to ensure that the values of the fields to be replaced are relevant. That is, if the type is a movie, then the media_name needs to be the movie name.
[0046] 4. Use the newly generated query corpus after extension for the training of the large language model.
[0047] 5. Perform inference and testing on the trained model to evaluate its performance.
[0048] The above embodiments show how to improve the efficiency and quality of model training through the data synthesis method while reducing costs. The experimental results show that it only takes 1 hour to generate 10,000 training corpora using this method, while the traditional manual annotation method takes 100 hours; the model training accuracy is increased by 15% (from 80% to 95%). This method is applicable to the media query tasks in the multimedia field and has a wide range of application prospects.
[0049] Of course, more abstract templates for queries can also be generated through the LLM according to the application scenario and the obtained basic query corpus, and then extended according to the data in the multimedia asset library, or different abstract templates can be generated by analyzing the common ways of different query corpora for different types of multimedia data, and then the data synthesis and extension can be performed, which are all feasible. It can also be extended according to the requirements of specific model training.
[0050] Therefore, the technical solution of the present invention has the following multiple advantages and beneficial effects:
[0051] The technical solution of the present invention is efficient and avoids the human consumption of data annotation. The present invention significantly improves the efficiency of data preparation and model training through an automated data synthesis process. In traditional methods, data annotation relies on a large amount of manpower, is time-consuming and error-prone. The method of this patent uses the large language model to automatically generate and expand the corpus without manual intervention, greatly reducing the labor cost and time consumption, and ensuring the efficiency of data preparation.
[0052] The technical solution of the present invention has low costs. Through automated and intelligent data synthesis technologies, the costs of model training are effectively reduced. Compared with traditional manual annotation methods, this method reduces the input of manpower and time, and at the same time avoids the additional costs caused by manual annotation errors. In addition, based on the corpus generation technology of large models, the costs of data preparation are further reduced, making model training more economical and efficient.
[0053] The method of the present invention has extremely high scalability and can adapt to media asset query tasks in different fields and scenarios. By flexibly adjusting the parameters and inputs of the large language model, corpus templates applicable to new fields and new tasks can be quickly generated. This flexibility and scalability enable this method to be widely applied to media asset query tasks in various multimedia fields and meet the needs of different users.
[0054] The method of the present invention uses a large language model to generate high-quality corpus templates, ensuring the accuracy and richness of the corpus templates. Large models have powerful language understanding and generation capabilities and can generate diverse corpus templates that conform to actual application scenarios. Such high-quality corpus not only improves the effect of model training but also enhances the generalization ability and robustness of the model.
[0055] The method of the present invention effectively compensates for the shortage of high-quality corpus in the original data through automated data synthesis technology. In traditional methods, the original data often has problems such as uneven corpus quality and incomplete coverage. This method generates corpus templates through a large model and then expands the corpus through engineering methods, providing a large amount of high-quality, diverse, scenario-based, and customized corpus, filling the deficiencies of the original data, and improving the overall quality of model training.
[0056] The method of the present invention realizes the technical combination of large model generation and engineering expansion of data synthesis, providing the ability of precise control and high customization. By flexibly adjusting the data synthesis process and parameters, the generation and expansion process of the corpus can be precisely controlled to meet the personalized needs of different users and scenarios. This engineering expansion and customization ability enable this method to flexibly handle various complex application scenarios and improve the flexibility and practicality of model training.
[0057] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A data synthesis method for model training of multimedia resource query tasks, characterized by: Firstly, the data of basic query corpus is obtained, and the LLM large language model is adopted to generate query corpus templates matching the multimedia retrieval scenario based on semantic analysis and scenario adaptation, including: semantically parsing the query corpus through LLM, extracting key features, and generating diversified query abstract templates according to the multimedia retrieval scenario; based on the query abstract template, the media data of the multimedia media library is combined and expanded as the entity object of the query abstract template to generate new query corpus data, and finally the new query corpus data is used as training data for model training of media query tasks.
2. The data synthesis method for model training of multimedia resource query tasks as claimed in claim 1, characterized in that: The data of the basic query corpus includes data of keywords obtained through speech recognition.
3. The data synthesis method for model training of multimedia resource query tasks according to claim 1 or 2, characterized in that: The query abstract template includes sentence templates of various feature combinations formed by the extracted keyword data, and query sentence templates generated based on multimedia retrieval scenarios; The template includes a task part template and an entity object part template. The task part template is generated by using the generation capability of a large language model and performing diversified expansion generation through a large language model according to multimedia retrieval scenario requirements, or by performing customized and scenario-based expansion generation through a large language model, and performing type matching according to the media type of a multimedia media library; the entity object part template is generated by using data that matches the media type of a corresponding multimedia media library for feature replacement and expansion.
4. The data synthesis method for model training of multimedia resource query tasks as claimed in claim 3, characterized in that: The entity object partial template is generated in the following way: the query is replaced and expanded with all or part of the commonly used data of the appropriate media type in the corresponding multimedia media library using an abstract template to match the query, thereby synthesizing a complete new query corpus, and the amount and type of data selected depend on the training task.
5. The data synthesis method for model training of multimedia resource query tasks as claimed in claim 3, characterized in that: The data in the multimedia asset library needs to be cleaned, sorted and prepared first, and the data is classified according to the media asset type, and it is ensured that the data of each field defined in each category is accurate, not empty, and not repeated.
6. The data synthesis method for model training of multimedia resource query tasks as claimed in claim 5, characterized in that: The data of the multimedia media library includes video media data, and the defined data fields contained in the data include: media_name, region, actor, director, year, type, tag; the data of the multimedia media library also includes channel media data, and the defined data fields contained in the data include: tv_channel; the data of the multimedia media library also includes music media data, and the defined data fields contained in the data include: media_name, region, singer, producer, year, type, tag.
7. The data synthesis method for model training of multimedia resource query tasks as claimed in claim 6, characterized in that: The data of the media asset type adopted by the entity object part template adopts the data field defined by the video media asset data, the channel media asset data or the music media asset data, and the data of the corresponding data field is replaced and expanded during expansion.
8. The data synthesis method for model training of multimedia resource query tasks as claimed in claim 7, characterized in that: After the expansion is completed, the relevance of the values of each data field of the replaced object data is automatically detected.
9. The data synthesis method for model training of multimedia resource query tasks as claimed in claim 1, characterized in that: The trained model for media resource query tasks is used for subsequent reasoning and testing.
10. A storage medium storing a program, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.