Intelligent media asset management and content production method and system, terminal and storage medium

By employing intelligent media asset management and content production methods, the problems of chaotic media asset management and lack of intelligent labeling have been solved, enabling efficient media asset management, rapid material retrieval, and automated audio and video synthesis, thereby improving content production efficiency and security.

CN120950707APending Publication Date: 2025-11-14GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ) +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511470128.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing media industry content management and generation systems suffer from chaotic media asset management and a lack of intelligent annotation, leading to difficulties for users in finding materials, low efficiency in audio and video synthesis, and an inability to meet user needs.

Method used

By employing intelligent media asset management and content production methods, including data cleaning, multi-agent collaborative annotation, multimodal coding, and retrieval, a vector index library is constructed to achieve efficient media asset management and automatic audio-visual synthesis.

Benefits of technology

It significantly improved media asset management efficiency, enhanced material retrieval and content generation capabilities, strengthened media asset understanding and review capabilities, and realized an intelligent content production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950707A_ABST
    Figure CN120950707A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and discloses an intelligent media asset management and content production method and system, a terminal and a storage medium, and the method comprises the steps: obtaining original media asset data, and carrying out the basic information extraction and data cleaning processing, and obtaining target media asset data; determining a multi-agent cooperation labeling module, and performing labeling processing on the target media asset data through the multi-agent cooperation labeling module to obtain a media asset label; encoding the target media asset data and the media asset label to obtain a visual index and a label index, and constructing a vector index database according to the visual index and the label index; obtaining a user query statement, and performing media asset retrieval in the vector index database according to the user query statement to obtain a media asset retrieval result; and obtaining synthetic data, and performing time sequence alignment processing on the synthetic data and the media asset retrieval result to obtain a target synthetic audio / video. The media resource data management efficiency and the audio and video synthesis efficiency can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to an intelligent media asset management and content production method, system, terminal, and computer-readable storage medium. Background Technology

[0002] Content assets generated by new media, or media assets for short, include various forms of media materials such as images, text, and audio. With the increasing growth of the new media industry, media asset data is characterized by large volume, high dimensionality and diversity, and high overhead and storage costs, which poses challenges to the management of media content.

[0003] Existing media industry content management and generation systems suffer from chaotic media asset management and a lack of intelligent annotation, which makes it very difficult for users to find relevant materials, and the efficiency of audio and video synthesis is low, making it difficult to meet users' needs for media asset data.

[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0005] The main objective of this invention is to provide an intelligent media asset management and content production method, system, terminal, and computer-readable storage medium. This invention aims to solve the problems in existing media industry content management and generation systems, such as chaotic media asset management and lack of intelligent annotation, which leads to difficulties for users in finding materials, low efficiency in audio and video synthesis, and inability to meet users' needs for media asset data.

[0006] To achieve the above objectives, the present invention provides an intelligent media asset management and content production method, which includes the following steps: Obtain raw media asset data, and perform basic information extraction and data cleaning processing on the raw media asset data to obtain target media asset data; A multi-agent collaborative annotation module is identified, and the target media asset data is annotated using the multi-agent collaborative annotation module to obtain media asset tags; The target media asset data and the media asset tags are encoded to obtain a visual index and a tag index, and a vector index library is constructed based on the visual index and the tag index. Obtain the user's query statement, and perform media asset retrieval in the vector index database based on the user's query statement to obtain the media asset retrieval results; Acquire the synthesized data, and perform time-series alignment processing on the synthesized data and the media asset retrieval results to obtain the target synthesized audio and video.

[0007] Optionally, the intelligent media asset management and content production method, wherein acquiring raw media asset data and performing basic information extraction and data cleaning processing on the raw media asset data to obtain target media asset data specifically includes: Obtain raw media asset data, wherein the raw media asset data includes raw media asset images and raw media asset videos; The original media asset images and original media asset videos are subjected to basic information extraction processing to obtain basic media asset data. The basic information extraction processing includes format verification processing, file type identification processing, and basic data extraction processing. The basic media asset data is cleaned to obtain the target media asset data. The data cleaning process includes deduplication, abnormal data removal, data compliance review, and data quality control.

[0008] Optionally, in the intelligent media asset management and content production method, the step of determining a multi-agent collaborative annotation module and using the multi-agent collaborative annotation module to annotate the target media asset data to obtain media asset tags specifically includes: Identify the multi-agent collaborative annotation module and determine the media asset type of the target media asset data; The task scheduler assigns the target media asset data to the multi-agent collaborative annotation module according to the media asset type, and the multi-agent collaborative annotation module performs annotation processing on the target media asset data to obtain media asset tags.

[0009] Optionally, in the intelligent media asset management and content production method, the multi-agent collaborative annotation module includes a semantic understanding agent, a domain calibration agent, a granularity and synonym specification agent, and a proofreading agent. The step of annotating the target media asset data through the multi-agent collaborative annotation module to obtain media asset tags specifically includes: The semantic understanding agent invokes the image understanding big model or the video understanding big model, and performs semantic parsing processing on the target media asset data through the image understanding big model or the video understanding big model to obtain candidate tags; A domain-specific vocabulary is obtained, and the candidate labels are corrected and supplemented based on the domain-specific vocabulary by the domain calibration agent to obtain calibration labels. A specific rule base is obtained, and the calibration label is refined in granularity and expanded in synonyms according to the specific rule base by the granularity and synonym specification agent to obtain the specification label; A quality control strategy is determined, and the standard label is verified by a review agent according to the quality control strategy to obtain the media asset label.

[0010] Optionally, the intelligent media asset management and content production method, wherein encoding the target media asset data and the media asset tags to obtain a visual index and a tag index, and constructing a vector index library based on the visual index and the tag index, specifically includes: A multimodal vector model is determined, wherein the multimodal vector model includes a multimodal visual vector model and a multimodal text vector model; The target media asset data is processed by multimodal visual encoding using the multimodal visual vector model to obtain a visual index; The media asset tags are processed using the multimodal text vector model to perform multimodal text encoding to obtain the tag index; A vector index library is constructed based on the visual index and the label index.

[0011] Optionally, the intelligent media asset management and content production method, wherein obtaining the user query statement and performing media asset retrieval in the vector index database based on the user query statement to obtain media asset retrieval results specifically includes: Obtain the user query statement and encode the user query statement to obtain the target query vector, wherein the user query statement includes a text query statement or an image query statement; The target query vector is compared and retrieved with the visual index and the tag index in the vector index library using the approximate nearest neighbor search method to obtain multiple image matching results or multiple video matching results; Calculate the similarity between multiple image matching results or multiple video matching results and the user query statement to obtain a similarity result; Based on the similarity results, a preset number of target image matching results or target video matching results are selected from multiple image matching results or multiple video matching results in descending order to obtain media asset retrieval results.

[0012] Optionally, in the intelligent media asset management and content production method, the synthesized data includes narration and background music; The step of acquiring synthetic data and performing time-series alignment processing on the synthetic data and the media asset retrieval results to obtain the target synthetic audio and video specifically includes: The user query is input into a preset large language model to obtain the spoken text, and the spoken text is input into a TTS model to obtain the narration and voice-over. The user query is input into the music generation model to obtain background music; The narration, background music, and media asset retrieval results are aligned and overlaid along the same timeline to obtain the target synthesized audio and video.

[0013] Furthermore, to achieve the above objectives, the present invention also provides an intelligent media asset management and content production system, wherein the intelligent media asset management and content production system includes: The data cleaning module is used to acquire raw media asset data and perform basic information extraction and data cleaning processing on the raw media asset data to obtain target media asset data. A data annotation module is used to determine a multi-agent collaborative annotation module and to annotate the target media asset data through the multi-agent collaborative annotation module to obtain media asset tags. An index building module is used to encode the target media asset data and the media asset tags to obtain a visual index and a tag index, and to build a vector index library based on the visual index and the tag index. The media asset retrieval module is used to obtain the user's query statement and perform media asset retrieval in the vector index database according to the user's query statement to obtain the media asset retrieval results; The video output module is used to acquire the synthesized data and perform time-series alignment processing on the synthesized data and the media asset retrieval results to obtain the target synthesized audio and video.

[0014] In this invention, raw media asset data is acquired, and basic information extraction and data cleaning are performed on the raw media asset data to obtain target media asset data. A multi-agent collaborative annotation module is determined, and the target media asset data is annotated through the multi-agent collaborative annotation module to obtain media asset tags. The target media asset data and the media asset tags are encoded to obtain a visual index and a tag index, and a vector index library is constructed based on the visual index and the tag index. A user query statement is acquired, and media asset retrieval is performed in the vector index library based on the user query statement to obtain media asset retrieval results. Synthetic data is acquired, and the synthetic data and the media asset retrieval results are time-series aligned to obtain the target synthetic audio and video. This invention, by cleaning and annotating the raw media asset data to construct a vector index library, can effectively improve the efficiency of user management and retrieval of media asset data. Furthermore, by time-series aligning the synthetic data with the media asset retrieval results, automatic audio and video synthesis can be achieved, significantly reducing manual editing operations and improving the production efficiency of audio and video content. Attached Figure Description

[0015] Figure 1 This is a flowchart of a preferred embodiment of the intelligent media asset management and content production method of the present invention; Figure 2 This is a schematic diagram of the overall architecture of the implementation process of a preferred embodiment of the intelligent media asset management and content production method of the present invention; Figure 3 This is a schematic diagram of the automated media asset cleaning process of a preferred embodiment of the intelligent media asset management and content production method of the present invention; Figure 4 This is a schematic diagram of the automated media asset annotation process, which is a preferred embodiment of the intelligent media asset management and content production method of the present invention. Figure 5 This is a schematic diagram of the media asset multimodal retrieval process, which is a preferred embodiment of the intelligent media asset management and content production method of the present invention. Figure 6 This is a schematic diagram of a preferred embodiment of the image-text comparison learning framework of the intelligent media asset management and content production method of the present invention; Figure 7 This is a schematic diagram of the LoRA principle of a preferred embodiment of the intelligent media asset management and content production method of the present invention; Figure 8 This is a schematic diagram of the content production process of a preferred embodiment of the intelligent media asset management and content production method of the present invention; Figure 9 This is a schematic diagram of the data security review process of a preferred embodiment of the intelligent media asset management and content production method of the present invention; Figure 10 This is a structural diagram of a preferred embodiment of the intelligent media asset management and content production system of the present invention; Figure 11 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0017] Traditional media industry content management and generation systems have the following main drawbacks: 1. Chaotic media asset management: Massive media resources such as images and videos are scattered in various folders or storage media, lacking a unified management platform, resulting in serious problems such as resource duplication, loss, and misuse.

[0018] 2. Lack of intelligent annotation: Most media assets are not structured or semantically annotated after collection, making efficient retrieval via keywords and semantics impossible and severely restricting content retrieval efficiency. 3. Low material retrieval efficiency: Publicity departments often need to manually search through historical files when using materials, resulting in low retrieval efficiency and difficulty in meeting the needs of rapid content production. 4. Cumbersome content creation process: Traditional video editing relies on manual material selection, editing, and voice-over, which is complex and time-consuming, making it difficult to support the growing demand for short video publicity. 5. Lack of content security review mechanism: During the management and use of media assets, there is a lack of automated review capabilities for sensitive content, posing security and compliance risks.

[0019] To address the problems of chaotic media asset management, low material retrieval efficiency, complex content generation processes, and lack of semantic understanding and security review capabilities in traditional media industry content management and generation systems, this invention provides an intelligent media asset management and content production method. Specific objectives include: 1. Improving media asset management efficiency: By introducing automatic media asset cleaning, automatic annotation, and structured storage, significantly improving the availability and management efficiency of media resources and reducing the burden of manual processing. 2. Enhancing material retrieval capabilities: By constructing a multimodal semantic alignment model, supporting intelligent retrieval through various methods such as text-based image search and image-based video search, achieving faster and more accurate material acquisition. 3. Improving content generation efficiency: By integrating steps such as copywriting generation, media asset retrieval and generation, speech synthesis, background music generation, and video mixing and synthesis, constructing a complete automated video mixing process, significantly reducing manual editing operations and improving content production efficiency. 4. Enhancing media asset understanding and review capabilities: Based on multimodal large-scale model technology, performing semantic-level identification and analysis of media assets, possessing the ability to initially identify and automatically review sensitive and illegal content, thus improving content security. 5. Build an intelligent content production process: Integrate natural language processing, image understanding, speech synthesis and AIGC (Artificial Intelligence Generated Content) generation capabilities to create an integrated and scalable intelligent content creation platform, and promote the upgrading of content production in the media industry towards automation and intelligence.

[0020] This invention relates to the following fields: 1. Multimodal Artificial Intelligence: used to achieve semantic alignment, intelligent retrieval, and automatic review of multimodal data such as text, images, and videos. 2. Natural Language Processing: used to automatically generate, understand, and review text, improving content production efficiency and quality. 3. Audio-Visual Synthesis: used to synthesize multimodal data such as images, videos, audio, and music into complete video content. 4. Text-to-Speech (TTS): used to convert text into video narration suitable for dubbing. 5. AIGC Technology: used to generate video footage, images, and background music, supporting automated content creation.

[0021] The intelligent media asset management and content production method described in the preferred embodiment of the present invention, such as... Figure 1 As shown, the intelligent media asset management and content production method includes the following steps: Step S10: Obtain the original media asset data, and perform basic information extraction and data cleaning processing on the original media asset data to obtain the target media asset data.

[0022] like Figure 2 As shown, this invention proposes an intelligent media asset management and content production system, which achieves efficient collaboration from media asset collection, management, retrieval to content generation through intelligent media asset management, automated content production, and full-process data review.

[0023] like Figure 3 As shown, this involves automated cleaning and storage of media asset data: collecting media asset data such as images and videos from the media industry, and using an automated cleaning module to identify and clean up resolution anomalies, data corruption, data duplication, and illegal content, thereby building a high-quality media asset database.

[0024] Specifically, the process involves acquiring raw media asset data, which includes raw media asset images and raw media asset videos; and performing basic information extraction processing on the raw media asset images and raw media asset videos to obtain basic media asset data, wherein the basic information extraction processing includes format verification processing, file type identification processing, and basic data extraction processing.

[0025] The core technology process of automated media asset cleaning can be summarized into the following main stages: 1. Media asset data source collection: Collect original images and videos (i.e., original media asset images and original media asset videos in this invention) from various media data sources (such as cultural and tourism industry media assets).

[0026] 2. Basic Information Extraction: Perform preliminary format verification, file type identification, and basic data extraction on the collected raw media data (including file name, file size, file format (such as JPEG, PNG, MP4, MOV, etc.), encoding method (such as H.264, H.265, etc.), duration, resolution, frame rate, and file hash value (MD5, SHA256, etc.), and detect corrupted, incorrectly encoded, or incomplete data.

[0027] The specific process for data corruption detection and handling is as follows: Technical principle: It utilizes technologies such as file integrity verification (e.g., MD5, CRC), encoder parsing detection, and frame or pixel anomaly detection.

[0028] The processing mechanism is as follows: File corruption: Detects issues such as corrupted file headers, incomplete downloads, and encoding errors that prevent playback or display. Data redundancy or missing data: For videos, detects the presence of numerous duplicate frames, skipped frames, black frames, or still frames. Anomaly handling: For minor corruption, attempts automatic repair (e.g., re-encoding, skipping corrupted frames); for severe corruption, direct deletion or manual processing.

[0029] The basic media asset data is cleaned to obtain the target media asset data. The data cleaning process includes deduplication, abnormal data removal, data compliance review, and data quality control.

[0030] 3. Automated cleaning and quality inspection, the specific process is as follows: Data duplication detection: Identify and eliminate completely duplicate or highly similar media assets. Resolution anomaly detection: Identify and process media assets with low resolution, high resolution overflow, or abnormal aspect ratios. Data compliance review: Utilize intelligent models to detect and process sensitive or non-compliant content. Data quality control: Remove blurry, abnormally exposed, or over-compressed media assets. High-quality media asset storage: Store media assets that pass cleaning and manual quality inspection into the database.

[0031] 4. Data duplication detection and processing, the specific process is as follows: Technical principle: Combining multiple deduplication strategies, including hash comparison, perceptual hashing, and image or video feature vector comparison.

[0032] The processing mechanism is as follows: Complete duplicate detection: By calculating the SHA256 hash value of the file, quickly identify and remove media assets that are exactly the same at the byte level, keeping only one copy.

[0033] Perceptual duplication detection: For visual content, a perceptual hash algorithm is used to generate a "fingerprint" of media assets, which can identify visually highly similar media assets even after cropping, compression or watermarking.

[0034] Feature vector comparison: Using deep learning models to extract feature vectors from images or videos, and calculating the vector space distance to determine content similarity.

[0035] Deduplication strategy: For completely duplicate media assets, only one copy is retained. For partially duplicate media assets, the highest quality version (e.g., without watermark) is manually processed and retained.

[0036] 5. Resolution anomaly detection and handling, the specific process is as follows: Technical principle: Based on image / video analysis technology, extract the resolution and aspect ratio information of media assets.

[0037] The processing mechanism is as follows: Low-resolution identification: Set thresholds (e.g., image width less than 640px, video height less than 480px) to automatically identify and label low-resolution media assets. Options include super-resolution processing (e.g., AI enhancement) or marking them as "pending review or low quality" for manual processing before direct entry into the database.

[0038] High-resolution overflow: Identify ultra-high-resolution media assets that exceed the system's processing capacity or storage specifications, and perform compression downsampling or manual processing.

[0039] Abnormal aspect ratio: Detect media assets with non-standard aspect ratios (such as 16:9, 4:3, 1:1) and crop, fill, or mark them to suit content production needs. Detect media assets with excessively large aspect ratios (such as 1:5) and clean or manually process them.

[0040] 6. Data compliance review, the specific process is as follows: Technical principle: Integrating a multimodal content security AI model, including image recognition (identifying sensitive content), video behavior analysis, OCR (text recognition) and ASR (speech recognition), combined with NLP (natural language processing) to detect security risks in the text and audio content of the video.

[0041] The processing mechanism is as follows: Categorized detection: Risk identification of sensitive or illegal content in media assets.

[0042] Risk level assessment: Based on the identification results, a risk level (high, medium, or low) is given.

[0043] Automated processing: High-risk content is directly prohibited from entering the database; medium- and low-risk content can be automatically labeled with a warning and handed over to manual processing for secondary review. This ensures that only compliant content can enter the media asset database.

[0044] Automated media asset cleaning is a preliminary step in the entire intelligent media asset library process. It can intelligently identify and process massive amounts of original media assets, including: Data quality verification and repair: automatically detecting and handling issues such as resolution anomalies, data corruption, and incomplete files; Content deduplication: accurately identifying and eliminating visually or content-duplicated media assets; and Illegal content filtering: detecting and filtering sensitive or illegal content through multimodal AI models. These functions enable high-quality, highly secure, and non-redundant media asset import, providing a clean and reliable material foundation for subsequent intelligent annotation, retrieval, and content production.

[0045] Step S20: Determine the multi-agent collaborative annotation module, and use the multi-agent collaborative annotation module to annotate the target media asset data to obtain media asset tags.

[0046] like Figure 4 As shown, automated annotation and multi-dimensional management of media assets are performed: automated annotation is performed on the cleaned media assets, including video frame extraction annotation, image content recognition, multi-dimensional annotation based on word segmentation type (general dimension and specific media industry dimension) and structured annotation, to achieve semantic and hierarchical management of media assets.

[0047] like Figure 4 As shown, the core technical process of automated media asset annotation includes the following stages: 1. High-quality media asset input: Qualified media assets from the automated media asset cleaning process are entered into the annotation process. The input includes direct image samples and video keyframe images extracted by automatic video frame extraction.

[0048] 2. Annotation Task Scheduling: The annotation task scheduler uniformly distributes images or keyframes to the multi-Agent collaborative annotation module (Agent represents intelligent agent), and assigns them to the corresponding model for processing according to the media asset type.

[0049] 3. Multi-agent collaborative annotation, including: Semantic Understanding Agent (i.e., Semantic Understanding Intelligent Agent): Calls on large-scale image understanding models or large-scale video understanding models to perform semantic parsing on media asset content and generate a set of candidate tags.

[0050] Domain calibration agent (i.e., domain calibration intelligent agent): Based on a specific domain lexicon (including general dimensions and cultural and tourism industry-specific dimensions), it corrects and supplements labels to ensure domain coverage and label accuracy.

[0051] Granularity and Synonym Specification Agent (i.e., granularity and synonym specification intelligent agent): Based on the rule base, the label is refined in granularity and synonyms are expanded to avoid overgeneralization (e.g., "skyscraper" cannot be labeled as "building").

[0052] Review Agent (i.e., review intelligent agent): Based on the quality control strategy, it performs consistency, logic and standardization checks on the tag set, and if necessary, falls back to the previous Agent for correction.

[0053] 4. Output and Storage: The approved tag set is output in JSON format to form standardized structured data containing information such as asset ID, type, and tag dimensions (general + specific), and then written to the tag library.

[0054] 5. Log recording and monitoring: Record key indicators such as time consumption, success rate, and error type during the execution of labeled tasks to facilitate subsequent performance analysis, error tracing, and system optimization.

[0055] Specifically, a multi-agent collaborative annotation module is determined, and the media asset type of the target media asset data is determined. The annotation task scheduler allocates the target media asset data to the multi-agent collaborative annotation module according to the media asset type. The semantic understanding agent calls the image understanding big model or the video understanding big model, and performs semantic parsing processing on the target media asset data through the image understanding big model or the video understanding big model to obtain candidate labels.

[0056] The specific process of the multi-agent collaborative labeling mechanism is as follows: Technical Principle: Employing a task orchestration and multi-agent collaboration model, the media asset annotation process is broken down into multiple functionally independent agents, including modules for semantic understanding, domain calibration, granularity and synonym specification, and quality review. These agents form sequential collaboration and feedback loops, allowing different agents to access different models, knowledge bases, and rules, achieving a clearly defined and complementary annotation process. This avoids the comprehension biases that may arise from using a single model in complex scenarios.

[0057] The processing mechanism is as follows: 1. The task scheduling module distributes tasks to multiple agent engines based on media asset type. 2. Each agent independently executes its own task and passes the results to the next agent.

[0058] 3. If the review agent finds label conflicts, logical errors, or insufficient confidence, it can roll back to the upstream agent to regenerate or correct the labels, forming a closed-loop process. 4. Multiple agents improve the accuracy and completeness of label generation by sharing contextual information and candidate label sets.

[0059] A domain-specific vocabulary is obtained, and the candidate labels are corrected and supplemented based on the domain-specific vocabulary by the domain calibration agent to obtain calibration labels.

[0060] The specific process of tag calibration based on domain lexicon is as follows: Technical Principle: The goal of tag calibration is to supplement and standardize industry-specific dimension tags, ensuring that media asset tagging is not only universal but also covers core elements of specific industries such as culture and tourism. In the tagging process, the semantic understanding agent first uses a multimodal large model to perform general dimension tagging on media assets (such as people, animals, natural landscapes, building types, actions, etc.). These tags belong to a "general tag set," but may lack industry-specific tags such as landmarks, activities, culture, seasons, and clothing styles. At this point, the domain calibration agent combines the generated general dimension tags with the original media asset content, and again calls upon specific models and domain vocabularies to supplement the tagging, making the tagging system complete and aligned with industry needs.

[0061] The processing mechanism is as follows: The semantic understanding agent performs general label generation on media assets (the semantic understanding agent is based on a large image understanding model, where the large image understanding model works by inputting an image and providing a prompt (e.g., "Please describe the content of the image, outputting descriptive words"), which in turn outputs the image content. While a large language model takes text as input and outputs text, the large image understanding model takes an image as input and outputs text, providing answers based on both the image and the user's question), for example, labeling "ancient buildings," "mountains," and "lakes." These general dimensional labels serve as the input basis for domain calibration.

[0062] After receiving general tags and original media images or video frames, the Domain Calibration Agent identifies missing industry-specific tags, such as "ancient city (landmark attraction)", "Water Splashing Festival (festival activity)", "Tang suit (clothing style)" and "spring (season)" through a domain vocabulary and multimodal model.

[0063] During the recognition process, the domain calibration agent uses general labels for contextual constraints. For example, if the general label is "snow mountain" and a specific terrain feature is detected, the label "a certain snow mountain" in the domain vocabulary is matched; if the general label is "street performance" and ethnic costume features are identified, the label "a certain dance" may be added.

[0064] Each item in the domain thesaurus consists of a tag category, tag name, tag description, synonym mapping, and sample image, for example: { "category": "landmark attractions" "label": "Ancient City" "description": "A famous ancient city located in a certain city of a certain province, a World Cultural Heritage site", "synonyms": ["a certain city", "a certain ancient town"], "examples": ["XX_old_town.jpg"], }

[0065] The domain terminology list is updated through regular manual review and when new industry events, attractions, or holidays are added.

[0066] A specific rule base is obtained, and the calibration label is refined in granularity and expanded in synonyms according to the specific rule base by the granularity and synonym specification agent to obtain the specification label.

[0067] The specific process for label granularity control and synonym expansion is as follows: Technical Principle: This agent receives the general tag candidate set (i.e., the calibration tags in this invention) and media content (images or keyframes) generated in the previous stage (semantic understanding or domain calibration). It refines overly coarse tags; for example, "skyscraper" and "residential building" should not be confused with "building," otherwise it will lead to irrelevant content being retrieved. Then, after the fine-grained tags are determined, common synonyms are automatically added (e.g., "woman" is supplemented with "female" or "women") to improve the retrieval recall rate and cross-language compatibility, avoiding a decrease in retrieval accuracy caused by overly coarse granularity or inconsistent expressions.

[0068] The processing mechanism is as follows: Construct a hierarchical tag system (rule base), such as "Building → High-rise Building → Skyscraper / Residential Building / Historical Building / Commercial Complex…", "People → Women / Men / Children / Elderly…", "Vehicles → Bus / Sedan / SUV…". Coarse-grained general concepts serve as parent nodes, and leaf nodes are fine-grained entities.

[0069] After receiving the tags generated in the previous stage, the Agent will check whether there are tags at the parent node (such as "building", "person", "animal") according to the tag hierarchy. If a coarse-grained tag is detected, it will combine the image / video content features of the media asset and match the leaf nodes in the tag tree to generate more fine-grained tags, such as refining "building" into "skyscraper", "residential building", and "office building".

[0070] After refining the tag generation, the Agent calls the standard tag thesaurus and synonym mapping table to find common synonyms for the tag (including Chinese, English, and industry-specific expressions). If a synonym tag exists (e.g., "woman" → "female", "women"), it is added to the tag set without replacing the original tag, thus improving the retrieval's diverse matching capabilities while maintaining semantic accuracy.

[0071] A quality control strategy is determined, and the standard label is verified by a review agent according to the quality control strategy to obtain the media asset label.

[0072] The quality review process is as follows: Technical Principle: The Quality Review Agent performs a systematic quality check on the generated tag set (i.e., the standardized tags in this invention) based on consistency detection, standardization checks, and confidence assessment. Its core idea is to utilize a rule base and tag knowledge graph to automatically verify the consistency and standardization of tags, and, when necessary, revert to the preceding Agent for regeneration or correction.

[0073] The processing mechanism is as follows: 1. Consistency check: Detects whether there are semantic conflicts or mutual exclusions between tags (such as "day" and "night" existing at the same time, "male" and "female" appearing side by side).

[0074] 2. Standardization Check: Based on the label naming conventions, check the label format and language consistency (e.g., avoid mixing Chinese and English, inconsistent capitalization, and using non-standard word forms). Check the label set for redundant punctuation marks, spaces, or encoding errors. 3. Confidence assessment and rollback mechanism: The confidence weight of the tags is calculated (combining the results of consistency check and standardization check). If the overall confidence is lower than the threshold, the rollback process is triggered, and the relevant media assets and tag set are sent back to the corresponding preceding Agent (such as granularity control Agent or domain calibration Agent) for reprocessing.

[0075] The automated media asset annotation module aims to achieve efficient and standardized tag generation for large-scale image and video media assets. Its function is to automatically parse raw media asset content into structured tag data usable for downstream retrieval and analysis through a multi-agent collaboration mechanism and a domain-specific vocabulary calibration system. Specifically, the module receives cleaned, qualified image samples and video keyframes as input, automatically schedules annotation tasks, and drives multiple agents—including semantic understanding, domain calibration, granularity and synonym expansion, and quality review—to collaborate sequentially. This generates a complete tag set that includes both general semantic dimensions (such as people, scenes, objects, and actions) and specific dimensions of the cultural and tourism industry (such as landmarks, festivals, clothing styles, and seasonal climates). At the output, the module standardizes the reviewed tag results into a JSON structure, writes it to a tag library, binds it to the media asset ID, and records log data to support performance monitoring and quality backtracking. Ultimately, the module outputs a high-quality tag system covering both general and industry-specific semantics, providing data support for subsequent multimodal retrieval.

[0076] Step S30: Encode the target media asset data and the media asset tags to obtain a visual index and a tag index, and construct a vector index library based on the visual index and the tag index.

[0077] like Figure 5 The diagram illustrates the process of multimodal retrieval of media assets: based on a multimodal large model and cross-modal semantic coding technology, it realizes multimodal retrieval methods such as text-to-image search, text-to-video search, image-to-image search, and image-to-video search, and improves retrieval accuracy and efficiency through vector model training and intra-domain fine-tuning.

[0078] The core technical process of media asset multimodal retrieval includes the following stages: 1. Media Asset Vectorization and Storage: Images and videos (with automatic keyframe extraction from videos) enter the vectorization stage: visual content is encoded using a multimodal visual vector model, and tag text is encoded using a multimodal text vector model, and then written into the visual index and tag index of the Milvus vector database, respectively.

[0079] Specifically, a multimodal vector model is determined, wherein the multimodal vector model includes a multimodal visual vector model and a multimodal text vector model; the target media asset data is subjected to multimodal visual encoding processing using the multimodal visual vector model to obtain a visual index; the media asset tags are subjected to multimodal text encoding processing using the multimodal text vector model to obtain a tag index; and a vector index library is constructed based on the visual index and the tag index.

[0080] The specific process of training and fine-tuning a multimodal vector model is as follows: Multimodal vector models jointly train image and text encoders through contrastive learning. The core objective is to make semantically related image-text pairs as close as possible in the vector space, and to keep unrelated pairs as far apart as possible. Given a set of images... and the corresponding text set Image encoder With text encoder Mapped to the same latent space respectively: ; in, For the encoded first i Image vectors, For the encoded first i A text vector, , For vector dimensions, It's a mathematical symbol that represents a real number vector. This means d A dimensional vector.

[0081] Furthermore, we calculate the similarity using cosine similarity, expressed as: ; for j A text vector, for and Cosine similarity between them.

[0082] loss function Using InfoNCE Loss (contrastive learning loss), the expression is: ; in, for and The similarity score between them, that is, the similarity between positive sample pairs (i.e., and (This is a real pair of image and text matching data). :express and The similarity score between the two is used to measure the degree of matching from the perspective of "image-to-text retrieval". :express and The similarity score between them is used to measure the degree of matching from the perspective of "text-retrieval images". For temperature coefficient, Given the total number of image-text pairs, this loss achieves image-text alignment, maximizing the similarity of the same image-text pair and distancing different pairs. The contrastive learning framework diagram is as follows: Figure 6 As shown.

[0083] To adapt to media assets in specific industries (such as cultural tourism data), intra-domain fine-tuning is introduced on the basis of pre-training. The LoRA (Low-Rank Adaptation) method is adopted, which inserts a low-rank decomposition matrix only into the model weights and performs small-scale training on specific domain samples (labeled media asset library) to achieve domain transfer effect and avoid large-scale full parameter updates.

[0084] Let the weight matrix in the pre-trained model be... In LoRA, updates are not performed directly. Instead, it represents its incremental update as The expression is: ; in, and Both are low-rank decomposition matrices. , , ,r It is the rank value of the low-rank decomposition. and All dimensions are vector-based, ensuring that incremental weights are low-rank during training and updated accordingly. A and B And frozen , The incremental update matrix representing the weights is obtained through low-rank decomposition, also known as LoRA training. This means that during training, only the weights are updated... No updates .

[0085] The schematic diagram of LoRA is as follows: Figure 7 As shown, Figure 7 In this context, h is the output vector after weight matrix transformation, X is the input feature vector (here, features are an abstract concept that exists during the calculation process), and N... For: Normal distribution, representing the matrix A Initialization method, It is the symbol for variance.

[0086] The processing mechanism is as follows: 1. **Basic Pre-training:** Train a dual-tower visual encoder (ViT, ResNet) and a text encoder (Transformer or BERT) on millions of general image-text pairs. 2. **Domain Data Preparation:** Extract image / video frames and corresponding labels from a media asset library to construct domain-specific image-text pairs. 3. **Intra-Domain Fine-tuning:** Freeze the backbone parameters and only use LoRA to adapt the weights of the attention layer during training. 4. **Vectorized Output:** Obtain a unified vector representation space, supporting cross-modal retrieval such as text → image or video, and image → image or video.

[0087] Step S40: Obtain the user query statement and perform media asset retrieval in the vector index database according to the user query statement to obtain the media asset retrieval results.

[0088] The media asset retrieval process is as follows: 1. User query vectorization: After the user inputs text or image queries, the data is encoded into vectors using the corresponding multimodal model. 2. Dual-channel retrieval: For text queries, the text vector is compared and retrieved against the visual index and tag index of the vector index library. For image queries, the visual vector is also compared and retrieved against the visual index and tag index of the vector index library. 3. Retrieval result integration: The dual-channel retrieval results are output in parallel. The front end returns results in two paths: "retrieval results based on text tags" and "retrieval results based on visual content," ensuring clear source of results and facilitating user understanding and selection.

[0089] Specifically, the process involves obtaining a user query statement and encoding it to obtain a target query vector, wherein the user query statement includes a text query statement or an image query statement; using an approximate nearest neighbor search method, the target query vector is compared and retrieved with the visual index and the tag index in the vector index library to obtain multiple image matching results or multiple video matching results; the similarity between the multiple image matching results or multiple video matching results and the user query statement is calculated to obtain a similarity result; based on the similarity result, a preset number of target image matching results or target video matching results are selected from the multiple image matching results or multiple video matching results in descending order to obtain media asset retrieval results.

[0090] The specific process of vector retrieval and indexing mechanisms is as follows: Technical Principle: Vector retrieval relies on efficient Approximate Nearest Neighbor (ANN) search. Milvus is chosen as the vector database, whose core mechanisms include support for efficient indexing methods such as IVF (Inverted File), HNSW (Hierarchical Navigable Small World), and PQ (Product Quantization) to achieve rapid retrieval of massive vectors. In the recall phase, the ANN algorithm is used to quickly locate the candidate set, and then precise calculations are performed to obtain the TopK similarity results.

[0091] The processing mechanism is as follows: 1. Index Construction: Indexes are built for both visual vectors and label vectors, using an IVF+PQ structure to balance accuracy and efficiency. 2. Query Retrieval: After the user input is encoded as a vector, it enters the corresponding index channel for ANN search. 3. Top-K Return: The top-ranked candidate results with the highest similarity are retrieved. 4. Dual-Channel Integration: The label channel results and visual channel results are output in parallel, achieving a dual-path return of "text retrieval of text labels + text retrieval of visual content".

[0092] This invention addresses the multimodal retrieval needs of large-scale media asset libraries, focusing on constructing a unified vectorized representation system and an efficient vector index structure. Specifically, it first uses a multimodal visual vector model trained with a contrastive learning strategy to encode the visual content of media assets (images, video keyframes). Simultaneously, it uses a multimodal text vector model to encode the labeled text, and improves retrieval performance in specific fields such as culture and tourism through domain-specific fine-tuning (e.g., LoRA). Then, it establishes a high-dimensional index based on vector databases such as Milvus, uniformly storing both types of vectors (visual vectors and label vectors) and constructing an efficient retrieval path. This allows for dual-path retrieval results when users input text or images: one type is based on label semantic matching, ensuring precise conceptual recall, such as "skyscraper" or "ancient city wall"; the other type is based on visual feature matching, ensuring recall based on similar styles or scenes, such as similar-looking buildings or visual styles. Ultimately, this invention achieves unified retrieval capabilities for text-based image or video search and image-based image or video search, with output retrieval results covering both semantic and visual dimensions, thereby improving the accuracy, robustness, and interpretability of the retrieval.

[0093] Step S50: Obtain the synthesized data, and perform time-series alignment processing on the synthesized data and the media asset retrieval results to obtain the target synthesized audio and video.

[0094] like Figure 8 As shown, the specific process of the target synthesized audio and video content production workflow is as follows: 1. Image-Video Composition Paradigm: Through text generation, speech synthesis, text-based media asset retrieval, background music (BGM) generation, and audio-video mixing, image and video materials are spliced ​​together to form a video. 2. Digital Human Voiceover Video Paradigm: Through text generation, speech synthesis, voice-driven digital human video generation, background music (BGM) generation, and audio-video mixing, digital human voiceover videos are output.

[0095] Specifically, the user query is input into a preset large language model to obtain the narration script, and the narration script is input into a TTS model to obtain the narration voice-over; the user query is input into a music generation model to obtain background music; the narration voice-over, the background music, and the media asset retrieval results are aligned and superimposed on the same timeline to obtain the target synthesized audio and video.

[0096] The specific process for generating target synthesized audio and video is as follows: like Figure 8 As shown, the process of editing a video into a final product includes: 1. **Script Generation:** Based on a Large Language Model (LLM), inputting a topic or keywords automatically generates a spoken script. 2. **Speech Synthesis:** Using a Text-to-Speech (TTS) model, the script (spoken script) is converted into a natural and fluent narration. Optional "voice cloning" technology allows for customized voice timbre (e.g., specifying a particular person's voice). 3. **Media Asset Retrieval:** A multimodal retrieval module is invoked to automatically retrieve matching image or video materials (i.e., media asset retrieval results in this invention) based on the script content (spoken script), and these are then stitched together into a sequence of segments along a timeline. 4. **Background Music Generation:** Using a music generation model, background music is generated using natural language (e.g., inputting "upbeat BGM"). 5. **Audio-Video Synthesis:** A synthesis module (based on the ffmpeg framework, a multimedia processing framework) is invoked to align images, video clips, narration, and background music in sequence, and render subtitles, on-screen text, and transition effects to output a complete video.

[0097] like Figure 8 As shown, the digital human video generation process is as follows: 1. Script Generation: LLM is used to generate the narration script. 2. Speech Synthesis: Voiceover is synthesized through a TTS model, supporting voice cloning. 3. Digital Human Video Generation: Based on the digital human generation model, the synthesized speech is input to drive the digital human's lip movements, expressions, and actions to generate continuous videos. 4. Media Asset Supplementation and BGM Generation: Retrieved images or video materials can be inserted into the digital human's narration segments, and matching BGM is generated. 5. Audio-Video Synthesis: Based on ffmpeg, the digital human video, material clips, BGM, subtitles, and on-screen text are synthesized in a unified time sequence to generate the final video.

[0098] The temporal composition process is as follows: Unified temporal composition refers to aligning and overlaying media elements of different modalities (video clips, digital human videos, background music, narration, subtitles, on-screen effects, etc.) along the same timeline to generate a complete multi-track video. The specific process is as follows: 1. The narration serves as the baseline timeline, determining the start and end times of each segment. 2. Corresponding source video or digital human video clips are inserted into the same time period. 3. Subtitles and on-screen effects are automatically aligned according to the timestamps of the narration. 4. Using ffmpeg's filter_complex function, the above multi-track content is mixed and overlaid on the same timeline, achieving synchronized output of audio, video, subtitles, and effects, ultimately generating a video file in a unified format.

[0099] The technical principles behind digital human video generation are as follows: The digital human avatar consists of two stages: image training and audio-driven generation. 1. Image Training: Inputting a reference video, the system models the target person's facial features, facial expressions, and lip movements to generate a dynamic digital avatar. This is achieved using Neural Rendering technology. 2. Audio-Driven Generation: Inputting narration audio, the system first extracts speech features, generates a lip-sync sequence using a lip-sync prediction model, and then combines this with the trained digital avatar to render facial expressions and head movements. Finally, a video clip highly matched to the audio sequence is synthesized.

[0100] The technical principles of audio-visual compositing (based on ffmpeg) are as follows: 1. ffmpeg provides powerful audio and video encoding, decoding, and processing capabilities, enabling the splicing of multi-source materials, filter rendering, subtitle overlay, and audio track mixing. 2. Video splicing and transitions: Retrieved image / video clips and digital human videos are spliced ​​along a timeline (based on SRT subtitle files), and natural transitions are achieved through filters. 3. Audio track compositing: The narration is time-aligned with the generated background music. 4. Subtitle and on-screen text rendering: SRT subtitle files are automatically generated based on the text and overlaid onto the video track using ffmpeg. 5. Export optimization: Video resolution, frame rate, and bitrate are unified to ensure playback output quality.

[0101] like Figure 9 As shown, the present invention also sets up a full-process data review: at each stage of media asset entry and content production, a multimodal review model is used to detect and filter sensitive information or illegal content to ensure the compliance and security of media assets and generated content.

[0102] The technical process is as follows: 1. Multimodal data input: The system supports the review of four types of data: text, audio, images, and video, originating from media asset storage and content production stages, respectively. 2. Modal preprocessing: Text: Directly input into the NLP review model. Audio: Transcribed into text using an ASR model, then input into the NLP review model, while simultaneously using audio feature detection to identify high-risk events. Images: Extracting text from images using OCR and inputting it into the text review, while simultaneously using an image multimodal recognition model to detect sensitive or illegal data. Video: Keyframe extraction is performed, the image is input into the image review model, and the audio is input into the audio review model, supporting similarity retrieval between keyframes and the violation database.

[0103] 3. Multi-dimensional Violation Detection: Pre-processed data is uniformly entered into a multi-dimensional detection module, covering various sensitive and illegal data. False Information Detection: Factual verification of text / images is performed, comparing them with an authoritative knowledge base. 4. Filtering and Output: Compliant samples passing the above review dimensions enter the subsequent media asset storage and content production process. Non-compliant samples are intercepted, marked, or stored in the violation database. The final output is the media assets or generated content after compliance filtering.

[0104] The data security review module's function and output are to provide compliance and security guarantees for the entire process of media asset entry and content production. Its main role is to conduct comprehensive detection and filtering of text, audio, images and videos through a multimodal review model. It uses technologies such as OCR, ASR, image recognition, and semantic comparison retrieval to automatically identify various sensitive and illegal data, and to intercept or roll back when necessary, ensuring that illegal content does not enter the media asset library or is used to generate videos.

[0105] This invention fully integrates multimodal large models, AIGC technology, and intelligent review capabilities, significantly improving the intelligence level of media asset management and the automation efficiency of content production, providing the media industry with an efficient, accurate, and secure content generation solution.

[0106] Technical effects: 1. Improve media asset management efficiency: Through automatic media asset cleaning, automatic labeling and structured storage, the problems of chaos and inefficiency caused by traditional manual sorting are solved, and high-quality management and unified access to resources are achieved.

[0107] 2. Enhanced material retrieval capabilities: Based on multimodal vector models and vector retrieval technology, it supports cross-modal retrieval methods such as text-to-image search, image-to-video search, and image-to-image search, significantly improving the accuracy and efficiency of material retrieval.

[0108] 3. Improve content production efficiency: By integrating technologies such as copywriting generation, speech synthesis, background music generation, video editing, and digital human generation, a high degree of automation is achieved from media asset retrieval to video production, significantly shortening the content production cycle and reducing manual operation costs.

[0109] 4. Ensure content compliance and security: Utilize a multimodal review model to perform multi-dimensional detection and filtering of various sensitive and illegal data in text, audio, images, and videos to ensure the compliance and security of media assets entering the database and generated content.

[0110] 5. Promote the intelligent upgrading of the media industry: Integrate natural language processing, image understanding, speech synthesis and AIGC generation capabilities to build an integrated and scalable intelligent media asset and content production platform, and promote the transformation of the media industry from manual to intelligent and automated.

[0111] This invention proposes an intelligent media asset management and content production method, which has the following innovative features and points to be protected: 1. A pioneering multimodal content production and media asset management platform in the media industry: This invention is the first in the industry to build a closed-loop system integrating automatic media asset cleaning, structured annotation, multimodal semantic retrieval, automated content production, and multimodal security review. It breaks through the limitations of traditional systems such as "fragmented processing and high dependence on manual labor" and realizes full-process automation from media asset storage to compliant content output.

[0112] 2. Multi-Agent Collaborative Annotation Mechanism: It innovatively proposes a multi-agent collaborative annotation process, including semantic understanding, domain calibration, granularity or synonym specification, and review. By combining large models with rules, it solves the problems of semantic inconsistency, chaotic granularity, and poor domain adaptability in traditional media asset annotation, and significantly improves the quality and standardization of media asset tags.

[0113] 3. Multimodal Semantic Vector Retrieval and Intra-Domain Fine-tuning: This invention employs a CLIP-like semantic alignment training framework in multimodal retrieval, combined with LoRA for intra-domain fine-tuning. It optimizes for media industry-specific corpora and promotional materials, thereby enhancing the model's understanding of industry terminology and contextual style, achieving high-precision semantic retrieval between text, images, and videos. Compared to general models, this invention is better adapted to media scenarios in media asset retrieval.

[0114] 4. Intelligent content generation process: It proposes two processes: "image-video hybrid editing" and "digital human video generation". For the first time, it combines text generation, TTS speech synthesis, BGM generation, video synthesis technology with digital human driving technology and applies it to the media industry, forming an integrated content production solution for the publicity and media industry, which greatly improves the production efficiency of short videos and promotional videos.

[0115] 5. Multimodal Compliance Review Model: This invention constructs a multimodal review mechanism covering text, audio, images, and videos, integrating technologies such as OCR, speech recognition, visual understanding, and comparative retrieval to detect and filter risky content such as various sensitive and illegal data. For the first time, it achieves full-chain security review within the same system, effectively ensuring the compliance and security of content production.

[0116] Furthermore, such as Figure 10 As shown, based on the above-described intelligent media asset management and content production method, the present invention also provides an intelligent media asset management and content production system, wherein the intelligent media asset management and content production system includes: The data cleaning module 51 is used to acquire raw media asset data and perform basic information extraction and data cleaning processing on the raw media asset data to obtain target media asset data. Data annotation module 52 is used to determine the multi-agent collaborative annotation module, and to perform annotation processing on the target media asset data through the multi-agent collaborative annotation module to obtain media asset tags; The index building module 53 is used to encode the target media asset data and the media asset tags to obtain a visual index and a tag index, and to build a vector index library based on the visual index and the tag index. The media asset retrieval module 54 is used to obtain the user's query statement and perform media asset retrieval in the vector index library according to the user's query statement to obtain the media asset retrieval results; The video output module 55 is used to acquire the synthesized data and perform time-series alignment processing on the synthesized data and the media asset retrieval results to obtain the target synthesized audio and video.

[0117] Furthermore, such as Figure 11 As shown, based on the above-mentioned intelligent media asset management and content production method and system, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 11 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0118] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores an intelligent media asset management and content production program 40, which can be executed by the processor 10 to implement the intelligent media asset management and content production method of this application.

[0119] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the intelligent media asset management and content production method.

[0120] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface.

[0121] In one embodiment, the steps of the intelligent media asset management and content production method are implemented when the processor 10 executes the intelligent media asset management and content production program 40 in the memory 20.

[0122] In summary, this invention provides an intelligent media asset management and content production method, system, and terminal. The method includes: acquiring raw media asset data, performing basic information extraction and data cleaning on the raw media asset data to obtain target media asset data; determining a multi-agent collaborative annotation module, and using the multi-agent collaborative annotation module to annotate the target media asset data to obtain media asset tags; encoding the target media asset data and the media asset tags to obtain a visual index and a tag index, and constructing a vector index library based on the visual index and the tag index; acquiring a user query statement, and performing media asset retrieval in the vector index library based on the user query statement to obtain media asset retrieval results; acquiring synthesized data, and performing time-series alignment processing on the synthesized data and the media asset retrieval results to obtain target synthesized audio and video. This invention, by cleaning and annotating the raw media asset data to construct a vector index library, can effectively improve the efficiency of user management and retrieval of media asset data. Furthermore, by performing time-series alignment processing on the synthesized data and the media asset retrieval results, it can achieve automatic audio and video synthesis, significantly reducing manual editing operations and improving the production efficiency of audio and video content.

[0123] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.

[0124] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0125] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A method for intelligent media asset management and content production, characterized in that, The intelligent media asset management and content production method includes: Obtain raw media asset data, and perform basic information extraction and data cleaning processing on the raw media asset data to obtain target media asset data; A multi-agent collaborative annotation module is identified, and the target media asset data is annotated using the multi-agent collaborative annotation module to obtain media asset tags; The target media asset data and the media asset tags are encoded to obtain a visual index and a tag index, and a vector index library is constructed based on the visual index and the tag index. Obtain the user's query statement, and perform media asset retrieval in the vector index database based on the user's query statement to obtain the media asset retrieval results; Acquire the synthesized data, and perform time-series alignment processing on the synthesized data and the media asset retrieval results to obtain the target synthesized audio and video.

2. The intelligent media asset management and content production method according to claim 1, characterized in that, The process of acquiring raw media asset data and performing basic information extraction and data cleaning on the raw media asset data to obtain target media asset data specifically includes: Obtain raw media asset data, wherein the raw media asset data includes raw media asset images and raw media asset videos; The original media asset images and original media asset videos are subjected to basic information extraction processing to obtain basic media asset data. The basic information extraction processing includes format verification processing, file type identification processing, and basic data extraction processing. The basic media asset data is cleaned to obtain the target media asset data. The data cleaning process includes deduplication, abnormal data removal, data compliance review, and data quality control.

3. The intelligent media asset management and content production method according to claim 1, characterized in that, The step of determining a multi-agent collaborative annotation module and using this module to annotate the target media asset data to obtain media asset tags specifically includes: Identify the multi-agent collaborative annotation module and determine the media asset type of the target media asset data; The task scheduler allocates the target media asset data to the multi-agent collaborative annotation module according to the media asset type, and the multi-agent collaborative annotation module performs annotation processing on the target media asset data to obtain media asset tags.

4. The intelligent media asset management and content production method according to claim 3, characterized in that, The multi-agent collaborative annotation module includes a semantic understanding agent, a domain calibration agent, a granularity and synonym specification agent, and a review agent. The step of annotating the target media asset data through the multi-agent collaborative annotation module to obtain media asset tags specifically includes: The semantic understanding agent invokes the image understanding big model or the video understanding big model, and performs semantic parsing processing on the target media asset data through the image understanding big model or the video understanding big model to obtain candidate tags; A domain-specific vocabulary is obtained, and the candidate labels are corrected and supplemented based on the domain-specific vocabulary by the domain calibration agent to obtain calibration labels. A specific rule base is obtained, and the calibration label is refined in granularity and expanded in synonyms according to the specific rule base by the granularity and synonym specification agent to obtain the specification label; A quality control strategy is determined, and the standard label is verified by a review agent according to the quality control strategy to obtain the media asset label.

5. The intelligent media asset management and content production method according to claim 1, characterized in that, The process of encoding the target media asset data and the media asset tags to obtain a visual index and a tag index, and constructing a vector index library based on the visual index and the tag index, specifically includes: A multimodal vector model is determined, wherein the multimodal vector model includes a multimodal visual vector model and a multimodal text vector model; The target media asset data is processed by multimodal visual encoding using the multimodal visual vector model to obtain a visual index; The media asset tags are processed using the multimodal text vector model to perform multimodal text encoding to obtain the tag index; A vector index library is constructed based on the visual index and the label index.

6. The intelligent media asset management and content production method according to claim 1, characterized in that, The process of obtaining the user's query statement and performing media asset retrieval in the vector index based on the user's query statement to obtain the media asset retrieval results specifically includes: Obtain the user query statement and encode the user query statement to obtain the target query vector, wherein the user query statement includes a text query statement or an image query statement; The target query vector is compared and retrieved with the visual index and the tag index in the vector index library using the approximate nearest neighbor search method to obtain multiple image matching results or multiple video matching results; Calculate the similarity between multiple image matching results or multiple video matching results and the user query statement to obtain a similarity result; Based on the similarity results, a preset number of target image matching results or target video matching results are selected from multiple image matching results or multiple video matching results in descending order to obtain media asset retrieval results.

7. The intelligent media asset management and content production method according to claim 1, characterized in that, The synthesized data includes narration and background music; The step of acquiring synthetic data and performing time-series alignment processing on the synthetic data and the media asset retrieval results to obtain the target synthetic audio and video specifically includes: The user query is input into a preset large language model to obtain the spoken text, and the spoken text is input into a TTS model to obtain the narration and voice-over. The user query is input into the music generation model to obtain background music; The narration, background music, and media asset retrieval results are aligned and overlaid along the same timeline to obtain the target synthesized audio and video.

8. An intelligent media asset management and content production system, characterized in that, The intelligent media asset management and content production system includes: The data cleaning module is used to acquire raw media asset data and perform basic information extraction and data cleaning processing on the raw media asset data to obtain target media asset data. A data annotation module is used to determine a multi-agent collaborative annotation module and to annotate the target media asset data through the multi-agent collaborative annotation module to obtain media asset tags. An index building module is used to encode the target media asset data and the media asset tags to obtain a visual index and a tag index, and to build a vector index library based on the visual index and the tag index. The media asset retrieval module is used to obtain the user's query statement and perform media asset retrieval in the vector index database according to the user's query statement to obtain the media asset retrieval results; The video output module is used to acquire the synthesized data and perform time-series alignment processing on the synthesized data and the media asset retrieval results to obtain the target synthesized audio and video.

9. A terminal, characterized in that, The terminal includes: a memory, a processor, and an intelligent media asset management and content production program stored in the memory and executable on the processor. When the intelligent media asset management and content production program is executed by the processor, it implements the steps of the intelligent media asset management and content production method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an intelligent media asset management and content production program, which, when executed by a processor, implements the steps of the intelligent media asset management and content production method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Video processing method, computing device, computer storage medium and computer program product

    CN118972671A

  • All-media fusion method and system based on complementary fusion

    CN120123970A

  • Supervisory control and data acquisition (SCADA) system label generation method and system based on structured multi-source information fusion

    CN120144567A

  • Multi-modal comment data automatic annotation and classification method for software development

    CN120654118A