Intelligent multimedia data management method and device based on fusion controlled text generation

By adopting a method based on converged controlled text generation in the multimedia data management system, the problems of redundancy of tags and inaccurate search results in multimodal data management are solved, and more efficient and accurate multimedia data management is achieved.

CN120123533AInactive Publication Date: 2025-06-10北京中电慧声科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510622378.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the existing multimedia data management system processes multimodal data, the tags generated automatically are redundant and unstructured, resulting in insufficient accuracy and credibility of search results, and the search results are scattered, making it difficult to meet management needs.

Method used

An intelligent multimedia data management method based on fused controlled text generation is adopted. By receiving multimedia data, audio and view text content is extracted and structured fragment index data is formed, multiple language models are input to generate controlled text, and data labels are generated based on controlled text for management.

Benefits of technology

It improves the accuracy and credibility of search results in multimedia data management, reduces the dispersion of search results, and enhances the efficiency of data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123533A_ABST
    Figure CN120123533A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent multimedia data management method and device based on fusion controlled text generation, and relates to the technical field of multimedia data processing. The method comprises the following steps: receiving multimedia data, extracting audio and video image-text contents of the multimedia data, and storing the audio and video image-text contents according to a preset template to form structured fragment index data; inputting the structured fragment index data into a plurality of language models, generating texts according to the plurality of models, fusing the texts to form a controlled text, and generating a data label of the multimedia data according to the controlled text; and performing multimedia data management according to the data label. On the basis of automatically analyzing the multimedia data through an artificial intelligence algorithm, controlled text generation is introduced, and the information extracted from the multimedia data by artificial intelligence is processed into a unified standardized format, so that the search result precision and credibility are improved during multimedia data management, the search results are more aggregated and accurate, and the search efficiency is improved. And the multimedia data management efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimedia data processing, and particularly to an intelligent multimedia data management method, device, equipment and medium based on the generation of fused controlled text. Background Art

[0002] With the explosive growth of multimedia data, traditional multimedia data management systems mainly rely on the mode of manual annotation of tags, with low efficiency; especially when facing long-duration audio and video data, the integrity and accuracy of the manual mode decline, resulting in limited efficiency and accuracy of subsequent data search.

[0003] In the prior art known to the inventors, some multimedia data management solutions introduce artificial intelligence algorithms to automatically extract text information from multimedia data and generate preliminary tags. However, when dealing with multi-modal data, the tags automatically generated by general artificial intelligence algorithms have prominent problems of redundancy and unstructuredness. It is easy to miss key information, have logical errors or generate incoherent text expressions due to context understanding deviation or semantic ambiguity, and it is also easy to produce wrong inference conclusions, resulting in insufficient accuracy and credibility of system search results, scattered search results, and difficulty in meeting management requirements. Summary of the Invention

[0004] The present invention provides an intelligent multimedia data management method, device, equipment and medium based on the generation of fused controlled text, which solves the problems of insufficient accuracy and credibility of search results and scattered search results in multimedia data management.

[0005] To achieve the above object, the present application adopts the following technical solutions:

[0006] In a first aspect, there is provided an intelligent multimedia data management method based on the generation of fused controlled text, including:

[0007] Receiving multimedia data, extracting the audio, video and text content of the multimedia data and storing it according to a preset template to form structured fragment index data;

[0008] Inputting the structured fragment index data into multiple language models, generating text by multiple models and fusing it to form controlled text, and generating data tags of the multimedia data according to the controlled text;

[0009] Managing multimedia data according to the data tags.

[0010] In a first possible implementation manner of the first aspect, the receiving multimedia data, extracting the audio, video and text content of the multimedia data and storing it according to a preset template to form structured fragment index data specifically includes:

[0011] For audio data: extracting text information;

[0012] For video data: Separate the audio from it to extract text information; Identify the information in the frame images and convert it into text information; For the identified person and object information, capture the frame images and label the positioning tags and timestamps.

[0013] For picture data, identify the information in the image and convert it into text information.

[0014] In the second possible implementation manner of the first aspect, inputting the structured segment index data into multiple language models, generating text according to multiple models, fusing the generated text to form a controlled text, and generating a data label for the multimedia data according to the controlled text specifically includes:

[0015] Based on the historical performance of the models, assign weights to the trustworthy text features for the candidate texts generated by each model based on the decision tree algorithm;

[0016] Based on the decision tree algorithm, score the trustworthy text features based on the indicators of text smoothness, keyword coverage, key information fitting degree, and text information mixing degree;

[0017] Based on the results of the trustworthy text feature weighting and decision tree scoring, weighted-fuse the high-scoring segments to form a preliminary optimized text;

[0018] Based on the positioning tags and timestamps, align the preliminary optimized text with the original multimedia data through feature alignment technology, and generate the controlled text according to a preset same Prompt template.

[0019] In the third possible implementation manner of the first aspect, the multiple language models include models with different architectures, models in different training stages, and / or models trained based on different training sets.

[0020] In the second aspect, there is provided an intelligent multimedia data management device based on the generation of a fused controlled text, including:

[0021] A structured segment index data generation module, configured to receive multimedia data, extract the audio, video, and text content of the multimedia data and store it according to a preset template to form structured segment index data;

[0022] A controlled text and data label generation module, configured to input the structured segment index data into multiple language models, generate text according to multiple models, fuse the generated text to form a controlled text, and generate a data label for the multimedia data according to the controlled text;

[0023] A data management module, configured to manage the multimedia data according to the data label.

[0024] In the first possible implementation of the second aspect, the structured fragment index data generation module is specifically configured to:

[0025] For audio data: extract text information;

[0026] For video data: separate the audio therein and extract text information; identify information in the frame images and convert it into text information; for the identified person and object information, intercept the frame pictures and label positioning tags and timestamps;

[0027] For picture data, identify information in the image and convert it into text information.

[0028] In the second possible implementation of the second aspect, the controlled text and data label generation module is specifically configured to:

[0029] Based on the historical performance of the model, assign weights to the credible text features for the candidate texts generated by each model based on the decision tree algorithm;

[0030] Based on the decision tree algorithm, score the credible text features based on the indicators of text smoothness, keyword coverage, key information fitting degree, and text information mixing degree;

[0031] Based on the results of the credible text feature weighting and decision tree scoring, weighted-fuse the high-scoring segments to form a preliminary optimized text;

[0032] Based on the positioning tags and timestamps, align the preliminary optimized text with the original multimedia data through feature alignment technology, and generate the controlled text according to a preset same Prompt template.

[0033] In the third possible implementation of the second aspect, the multiple language models include models with different architectures, models in different training stages, and / or models trained based on different training sets.

[0034] In a third aspect, an electronic device is provided, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the intelligent multimedia data management method based on the fusion of controlled text as described in the first aspect are implemented.

[0035] In a fourth aspect, a readable storage medium is provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the intelligent multimedia data management method based on the fusion of controlled text as described in the first aspect are implemented.

[0036] The intelligent multimedia data management method based on the fusion of controlled text of the present invention has the following beneficial effects:

[0037] This application forms structured fragment index data for intelligent retrieval from audio, video, and text content, adopts a multi-model text generation technical route of controlled text generation technology, effectively organizes the information contained in the structured fragment index data, generates data tags, and constructs a perfect fragmented tag system, overcoming problems existing in traditional natural language processing methods such as missing key information, unsmooth sentence expressions, and incorrect inference conclusions; based on automatically analyzing multimedia data through artificial intelligence algorithms, this application introduces controlled text generation to process the information extracted by artificial intelligence from multimedia data into a unified standardized format, improving the accuracy and credibility of search results during multimedia data management, making the search results more concentrated and accurate, and improving the efficiency of multimedia data management.

[0038] The device, electronic device, and readable storage medium corresponding to the intelligent multimedia data management method based on the fusion of controlled text generation of the present invention can achieve the same technical effects, and will not be elaborated here to avoid repetition. Brief Description of the Drawings

[0039] Figure 1 It is a schematic flowchart of an intelligent multimedia data management method based on the fusion of controlled text generation provided by an embodiment of this application;

[0040] Figure 2 It is a schematic flowchart of another intelligent multimedia data management method based on the fusion of controlled text generation provided by an embodiment of this application;

[0041] Figure 3 It is a schematic diagram of the process of fusing multi-model generated texts to form controlled texts provided by an embodiment of this application;

[0042] Figure 4 It is a schematic structural diagram of an intelligent multimedia data management system based on the fusion of controlled text generation provided by an embodiment of this application;

[0043] Figure 5 It is a schematic structural diagram of an intelligent multimedia data management device based on the fusion of controlled text generation provided by an embodiment of this application;

[0044] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of this application;

[0045] Figure 7 It is a fine-tuning effect curve graph provided by an embodiment of this application. Detailed Description of the Embodiment

[0046] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the technical solutions in the embodiments of the present application are clearly described. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0047] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0048] In the description of the method process in the specification of the present application and the steps in the flowchart in the accompanying drawings of the present invention, it is not necessary to strictly execute according to the step numbers. The execution order of the method steps can be changed. Moreover, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

[0049] The following will be described in detail with reference to the accompanying drawings and preferred embodiments the intelligent multimedia data management method, device, equipment and medium provided by the embodiments of the present application.

[0050] Traditional multimedia data management systems often need to manage multimedia data by manually adding tags. This management method has two disadvantages. Manually adding tags often requires a large amount of human resources and is inefficient. Managers need to manually listen to, read, and view files such as text, voice, and video to identify the information contained in the media, and then summarize the media content. Especially for some multimedia data with a long duration, manual analysis takes too long and is often infeasible.

[0051] In recent years, some artificial intelligence algorithms have been introduced into multimedia data management for re-editing tags. However, this method requires processing a large amount of text data. After extracting the multimedia data information through artificial intelligence methods, traditional natural language processing methods often have poor effects, such as missing key information, unsmooth sentence expressions, and incorrect inference conclusions, and cannot present the data in a unified format. Managers often cannot comprehensively cover and summarize the media information with a large amount of content, resulting in incomplete search results when searching for multimedia data based on tags.

[0052] Based on this, the present application proposes an intelligent multimedia data management system based on the generation of fused controlled text. This system is developed based on the research and development of a rectification and optimization algorithm for controlled text generation under hard constraints at the model and policy levels, and designs a fusion selection and optimization algorithm for generating text by multiple models to solve problems such as information deviation and insufficient reliability in the text generation results. First, it combines the generative large model to automatically generate tags, comprehensively and effectively cover multimedia information, form high-precision management of multimedia data, and provide support for efficient data search.

[0053] Please refer to Figure 1-2 , the embodiment of the present application provides an intelligent multimedia data management method based on the generation of fused controlled text, as Figure 1-2 shown, including:

[0054] S1, Receive multimedia data, extract the audio, video, text, and image content of the multimedia data and store it according to a preset template to form structured segment index data.

[0055] Furthermore, the multimedia data includes: audio data, video data, and picture data. Specifically, this step S1 includes:

[0056] S11, For audio data: Extract the text information.

[0057] S12, For video data: Separate the audio therein and extract the text information; Based on frame image recognition, recognize the information in the image and convert it into text information; For the recognized person and object information, intercept the frame picture and mark the positioning label and timestamp.

[0058] S13, For picture data, recognize the information in the image and convert it into text information.

[0059] The audio data in the multimedia data is converted into text information by speech recognition; The audio in the video data extracts the text information contained in the speech by speech recognition, extracts the person information contained therein by face recognition, extracts the text information contained in the image by optical character recognition, and converts the information in the image into text information by object recognition and image recognition; The picture data also converts the information contained therein into text information by face recognition, optical character recognition, object recognition, and image recognition; Finally, the information in the multimedia data will be uniformly input into the controlled text generation system in the form of text information to generate data tags for data management.

[0060] S2, Input the structured segment index data into multiple language models, fuse the texts generated by multiple models to form controlled text, and generate data tags for the multimedia data according to the controlled text.

[0061] Generate text fusion based on multiple models to form controlled text. The controlled text generation technology adopts the technical route of multi-model text generation. Multi-model text generation refers to using a variety of different pre-trained models or algorithms to generate text, and then fusing and selecting these generated texts to obtain better generation results. Among them, the multiple language models can be generative models and / or discriminative models with different architectures, or models at different training stages, or models obtained by training based on different training sets.

[0062] Further, step S2 specifically includes:

[0063] S21, based on the historical performance of the model, assign weights to the trustworthy text features for the candidate texts generated by each model based on the decision tree algorithm.

[0064] This step can dynamically assign weights according to the historical accuracy and recall rate of the model, and preferentially select the generated content of high-trust models; it can perform feature selection and weighting on multiple features of the text through the gradient boosting tree algorithm combined with XGBoost, so as to select high-quality generated texts according to the size of the weighted average. Gradient Boosting Decision Tree (XGBoost) is an efficient gradient boosting decision tree algorithm. It is improved on the basis of the original GBDT, greatly improving the model effect.

[0065] S22, based on the decision tree algorithm, score the trustworthy text features based on the indicators of text smoothness, keyword coverage, key information fitting degree, and text information mixing degree.

[0066] Specifically, the exemplary code implementation of text smoothness is as follows:

[0067] def compute_ppl(text, keywords):

[0068] for keyword in keywords:

[0069] if keyword:

[0070] lac.add_word(keyword)

[0071] text = text.strip('\n')

[0072] segs_row = lac.run(text)

[0073] segs_result = [x for x in segs_row if not x inpplhander.stopwords]

[0074] sentence = ' '.join(segs_result)

[0075] res = re.sub(r'[。、!,]', pplhander.replace_chars, sentence)

[0076] res = re.sub(r"^,\s*", "", res)

[0077] res = re.sub(r"^。\s*", "", res)

[0078] res = res.replace('。 ,', '。')

[0079] ppl = pplhander.corrector_instance.ppl_score(res.split(' '))

[0080] return ppl

[0081] The n-gram language model is used to evaluate the smoothness of the text. Its core function is to predict whether a sequence of words in a text conforms to the habitual usage of the language. The model first learns the probability distribution of the word sequence in the text, and the n-gram model scores the given text to judge the relevance between the concerned words and whether the word combination conforms to the grammar rules and semantic logic.

[0082] The exemplary code implementation of the keyword coverage is as follows:

[0083] def compute_coverage(text: str, keywords: list):

[0084] for keyword in keywords:

[0085] if keyword:

[0086] lac.add_word(keyword)

[0087] segs_row = lac.run(text)

[0088] seg_list = [x for x in segs_row if not x in pplhander.stopwords]

[0089] keywords_count = dict(Counter(keywords))

[0090] seg_list_count = dict(Counter(seg_list))

[0091] sub_recall = 0

[0092] precision_numerate = 0

[0093] for k, v in keywords_count.items():

[0094] if k in seg_list_count:

[0095] sub_recall += min(seg_list_count[k], v)

[0096] precision_numerate += 1

[0097] precision = precision_numerate / len(set(keywords_count))

[0098] recall = sub_recall / len(keywords)

[0099] f1_score = 2 * precision * recall / (precision + recall)

[0100] return f1_score

[0101] Since the text generation result of this algorithm is subject to hard constraints and is text generated under control. This application scores the keyword coverage by calculating the occurrence frequency and distribution of keywords, thereby ensuring the quality of text generation.

[0102] The exemplary code implementation of the key information fitting degree is as follows:

[0103] def text_rank(sentences):

[0104] # Word segmentation

[0105] words = sum([word.lower() for word in re.findall(r'\b\w+\b', sentence)] for sentence in sentences)

[0106] unique_words = list(set(words))

[0107] # Build the graph

[0108] graph = {word: [] for word in unique_words}

[0109] for sentence in sentences:

[0110] for i in range(len(sentence)):

[0111] for j in range(i + 1, len(sentence)):

[0112] word1 = sentence[i].lower()

[0113] word2 = sentence[j].lower()

[0114] if word1 in graph and word2 in graph:

[0115] graph[word1].append(word2)

[0116] graph[word2].append(word1)

[0117] # Calculate weights

[0118] ranks = {word: 1 for word in unique_words} # Initialize the rank of all nodes to 1

[0119] for word in unique_words:

[0120] if graph[word]: # If the node has neighbors

[0121] ranks[word] = (1 - damping_factor) + (damping_factor / len(graph[word])) * sum(

[0122] [ranks[neighbor] for neighbor in graph[word]])

[0123] # Extract key information

[0124] sorted_words = sorted(unique_words, key=lambda x: ranks[x], reverse=True)

[0125] return sorted_words

[0126] An exemplary code implementation of the text information mixing degree is as follows:

[0127] def calculate_entropy(X):

[0128] p = np.array([np.sum(X == xi) for xi in np.unique(X)]) / len(X)

[0129] p = p[p > 0]

[0130] # Calculate the entropy value

[0131] entropy = -np.sum(p * np.log2(p))

[0132] return entropy

[0133] The algorithm extracts core information from the text through graph algorithms, assigns corresponding weights to the information, and measures the importance. For texts with higher quality, the key information revealed by the graph structure of the internal core information has a high matching degree with the key information of the original data. This algorithm calculates the matching degree between the key information extracted through graph algorithms and the key information of the original data to measure the text quality and ensure that the extracted information is as accurate as possible. In the code, first adjust the parameters of the graph algorithm, and then optimize the weight assignment to ensure that the extracted information represents the logical structure of the text and accurately grasps the core content of the text.

[0134] An exemplary code implementation of the text information mixing degree is as follows:

[0135] def calculate_entropy(X):

[0136] p = np.array([np.sum(X == xi) for xi in np.unique(X)]) / len(X)

[0137] p = p[p > 0]

[0138] # Calculate the entropy value

[0139] entropy = -np.sum(p * np.log2(p))

[0140] return entropy

[0141] This algorithm evaluates the text information mixing degree through information entropy. First, it counts the occurrence times of words, calculates the probability distribution of each word, and then calculates the entropy value. The higher the information entropy, the worse the normativity of the text in the generation of constrained controlled texts. The greater the information entropy gain of the generated text, the greater the content jump and the lower the credibility.

[0142] S23, based on the results of the trusted text feature weighting and decision tree scoring, weighted-fuse the high-scoring segments to form a preliminary optimized text.

[0143] By designing and implementing a series of evaluation algorithms, the text with the highest quality can be selected from multiple generation models to meet the text generation requirements in the field. Introduce a weight-based optimization method for selecting generation results, design multiple trusted text features, calculate the trusted text feature values for the output results of different models, and perform weighted processing on each trusted text feature value to obtain the overall text score, and finally select the text with the highest score as the final output.

[0144] S24, based on the positioning label and timestamp, align the preliminary optimized text with the original multimedia data through feature alignment technology, and generate a controlled text according to the preset same Prompt template.

[0145] Controlled text generation uses multiple large language models to generate diverse text results for prompts related to the same field, and the highest-quality text is selected and fused from them. Among them, the Prompt design needs to design a field-specific Prompt template according to the application scenario, clarify the format constraints of the generated text, and guide the model to output standardized content.

[0146] In the process of step S2 above, considering the structural differences of different models themselves, different training saturations, and the degree of model optimization effects, resulting in the problem that the text quality generated by different models is uneven. In order to better select and fuse texts of different qualities, learn from each other's strengths and weaknesses, and produce high-quality texts, as follows Figure 3As shown, the present invention mainly solves the problems of sentence smoothness, the fusion of optimal subsequences, and the dependency relationship between the content itself and restrictive keywords through three technologies.

[0147] 1) Fusion strategy, which can integrate texts generated by multiple models to improve the smoothness of sentences. It involves the weighting of credible text features and the construction and training of decision trees to ensure that the fused text is coherent both grammatically and logically.

[0148] 2) Optimization based on model diversity. Considering the diversity of different models, the present invention comprehensively evaluates through text smoothness, keyword coverage evaluation, key information fitting degree, and text information mixing degree to ensure that the generated text is richer and more accurate in content.

[0149] 3) Fusion based on context information. To further improve the relevance and accuracy of the text, the present invention proposes a fusion method based on context information. It includes Prompt design and feature alignment to ensure that the generated text is closely related to the given context information.

[0150] In addition, the LoRa fine-tuning technology is adopted. Using the low-rank adaptation method, the model is adjusted to better adapt to specific context information without significantly increasing the parameters.

[0151] The core idea of LoRA is that after freezing the weights of the pre-trained model, trainable low-rank decomposition matrices are injected into each layer of the Transformer architecture, thus greatly reducing the number of trainable parameters in downstream tasks. During inference, for models using LoRA, the original pre-trained model weights can be directly merged with the trained LoRA weights, so there is no additional overhead during inference.

[0152] The fine-tuning effect is proved through the Loss in the training process. The curve graph is as Figure 7 shown. After multiple rounds of iteration, the final loss converges within a certain range, proving that this method can obtain effective results.

[0153] After experimental verification, the average values of five metrics including Rouge-1, Rouge-2, Rouge-l, Bleu-4, and Coverage for complex unstructured text generation under high constraint conditions of this method reach 88.3%. The voice anti-spoofing metrics are continuously improved, with the voice audio authenticity recognition rate reaching 77.41% and the forgery means traceability F1-score reaching 83.12%.

[0154] Step S3, manage multimedia data according to data tags.

[0155] In the specific implementation process, refer to Figure 4, the management of multimedia data can include the following aspects:

[0156] 1) Resource management: The intelligent media asset management subsystem supports the management of imported materials and finished products, including production materials or finished products such as videos, audios, pictures, documents, etc., and other files for unified content management to meet the diverse content management needs and business scenarios of the system. The content management system also supports the management of videos, audios, pictures, and documents.

[0157] 2) Resource cataloging: Cataloging settings support cataloging different types of files on the platform and setting different cataloging templates. At the same time, content cataloging supports custom cataloging, where cataloging data and templates can be customized, and cataloging information can be configured according to different file types and columns.

[0158] 3) Cross-modal retrieval: Cross-modal retrieval is supported. Through face recognition, image recognition, speech recognition, text recognition, etc., audio-visual and text content is formed into structured fragment index data for intelligent retrieval, automatically identified and cataloged, and a complete fragmented label system is constructed.

[0159] 4) Intelligent tagging: Support for structuring tag annotation of video content. Support using an image recognition model to identify common objects in the video, and intercept frame images to annotate tags and timestamps, facilitating the positioning of the video location where the tag is located. Support using a face recognition model to annotate people in pictures and videos; extract relevant people and the frame timestamps from picture and video information through face features, facilitating the positioning of the video location.

[0160] 5) Intelligent topics: Support for the rapid aggregation of calendars, people, and events. It can customize person, event topics, and sub-topics. The background automatically realizes data aggregation according to the topic name, assisting in quickly viewing resource topics, and providing a specialized window for content production.

[0161] 6) Audio and video transcoding: Provide conversions between professional video formats, common video formats, common audio formats, and common picture formats.

[0162] 7) User management: Administrators uniformly manage user permissions, support adding, deleting, querying, and modifying user information and permissions, and allocate permissions to each user; support the administrator center, support setting user management permissions, and support setting platform users, roles, and organizational structures.

[0163] Based on the above technical solutions, the present application forms structured fragment index data for intelligent retrieval from audio-visual and text content, which is used to construct a complete fragmented tag system; adopts a multi-model text generation technology route of controlled text generation technology to effectively organize information such as objects, scenes, texts, and characters included in the structured fragment index data, and generate data tags. Compared with traditional natural language processing methods, the present application overcomes problems such as missing key information, unsmooth sentence expressions, and incorrect inference conclusions; based on automatically analyzing multimedia data through artificial intelligence algorithms, the present application introduces controlled text generation under hard constraints to process the information extracted by artificial intelligence from multimedia data into a unified standardized format, improving the accuracy and credibility of search results during multimedia data management, making the search results more concentrated and accurate, and improving the efficiency of multimedia data management.

[0164] See Figure 5 , corresponding to the above embodiment of the intelligent multimedia data management method based on fusion-controlled text generation, the embodiment of the present application provides an intelligent multimedia data management device based on fusion-controlled text generation, including:

[0165] A structured fragment index data generation module 1001, configured to receive multimedia data, extract the audio-visual and text content of the multimedia data and store it according to a preset template to form structured fragment index data;

[0166] A controlled text and data tag generation module 1002, configured to input the structured fragment index data into multiple language models, generate a controlled text by fusing texts generated by multiple models, and generate data tags for the multimedia data according to the controlled text;

[0167] A data management module 1003, configured to manage multimedia data according to the data tags.

[0168] Further, the structured fragment index data generation module 1001 is specifically configured to:

[0169] For audio data: extract text information;

[0170] For video data: separate the audio therein and extract text information; identify the information in the frame image and convert it into text information; for the identified character and object information, intercept the frame picture and mark the positioning tag and time stamp;

[0171] For picture data, identify the information in the image and convert it into text information.

[0172] Further, the controlled text and data tag generation module 1002 is specifically configured to:

[0173] Based on the historical performance of the model, weight the trustworthy text features for the candidate texts generated for each model using the decision tree algorithm;

[0174] Based on the decision tree algorithm, score the trustworthy text features based on indicators such as text smoothness, keyword coverage, key information fitting degree, and text information mixing degree;

[0175] Based on the results of the trustworthy text feature weighting and decision tree scoring, weighted-fuse the high-scoring segments to form a preliminary optimized text;

[0176] Based on the positioning label and timestamp, align the preliminary optimized text with the original multimedia data through feature alignment technology, and generate the controlled text according to a preset same Prompt template.

[0177] Further, the multiple language models include models with different architectures, models in different training stages, and / or models trained based on different training sets.

[0178] The intelligent multimedia data management device based on the fusion-controlled text implementation steps and each process of the intelligent multimedia data management method embodiment based on the fusion-controlled text as described above, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0179] See Figure 6 , corresponding to the intelligent multimedia data management method embodiment based on the fusion-controlled text as described above, an embodiment of the present application provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps and each process of the intelligent multimedia data management method embodiment based on the fusion-controlled text as described above, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0180] The memory 1009 can be used to store software programs and various data. The memory 1009 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1009 may include a volatile memory or a non-volatile memory, or the memory 1009 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchlink dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1009 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.

[0181] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1010 either.

[0182] Corresponding to the above embodiments of the intelligent multimedia data management method based on fusion-controlled text generation, the embodiments of the present application also provide a readable storage medium. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, the steps and various processes of the above embodiments of the intelligent multimedia data management method based on fusion-controlled text generation are implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0183] Among them, the processor is the processor in the electronic device described in the embodiments of the present application above. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.

[0184] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0185] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0186] It can be understood that the embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those skilled in the art know that without departing from the spirit and scope of the present invention, these features and embodiments can be variously changed or equivalently replaced. Additionally, those of ordinary skill in the art, under the inspiration or teaching of the present application, can modify these features and embodiments to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application belong to the scope protected by the present invention.

Claims

1. An intelligent multimedia data management method based on fusion controlled text generation, characterized in that: include: Receiving multimedia data, extracting audio, video, and text content of the multimedia data and storing them according to a preset template to form structured segment index data; Inputting the structured segment index data into multiple language models, fusing the text generated by the multiple models to form a controlled text, and generating data tags for the multimedia data based on the controlled text; Multimedia data management is performed according to the data tags.

2. The intelligent multimedia data management method based on fusion controlled text generation according to claim 1 is characterized in that: The receiving of multimedia data, extracting audio, video and text content of the multimedia data and storing them according to a preset template to form structured segment index data specifically includes: For audio data: extract text information; For video data: separate the audio and extract the text information; identify the information in the image based on the frame image and convert it into text information; for the identified person and object information, capture the frame image and annotate the location label and timestamp; For image data, the information in the recognized image is converted into text information.

3. The intelligent multimedia data management method based on fusion controlled text generation according to claim 1 is characterized in that: The step of inputting the structured segment index data into multiple language models, fusing the text generated by the multiple models to form a controlled text, and generating data tags for the multimedia data according to the controlled text specifically includes: According to the historical performance of the model, the candidate texts generated by each model are weighted based on the decision tree algorithm to determine the credible text features; Based on a decision tree algorithm, the credible text features are scored based on text smoothness, keyword coverage, key information fit, and text information confusion indicators; Based on the results of the credible text feature weighting and decision tree scoring, weighted fusion of high-scoring segments to form a preliminary optimized text; Based on the positioning tag and the timestamp, the preliminary optimized text is aligned with the original multimedia data through the feature alignment technology, and the controlled text is generated according to the same preset prompt template.

4. The intelligent multimedia data management method based on fusion controlled text generation according to claim 1 is characterized in that: The multiple language models include models of different architectures, models at different training stages, and / or models trained based on different training sets.

5. An intelligent multimedia data management device based on fusion controlled text generation, characterized in that: include: A structured segment index data generation module, used to receive multimedia data, extract audio, video, and text content of the multimedia data, and store them according to a preset template to form structured segment index data; A controlled text and data label generation module, used for inputting the structured segment index data into multiple language models, generating controlled text by fusing the texts generated by the multiple models, and generating data labels for the multimedia data according to the controlled text; The data management module is used to manage multimedia data according to the data tags.

6. The intelligent multimedia data management device based on fusion controlled text generation according to claim 5 is characterized in that: The structured segment index data generation module is specifically used to: For audio data: extract text information; For video data: separate the audio and extract the text information; identify the information in the image based on the frame image and convert it into text information; for the identified person and object information, capture the frame image and annotate the location label and timestamp; For image data, the information in the recognized image is converted into text information.

7. The intelligent multimedia data management device based on fusion controlled text generation according to claim 5 is characterized in that: The controlled text and data label generation module is specifically used for: According to the historical performance of the model, the candidate texts generated by each model are weighted based on the decision tree algorithm to determine the credible text features; Based on a decision tree algorithm, the credible text features are scored based on text smoothness, keyword coverage, key information fit, and text information confusion indicators; Based on the results of the credible text feature weighting and decision tree scoring, weighted fusion of high-scoring segments to form a preliminary optimized text; Based on the positioning tag and the timestamp, the preliminary optimized text is aligned with the original multimedia data through the feature alignment technology, and the controlled text is generated according to the same preset prompt template.

8. The intelligent multimedia data management device based on fusion controlled text generation according to claim 5 is characterized in that: The multiple language models include models of different architectures, models at different training stages, and / or models trained based on different training sets.

9. An electronic device, characterized in that: The electronic device comprises: a memory, a processor and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the intelligent multimedia data management method based on fused controlled text generation as described in any one of claims 1 to 4 are implemented.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the intelligent multimedia data management method based on fused controlled text generation as described in any one of claims 1 to 4 are implemented.