Prompt-based Data Content Extraction Method, System, Device and Medium

By designing and optimizing prompt words, multimodal large model is used to extract and generate video, image, audio and text data, which solves the complexity of multimodal data processing and lack of flexibility in labeling, and realizes efficient and flexible multimodal data analysis and labeling processing.

CN119691233BActive Publication Date: 2025-07-08FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510204090.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-08
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The prior art has problems such as high complexity, lack of flexibility in data labeling, and difficulty in efficient processing of large-scale data when processing multimodal data. Especially in the unified processing of video, image, audio and text data, there are high technical barriers and low efficiency.

Method used

By configuring the operating environment, designing and optimizing prompt words, using multimodal large models to extract and label video, image, audio and text data, and using large language models to calculate and store them to realize unified processing and labeling of multimodal data.

Benefits of technology

It lowers the technical threshold for multimodal data analysis, supports analysis in any dimension, improves label extraction efficiency and accuracy, is suitable for large-scale data processing, and builds a unified multimodal tag library to facilitate data storage and retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119691233B_ABST
    Figure CN119691233B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of multimodal data processing, and particularly relates to a method, system, device and medium for extracting data content based on prompt words. The present invention proposes a method, system, device and medium for extracting multimodal data content based on prompt words, and the method includes: configuring the operating environment, designing and optimizing prompt words, extracting multimodal data and generating labels, storing and managing labels, using a large language model to calculate the multimodal data to obtain original content description data, parsing the original data and storing it in a relational database, and indexing it by classification; batch data processing. It aims to solve the problem that the content of multimodal data cannot be extracted through prompt words in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal data processing, and particularly relates to a method, system, device and medium for extracting data content based on prompts. Background Art

[0002] With the rapid development of digital media platforms, user-generated content presents multimodal characteristics such as videos, images, audio, and text. Studying the behaviors and dissemination rules of Internet users has become an important topic in the fields of social science and journalism and communication. However, the existing technologies have the following problems when analyzing these data:

[0003] 1. High complexity in multimodal data processing: The characteristics of video, image, audio, and text data are significantly different. Traditional methods usually require separately designing extraction and analysis algorithms for each modality, resulting in high technical thresholds and difficulty in achieving unified processing of multimodal data.

[0004] 2. Lack of flexibility in data tagging: The tagging of data by existing systems is usually limited to predefined dimensions, lacking a flexible tag generation mechanism for different analysis objectives, which restricts the in-depth mining and application of data.

[0005] 3. Challenges brought by data scale: Facing the massive user content on digital media platforms, existing technologies are difficult to efficiently process and analyze multimodal data, especially inefficient in constructing large-scale and multi-dimensional tag libraries.

[0006] The Chinese patent application with the publication number CN106021234A was published on October 12, 2016, and discloses a method and system for extracting tags, belonging to the technical field of language recognition, and capable of achieving relatively accurate tag extraction. The tag extraction method includes: obtaining comments from a database; annotating the part-of-speech of words in the comments; extracting keywords in each comment according to the part-of-speech annotation results; and generating phrases containing the keywords based on the extracted keywords. The embodiments of this patent application can be applied to tag extraction of things with higher complexity such as music and commodities, but it mainly extracts for comment texts, with great limitations and fails to achieve unified processing and analysis of multimodal data such as videos, audios, and images.

[0007] Therefore, how to design an efficient and intelligent method for multimodal data extraction and tagging analysis to meet the structured requirements for multimodal data in the fields of social science and journalism and communication, while reducing the technical threshold of data analysis, has become an urgent problem to be solved. Summary of the Invention

[0008] The object of the present invention is to overcome the deficiencies of the above-mentioned prior art and provide a method, system, device and medium for extracting multimodal data content based on prompts, aiming to solve the problem in the prior art that the content of multimodal data cannot be extracted through prompts.

[0009] To achieve the above object, the present application proposes the following solutions:

[0010] In the first aspect, the present invention proposes a method for extracting multimodal data content based on prompts, and the method includes:

[0011] S1, configure the operating environment to provide a basis for extracting prompt content;

[0012] S2, design and optimize prompts, including designing a general prompt template, theme customization of prompts, and cyclic debugging and optimization of prompts;

[0013] S3, multimodal data extraction and label generation, including extracting content labels for videos, images, audio, and text;

[0014] S4, label storage and management, use a large language model to calculate multimodal data to obtain original content description data, parse the original data and store it in a relational database, and index it by classification;

[0015] S5, batch data processing, after completing the calculation of a single prompt and a single video, image, audio, and text, based on the optimized prompt in step S2, loop through the data to be processed according to the above steps S3 and S4, and store multiple label data in the relational database in sequence according to the results produced by each group of data, and index it by classification.

[0016] Preferably, the steps of configuring the operating environment in S1 include:

[0017] 1), Prepare a high-performance computer host for deploying a large language model environment;

[0018] 2), Build a visual recognition environment (including videos and images) based on an open-source multimodal large model (such as: MiniCPM-2.6-V), and provide an API interface through encoding;

[0019] 3), Build a speech recognition environment based on an open-source speech model (such as: Whisper-large), and provide an API interface through encoding;

[0020] 4), Build a large model text inference environment based on an open-source text large language model (such as: DeepSeek), and provide an API interface through encoding;

[0021] 5), Develop a program to integrate the API interfaces of the above three steps and provide an interface for users to operate.

[0022] Preferably, the steps of designing and optimizing the prompt words in S2 include:

[0023] 1), A general prompt word template, design general prompt words, suitable for basic label extraction of multimodal data;

[0024] 2), Theme-customized prompt words, according to the analysis dimensions of cognition, emotion, narrative, atmosphere, and value judgment, optimize the prompt words to analyze these dimensions and support the data labeling of the above dimensions;

[0025] 3), Loop debugging and optimization, conduct multiple rounds of optimization on the prompt words through the prompt word debugging interface until the analysis requirements are met.

[0026] Preferably, the steps of multimodal data extraction and label generation in S3 include:

[0027] 1), Video processing, pass the optimized prompt words in the above step S2 into the multimodal large model through the interface, and upload a video to be analyzed. The program will parse the key frames of the video and submit the prompt words and key frame image data to the large model. Extract the video content according to the output dimensions specified in the prompt words and generate structured labels in JSON format; the specified output dimensions are the analysis dimensions of cognition, emotion, narrative, atmosphere, and value judgment, and optimize the prompt words to analyze these dimensions;

[0028] 2), Image processing, pass the optimized prompt words in the above step S2 into the multimodal large model through the interface, and upload an image file to be analyzed. The program will submit the prompt words and image data to the large model, extract the image content according to the output dimensions specified in the prompt words, and generate structured labels in JSON format;

[0029] 3), Audio processing, use a speech recognition model to transcribe the audio content, extract the speech information in the audio. For music audio, we use a music analysis model to extract information such as pitch, intensity, rhythm, and spectrum of the audio, so as to form structured text content, and store the structured text content;

[0030] 4), Text analysis, calculate the text content based on the large language model, provide the structured text content output in the above step 3) to the text large language model, and extract content labels according to the given prompt words in different dimensions and generate structured label data. The dimensions are the analysis dimensions of cognition, emotion, narrative, atmosphere, and value judgment.

[0031] Preferably, the specific method for storing and managing the middle tags in step S4 is as follows: through the model calculation in step S3, the original tag data is obtained. We parse this original data to obtain the key-value structure in the analysis dimensions of cognition, emotion, narrative, atmosphere, and value judgment, forming structured data. This structured data is the required tag data, which is then stored in a relational database and indexed by category.

[0032] In a second aspect, the present invention proposes a multi-modal data content extraction system based on prompt words, and the system includes:

[0033] A prompt word design and optimization module, which can continuously clarify the key information to be extracted and improve the accuracy of large model recognition;

[0034] A multi-modal data extraction and tag generation module, which can analyze, extract and generate tags for video, image, audio and text data;

[0035] A tag storage and management module, which uses a large language model to calculate multi-modal data to obtain the original content description data and original tag data, parses the original content description data to obtain the key-value structure in the analysis dimensions of cognition, emotion, narrative, atmosphere, and value judgment, thereby forming structured data. This structured data is the required tag data, which is then stored in a relational database and indexed by category;

[0036] A batch data processing module, based on the optimized prompt words in the prompt word design and optimization module, runs the data to be processed in a loop according to the above multi-modal data extraction and tag generation module and tag storage and management module, and sequentially stores multiple tag data in the relational database and indexes them by category according to the results produced by each group of data.

[0037] Furthermore, the prompt word design and optimization module includes a general prompt word template module, a theme-customized prompt word module, and a loop debugging and optimization module;

[0038] The general prompt word template module is used to design general prompt words, which are suitable for extracting basic tags of multi-modal data;

[0039] The theme-customized prompt word module optimizes the prompt words according to the analysis dimensions of cognition, emotion, narrative, atmosphere, and value judgment to parse these dimensions and supports the data tagging of the above dimensions;

[0040] The loop debugging and optimization module performs multiple rounds of optimization on the prompt words through the prompt word debugging interface until the analysis requirements are met.

[0041] Further, the multi-modal data extraction and label generation module includes a video processing module, an image processing module, an audio processing module, and a text analysis module;

[0042] The video processing module passes the optimized prompt words in the prompt word design and optimization module into the multi-modal large model through an interface, and uploads a video to be analyzed. The program will parse the key frames of the video, submit the prompt words and key frame image data to the large model, extract the video content according to the output dimensions specified in the prompt words, and generate structured labels in JSON format;

[0043] The image processing module passes the optimized prompt words in the prompt word design and optimization module into the multi-modal large model through an interface, and uploads an image file to be analyzed. The program will submit the prompt words and image data to the large model, extract the image content according to the output dimensions specified in the prompt words, and generate structured labels in JSON format;

[0044] The audio processing module uses a speech recognition model to transcribe the audio content and extract the speech information in the audio. For music-type audio, we use a music analysis model to extract information such as the pitch, intensity, rhythm, and spectrum of the audio, so as to form structured text content, and store the structured text content;

[0045] The text analysis module calculates the text content based on a large language model, provides the structured text content output in the above audio processing module to the text large language model, and extracts content labels according to different dimensions based on the given prompt words, and generates structured label data.

[0046] In a third aspect, the present invention proposes a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect is implemented.

[0047] In a fourth aspect, the present invention proposes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0048] Compared with the prior art, the beneficial effects of a method, system, device, and medium for extracting data content based on prompt words provided by the present invention are as follows:

[0049] 1. Lower technical threshold: By adjusting and optimizing the prompt words, accurate extraction and multi-dimensional labeling analysis of video, image, audio, and text data can be performed without traditional coding, greatly reducing the technical difficulty of multi-modal analysis, making it easier for researchers and enterprise users in the fields of social science and journalism and communication to get started, and being suitable for wide applications in the fields of social science and journalism and communication.

[0050] 2. Support arbitrary - dimensional analysis: The present invention can customize the data - analysis dimension through prompt words, meeting the cognitive, emotional, narrative, atmosphere, and value - judgment analysis dimensions of users in various scenarios such as the study of dissemination laws, user behavior, screen style, related themes, and description extraction. Optimize the prompt words to analyze these dimensions, support the data tagging of the above - mentioned dimensions, greatly enhancing the flexibility and practicality of the analysis.

[0051] 3. Improve the efficiency and accuracy of tag extraction: Based on the unified processing of data by a multi - modal large - model, it can efficiently generate a structured tag library without complex feature engineering, achieving seamless integration between different - modal data. Through the prompt - word optimization mechanism, ensure that the tag - extraction results are highly relevant to the analysis theme.

[0052] 4. Optimize large - scale data processing: Support batch processing of large - scale data on social - media platforms, significantly improving the processing efficiency through multi - threading and automation technologies, and being applicable to high - frequency and real - time data - analysis scenarios.

[0053] 5. Enhance data availability: Build a unified multi - modal tag library, facilitating data storage, retrieval, and reuse, providing high - quality basic data support for subsequent social - science and journalism - communication research.

[0054] 6. Wide range of application scenarios: This technology is applicable to multiple directions such as social science, journalism communication, government publicity, medical industry, cultural tourism, advertising and public relations, and marketing. It is used for multi - modal content extraction and tagging analysis of user - behavior data on social - media platforms, and can also be extended to practical application scenarios such as network communication and correlation analysis.

[0055] 7. Empower scientific research and industry: Through intelligent and automated analysis methods, promote technological innovation in the fields of social science and journalism communication, and at the same time provide data support for media platforms, government agencies, and enterprise users, having significant academic value and industrialization potential.

[0056] The features and advantages of the present invention will be described in detail through embodiments in combination with the accompanying drawings. Description of the Drawings

[0057] Figure 1 is a schematic flow chart of the multi - modal data content extraction method based on prompt words in an embodiment of the present invention.

[0058] Figure 2 is a flow block diagram of the steps for configuring the operating environment in S1 of the present invention.

[0059] Figure 3 is a flow block diagram of the steps for designing and optimizing prompt words in S2 of the present invention.

[0060] Figure 4 It is a flow chart of the steps for multi-modal data extraction and label generation in S3 of the present invention.

[0061] Figure 5 It is a schematic diagram of the modules of the system of the present invention.

[0062] Figure 6 General prompt word examples.

[0063] Figure 7 Upload an image / video for analysis and obtain the raw data interface.

[0064] Figure 8 It is the interface of the label table data page formed after the raw data is parsed and labeled.

[0065] Figure 9 It is the interface for speech-to-text conversion.

[0066] Figure 10 It is an example of the speech label extraction result. Detailed implementation manners

[0067] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below through the accompanying drawings and embodiments. However, it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessarily confusing the concepts of the present invention.

[0068] The present invention is mainly used in the fields of social sciences, news communication, government publicity, medical industry, cultural tourism, advertising and public relations, marketing, etc. It is a multi-modal content extraction and labeling analysis technology for user behavior data on social media platforms. By adjusting and optimizing the prompt words, it simplifies the multi-modal element extraction and labeling process of video, image, audio, and text data, and supports flexible analysis of data in any dimension. Different from traditional methods, the present invention does not require complex coding and a separate algorithm for any one of the modalities of video, image, audio, and text. Only by optimizing the prompt words, it can achieve efficient and accurate content extraction and label generation, greatly reducing the technical threshold of multi-modal data analysis. This method provides an intelligent and automated solution for the fields of social sciences, news communication, government publicity, medical industry, cultural tourism, advertising and public relations, marketing, etc., and can quickly build a structured label library, promoting the development of applications such as user behavior analysis, communication mode research, and information monitoring.

[0069] Refer to Figure 1-4 , based on the above situation, an embodiment of the present invention provides a method for extracting multi-modal data content based on prompt words, and the method includes:

[0070] S1. Configure the operating environment to provide the basis for extracting prompt content;

[0071] S2. Design and optimize prompts, including designing a general prompt template, customizing the theme of the prompt, and circularly debugging and optimizing the theme words;

[0072] S3. Extract multi-modal data and generate labels, including extracting content labels from videos, images, audio, and text;

[0073] S4. Store and manage labels, use a large language model to calculate multi-modal data, obtain the original content description data, parse the original data and store it in a relational database, and index it by classification;

[0074] S5. Batch data processing, after completing the calculation of a single prompt and a single video, image, voice, or text, based on the optimized prompt in step S2, run the data to be processed in a loop according to steps S3 and S4 above, and store multiple label data in a relational database (such as MySQL) in sequence according to the results produced by each group of data, and index it by classification.

[0075] Furthermore, the present invention realizes the efficient analysis and labeling of multi-modal data of videos, images, audio, and text by adjusting the prompt. In this design, the following environment needs to be configured in advance (the following steps are for operating environment configuration, which is the basis for subsequent realization of content extraction by the prompt):

[0076] S1 is a prerequisite. The steps for configuring the operating environment in S1 include:

[0077] 1). Prepare a high-performance computer host for deploying the large language model environment;

[0078] 2). Build a visual recognition environment (including videos and images) based on an open-source multi-modal large model (such as: MiniCPM-2.6-V), and provide an API interface through coding;

[0079] 3). Build a speech recognition environment based on an open-source speech model (such as: Whisper-large), and provide an API interface through coding;

[0080] 4). Build a large model text inference environment based on an open-source text large language model (such as: DeepSeek), and provide an API interface through coding;

[0081] 5). Develop a program to integrate the API interfaces of the above three steps and provide a page for users to operate.

[0082] The operating environment can be configured through the above five steps. Only with a good operating environment can the solution of the present invention be better completed.

[0083] Furthermore, the steps of designing and optimizing the prompt words in S2 include:

[0084] 1). General prompt word template, design general prompt words, which are applicable to the extraction of basic labels for multi-modal data. The basic labels include actions, scenes, characters, texts, etc.;

[0085] 2). Theme customization prompt words, optimize the prompt words according to the analysis dimensions of cognition, emotion, narrative, atmosphere, and value judgment to parse these dimensions and support the data labeling of the above dimensions;

[0086] 3). Loop debugging and optimization, perform multiple rounds of optimization on the prompt words through the prompt word debugging interface until the analysis requirements are met.

[0087] The steps of designing and optimizing the prompt words in S2 are elaborated through a specific case. Example: The following are the prompt words optimized in three steps. It is required to extract the cartoon elements appearing in the video. The first step roughly gives the dimensions to be extracted. The second step clarifies the content to be extracted. The third step continues to add the description and definition of the extracted content. After these three steps of optimization of the prompt words, the key information to be extracted becomes increasingly clear, and the accuracy of recognition by the large model is also continuously improving.

[0088] 1). "Please extract the cartoon elements used in the image and output them as standard JSON format data."

[0089] 2). "Please extract the cartoon characters, cartoon scene descriptions, cartoon styles, cartoon background colors, and cartoon accessories used in the image and output them as standard JSON format data. If there are none, output null values. The output example is as follows: {"cartoon characters": [], "cartoon scene descriptions": [], "cartoon styles": [], "cartoon background colors": [], "cartoon accessories": []}.”

[0090] 3). "Please extract the cartoon characters, cartoon scene descriptions, cartoon styles, cartoon background colors, and cartoon accessories used in the image and output them as standard JSON format data. If there are no cartoon characters, output all null values. The output example is as follows: {"cartoon characters": [], "cartoon scene descriptions": [], "cartoon styles": [], "cartoon background colors": [], "cartoon accessories": []}, and set the definition of cartoon: Cartoon is a visual art form characterized by simplification and exaggeration, usually including paintings, illustrations or animations, with distinct colors, simple lines and vivid character designs.”

[0091] Further, the steps of multi-modal data extraction and label generation in S3 include:

[0092] 1) Video processing: Pass the optimized prompt in step S2 above into the multi-modal large model (such as MiniCPM-V2.6) through the interface, and upload a video to be analyzed. The program will parse the key frames of the video, submit the prompt and the key frame image data to the large model, extract the video content according to the output dimension specified in the prompt, and generate cartoon character, cartoon scene description, cartoon style, cartoon background color, and cartoon accessory labels in JSON format;

[0093] 2) Image processing: Pass the optimized prompt in step S2 above into the multi-modal large model (such as MiniCPM-V2.6) through the interface, and upload an image file to be analyzed. The program will submit the prompt and the image data to the large model, extract the image content according to the output dimension specified in the prompt, and generate cartoon character, cartoon scene description, cartoon style, cartoon background color, and cartoon accessory labels in JSON format;

[0094] 3) Audio processing: Use a speech recognition model (such as Whisper-large) to transcribe the audio content and extract the speech information in the audio. For music-type audio, we use a music analysis model to extract information such as pitch, intensity, rhythm, and spectrum of the audio, so as to form structured text content, and store the structured text content;

[0095] 4) Text analysis: Calculate the text content based on the large language model. Provide the structured text content output in step 3) above to the text large language model (such as DeepSeek), and extract content labels according to different dimensions based on the given prompt (another set of prompts for text content extraction), and generate structured label data.

[0096] Further, the specific method of label storage and management in step S4 is as follows: Through the model calculation in step S3, obtain the original label data, parse the original data to obtain the key-value structure in the cognitive, emotional, narrative, atmosphere, and value judgment analysis dimensions (such as the JSON structure data in the long text output by the model), form structured data, and this structured data is the required label data. Then store the label data in a relational database and index it by category.

[0097] Through the above method, the content of data in four modalities, namely video, image, voice, and text, is extracted. This process is completely based on the given prompt words. In the subsequent analysis by users, different analysis dimensions can be flexibly switched without any coding, and different tags can be extracted. This solution is based on a large language model and provides an efficient and intelligent multi-modal data processing method for fields such as social science and journalism and communication by simplifying the operation process, reducing the technical threshold, and enhancing the flexibility of data analysis. The solution of the present invention can also process images. The image processing method is the same as the video processing method. During video processing, frames are extracted to become images, and image processing reduces the frame extraction link.

[0098] Specific usage process: Taking the multi-modal data content extraction method based on prompt words for a single piece of data as an example, the page operations and the generated tag structure are described below. The same applies to batch data analysis, which is a background running program without a page display part.

[0099] 1. Upload an image / video for analysis, and on the result data page extracted according to the prompt words, the prompt words are as Figure 6 , and then perform multi-modal debugging such as Figure 7 .

[0100] 2. The tag table data page formed after label parsing of the original data, with the specific interface diagram Figure 8 .

[0101] 3. Submit the voice-to-text page (with different example data contents), and the display interface is as Figure 9 .

[0102] 4. Extract tags from the above voice-to-text result. The prompt words are as follows: "Please extract the themes, scenes, disease names, medical terms, symptoms, treatment methods, drug names, as well as adjectives and nouns related to a disease described in the content, and output them in JSON format. If there are none, output null values." The output result is as follows. Analyzing this result can obtain structured tag data such as Figure 10 .

[0103] 0 Refer to Figure 5 , in an optional embodiment, the present invention provides a multi-modal data content extraction system based on prompt words, and the system includes:

[0104] A prompt word design and optimization module, which can continuously clarify the key information to be extracted and improve the accuracy of large model recognition;

[0105] A multi-modal data extraction and tag generation module, which can analyze, extract, and generate tags for video, image, audio, and text data;

[0106] The label storage and management module uses a large language model to calculate multimodal data to obtain the original content description data. We parse this original content description data to obtain the key-value structure in the cognitive, emotional, narrative, atmosphere, and value judgment analysis dimensions (such as the JSON structure data in the long text output by the model), thereby forming structured data. This structured data is the label data we need. Then, we store this label data in a relational database (such as MySQL) and index it by category.

[0107] The batch data processing module, based on the optimized prompts in the prompt design and optimization module, runs the data to be processed in a loop according to the above multimodal data extraction and label generation module and label storage and management module. And according to the results produced by each group of data, multiple label data are sequentially stored in a relational database (such as MySQL) and indexed by category.

[0108] Furthermore, the prompt design and optimization module includes a general prompt template module, a theme customization prompt module, and a loop debugging and optimization module.

[0109] The general prompt template module is used to design general prompts, which are applicable to the extraction of basic labels for multimodal data (such as actions, scenes, characters, text, etc.).

[0110] The theme customization prompt module optimizes the prompts according to the cognitive, emotional, narrative, atmosphere, and value judgment analysis dimensions to parse these dimensions and supports the data labeling of these dimensions.

[0111] The loop debugging and optimization module performs multiple rounds of optimization on the prompts through the prompt debugging interface until the analysis requirements are met.

[0112] The working process of the prompt design and optimization module is illustrated through a specific case. For example: The following are the prompts optimized in three steps. It is required to extract the cartoon elements that appear in the video. The first step roughly gives the dimensions to be extracted. The second step clarifies the content to be extracted. The third step further adds the description and definition of the extracted content. After these three steps of optimization of the prompts, the key information to be extracted becomes increasingly clear, and the accuracy of recognition by the large model is also continuously improving.

[0113] 1). "Please extract the cartoon elements used in the image and output them in the standard JSON format data."

[0114] 2) "Please extract the cartoon characters, cartoon scene descriptions, cartoon styles, cartoon background colors, and cartoon accessories used in the image and output them as standard JSON format data. If there are none, output null values. The output example is as follows: {"cartoon characters": [], "cartoon scene descriptions": [], "cartoon styles": [], "cartoon background colors": [], "cartoon accessories": []}."

[0115] 3) "Please extract the cartoon characters, cartoon scene descriptions, cartoon styles, cartoon background colors, and cartoon accessories used in the image and output them as standard JSON format data. If there are no cartoon characters, output all null values. The output example is as follows: {"cartoon characters": [], "cartoon scene descriptions": [], "cartoon styles": [], "cartoon background colors": [], "cartoon accessories": []}. Set the definition of a cartoon: A cartoon is a visual art form characterized by simplification and exaggeration, usually including paintings, illustrations, or animations, with distinct colors, simple lines, and vivid character designs."

[0116] Furthermore, the multi-modal data extraction and label generation module includes a video processing module, an image processing module, an audio processing module, and a text analysis module;

[0117] The video processing module passes the optimized prompt words in the prompt word design and optimization module to the multi-modal large model (such as MiniCPM-V2.6) through an interface and uploads a video to be analyzed. The program will parse the key frames of the video, submit the prompt words and key frame image data to the large model, extract the video content according to the output dimensions specified in the prompt words, and generate labels for cartoon characters, cartoon scene descriptions, cartoon styles, cartoon background colors, and cartoon accessories in JSON format;

[0118] The image processing module passes the optimized prompt words in the prompt word design and optimization module to the multi-modal large model through an interface and uploads an image file to be analyzed. The program will submit the prompt words and image data to the large model, extract the image content according to the output dimensions specified in the prompt words, and generate labels for cartoon characters, cartoon scene descriptions, cartoon styles, cartoon background colors, and cartoon accessories in JSON format;

[0119] The audio processing module uses a speech recognition model to transcribe the audio content, extracts the speech information in the audio. For music audio, we use a music parsing model to extract information such as the pitch, intensity, rhythm, and spectrum of the audio, thereby forming structured text content, and stores the structured text content;

[0120] The text analysis module performs calculations on the text content based on a large language model, provides the structured text content output by the above audio processing module to a text large language model (such as DeepSeek), and extracts content tags according to different dimensions and generates structured tag data according to the given prompt words (another set of prompt words extracted for the text content).

[0121] In an optional embodiment, the present invention further provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it realizes Figure 1-4 the method in

[0122] In an optional embodiment, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it realizes Figure 1-4 the method in

[0123] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0124] It can be understood that the relevant features in the above methods and systems can be referred to each other. In addition, the "first", "second", etc. in the above embodiments are used to distinguish each embodiment, and do not represent the advantages and disadvantages of each embodiment.

[0125] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0126] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The structure required to construct such a system is obvious from the above description. In addition, this application is not directed to any specific programming language. It should be understood that the content of this application described herein can be implemented using various programming languages, and the descriptions made above for specific languages are for disclosing the best mode of this application.

[0127] In addition, the memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0128] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0129] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0130] . These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0131] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0132] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0133] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0134] A computer-readable medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0135] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0136] Those skilled in the art should understand that the computing models provided in the embodiments of the present application are not limited to MiniCPM, Whisper, and DeepSeek. The effects of the present invention can also be achieved by using other computing models in the same field for substitution. The above three models are only a preferred embodiment of the present invention.

[0137] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0138] . The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for extracting multi-modal data content based on prompts, characterized in that: The method includes:[ [ END ] ] S1. Configure the operating environment to provide the basis for extracting prompt content;[ [ END ] ] S2. Design and optimize the prompt, including designing a general prompt template, theme customization of the prompt, and iterative debugging and optimization of the prompt;[ [ END ] ] S3. Multimodal data extraction and label generation, including extracting content labels from videos, images, and texts;[ [ END ] ] S4. Label storage and management, using a large language model to calculate the multimodal data to obtain the original content description data, parsing the original data and storing it in a relational database, and indexing it by category;[ [ END ] ] S5. Batch data processing, after completing the calculation of a single prompt and a single video, image, and text, based on the optimized prompt in step S2, loop through the data to be processed according to steps S3 and S4 above, and store multiple label data in the relational database in sequence according to the results produced by each set of data, and index it by category;[ [ END ] ] The steps of multimodal data extraction and label generation in S3 include:[ [ END ] ] 1). Video processing: Pass the optimized prompt in step S2 above into the multimodal large model through an interface, and upload a video to be analyzed. The program will parse the key frames of the video and submit the prompt and key frame image data to the large model. Extract the video content according to the output dimension specified in the prompt and generate structured labels in JSON format;[ [ END ] ] 2). Image processing: Pass the optimized prompt in step S2 above into the multimodal large model through an interface, and upload an image file to be analyzed. The program will submit the prompt and image data to the large model. Extract the image content according to the output dimension specified in the prompt and generate structured labels in JSON format;[ [ END ] ] 3). Text analysis: Calculate the text content based on the large language model, provide the structured text content output in steps 1) and 2) above to the text large language model, and extract content labels according to different dimensions according to the given prompt, and generate structured label data.[ [ END ] ] 2. The multimodal data content extraction method based on prompts according to claim 1, wherein: The steps of configuring the operating environment in S1 include:[ [ END ] ] 1). Prepare a high-performance computer host for deploying the large language model environment;[ [ END ] ] 2). Build a visual recognition environment based on the open-source multimodal large model and provide an API interface through coding;[ [ END ] ] 3). Build a speech recognition environment based on the open-source speech model and provide an API interface through coding;[ [ END ] ] 4). Build a large model text inference environment based on the open-source text large language model and provide an API interface through coding;[ [ END ] ] 5). Develop a program to integrate the API interfaces of the above three steps and provide an interface for users to operate.[ [ END ] ] 3. The multimodal data content extraction method based on prompt words as described in claim 1, wherein: The steps of designing and optimizing the prompt in S2 include:[ [ END ] ] 1). General prompt template: Design a general prompt applicable to the basic label extraction of multimodal data;[ [ END ] ] 2). Thematic customized prompt: Optimize the prompt according to the cognitive, emotional, narrative, atmosphere, and value judgment analysis dimensions to parse these dimensions and support the data labeling of these dimensions;[ [ END ] ] 3). Iterative debugging and optimization: Perform multiple rounds of optimization on the prompt through the prompt debugging interface until the analysis requirements are met.[ [ END ] ] 4. A method for extracting multimodal data content based on prompt words according to claim 1, characterized in that: The specific method of storing and managing the middle tags in step S4 is as follows: through the model calculation in step S3, the original tag data is obtained. We parse this original data to obtain the key-value structure in the analysis dimensions of cognition, emotion, narrative, atmosphere, and value judgment, forming structured data. This structured data is the required tag data, which is then stored in a relational database and indexed by category.

5. A multi-modal data content extraction system based on prompting words, characterized in that: The system includes: A prompt design and optimization module that can continuously clarify the key information to be extracted and improve the accuracy of large model recognition; A multi-modal data extraction and tag generation module that can analyze, extract, and generate tags for video, image, and text data; A tag storage and management module that uses a large language model to calculate multi-modal data to obtain the original content description data. We parse this original content description data to obtain the key-value structure in the analysis dimensions of cognition, emotion, narrative, atmosphere, and value judgment, thereby forming structured data. This structured data is the required tag data, which is then stored in a relational database and indexed by category; A batch data processing module that, based on the optimized prompts in the prompt design and optimization module, runs the data to be processed in a loop according to the above multi-modal data extraction and tag generation module and tag storage and management module, and sequentially stores multiple tag data in the relational database and indexes them by category according to the results produced by each set of data; The multi-modal data extraction and tag generation module includes a video processing module, an image processing module, and a text analysis module; The video processing module passes the optimized prompts in the prompt design and optimization module into the multi-modal large model through an interface and uploads a video to be analyzed. The program will parse the key frames of the video and submit the prompts and key frame image data to the large model, extract the video content according to the output dimensions specified in the prompts, and generate structured tags in JSON format; The image processing module passes the optimized prompts in the prompt design and optimization module into the multi-modal large model through an interface and uploads an image file to be analyzed. The program will submit the prompts and image data to the large model, extract the image content according to the output dimensions specified in the prompts, and generate structured tags in JSON format; The text analysis module calculates the text content based on the large language model, provides the structured text content output by the above processing module to the text large language model, and extracts content tags according to different dimensions and generates structured tag data according to the given prompts.

6. The multimodal data content extraction system based on prompt words according to claim 5, characterized in that: The prompt design and optimization module includes a general prompt template module, a theme-customized prompt module, and a loop debugging and optimization module; The general prompt template module is used to design general prompts, which are applicable to the basic tag extraction of multi-modal data; The theme-customized prompt module optimizes the prompts according to the analysis dimensions of cognition, emotion, narrative, atmosphere, and value judgment to parse these dimensions and supports the data tagging of the above dimensions; The loop debugging and optimization module fine-tunes and optimizes the prompt multiple times through the prompt debugging interface until the analysis requirements are met.

7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the method described in any one of claims 1 to 4.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the method described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Label extraction method and system

    CN106021234A

  • Method and device for generating multimedia product of music and medium

    CN118427372A