Method, system and equipment for synthesizing and processing four-mode alignment data
By designing fine and simple prompt words, combining video data to extract multiple modal data, four-modal alignment is achieved, and cosine similarity evaluation, the problem of poor alignment effect of multimodal data in the prior art is solved, significantly improving alignment accuracy and adaptability.
Patent Information
- Application Number
- CN202411923946.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-27
AI Technical Summary
When facing complex scenarios and multimodal data alignment, existing multimodal models are difficult to effectively capture the deep relationship between the four modes, resulting in inaccurate and coherent alignment effects, and lack of high-quality four-modal data sets, which limits the performance of the model in the Chinese environment.
By designing fine prompt words and brief prompt words, the alignment of text modal data to images, video and audio data is achieved, text, audio and image data is extracted using video to achieve four-modal alignment with video as the core, and the data alignment effect is evaluated through cosine similarity.
It improves the accuracy and adaptability of alignment between multimodal data, expands the application scope of multimodal alignment, especially in complex video content analysis, significantly improves the performance of the model, and greatly improves data processing efficiency through automated processing.
Smart Images

Figure CN120046091A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and more specifically, relates to a method, system, and device for four-modal alignment data synthesis and processing. Background Art
[0002] With the rapid development of artificial intelligence technology, multi-modal systems have become an important research direction in the field of AI. Multi-modal understanding and generation technologies can process and fuse data from different modalities (such as images, texts, audio, and videos) to improve the machine's ability to understand and generate complex information. This technology has shown great application potential in many fields such as intelligent assistants, autonomous driving, virtual reality, medical diagnosis, and content recommendation. However, existing multi-modal models and technologies still face many challenges and limitations when dealing with complex scenarios and multi-modal data alignment, especially in the full alignment and consistency evaluation of data in four modalities: images, texts, audio, and videos.
[0003] Currently, the application of four-modal data of images, texts, audio, and videos is becoming increasingly widespread. Many practical application scenarios, such as video analysis, cross-modal search, and automatic subtitle generation, require precise alignment and conversion between these four modalities. However, the technology for achieving such multi-modal alignment still faces huge challenges. Especially when dealing with the complex semantic associations between the four modalities, existing technologies often cannot effectively capture the deep relationships between these modalities. This alignment problem is particularly obvious in complex multi-modal scenarios. The model often cannot accurately identify key frames in videos, semantic intentions in texts, emotional information in audio, or detailed content in images, resulting in inaccurate and incoherent final generated results.
[0004] More critically, large-scale four-modal datasets that can be used to train and evaluate these multi-modal alignment technologies are extremely scarce. Currently, most common multi-modal datasets only contain two modalities: images and texts, lacking the support of high-quality audio and video data. For example, datasets such as COCO and Flickr30k are widely used in image-text alignment tasks, but they are mainly oriented towards English scenarios and do not include multi-modal information of audio and video. The lack of large-scale Chinese multi-modal datasets further restricts the performance of existing models in the Chinese environment and cannot fully meet the needs of multi-modal fusion. Such data scarcity directly affects the model's understanding and generation capabilities in Chinese multi-modal tasks, especially in scenarios that require processing complex semantic information, where the model's performance significantly degrades.
[0005] In view of this, overcoming the technical defects existing in the above-mentioned prior art is an urgent problem to be solved in this technical field. Summary of the Invention
[0006] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides a method, system and device for four-modal alignment data synthesis and processing, aiming to solve the technical problems of poor flexibility and low accuracy in cross-modal data alignment and semantic consistency, and the lack of training and evaluation data sets.
[0007] To achieve the above object, according to one aspect of the present invention, there is provided a method for four-modal alignment data synthesis and processing, the method comprising:
[0008] By separately designing fine-grained prompts and brief prompts, realizing the alignment of text-modal data with image, video and audio data;
[0009] Then extracting text, audio and image data from the video, aligning the text with the video, the image with the video, and the audio with the video, so as to achieve four-modal alignment with the video as the core;
[0010] After the four-modal alignment with the video as the core, calculating the cosine similarity between two single modalities, and evaluating the data alignment effect according to the cosine similarity.
[0011] As a further improvement and supplement to the above solution, the present invention further includes the following additional technical features.
[0012] Preferably, the method of realizing the alignment of text-modal data with image, video and audio data by separately designing fine-grained prompts and brief prompts comprises:
[0013] In image-text alignment, the fine-grained prompt describes the image details, and the brief prompt summarizes the main information of the image;
[0014] In video-text alignment, the fine-grained prompt analyzes the video details, and the brief prompt obtains the core information of the video;
[0015] In audio-text alignment, the fine-grained prompt analyzes the emotion and intonation of the audio, and the brief prompt extracts the core content of the audio.
[0016] Preferably, the method of extracting text from the video and aligning the text with the video comprises:
[0017] Performing frame-by-frame processing on the video, extracting the text region therefrom, recognizing the text, aligning the recognized text with the time axis of the video, generating a subtitle file, and realizing the two-modal alignment of the text and the video.
[0018] Preferably, the method of extracting text from the video and aligning the text with the video further comprises:
[0019] Extracting the audio track of the video, converting the speech into text, and realizing the two-modal alignment of the text and the video.
[0020] Preferably, the method for extracting image data from a video and aligning the image with the video includes:
[0021] Analyze high-energy points from the text corresponding to the video, where the high-energy points correspond to key plots or moments in the video, and identify image frames representing the video content through the high-energy points, thereby achieving bimodal alignment of the image and the video.
[0022] Preferably, the method for extracting audio data from a video and aligning the audio with the video includes:
[0023] Extract the original audio in the video, calculate the semantic consistency between the original audio and the video to determine whether the original audio meets the context requirements of the video, and calculate the similarity between the semantics of the original audio and the video;
[0024] If the similarity exceeds a first threshold, the original audio is consistent with the video content; otherwise, an audio matching the video is automatically generated, thereby achieving bimodal alignment of the audio and the video.
[0025] Preferably, there are at least two types of the original audio.
[0026] Preferably, the method for calculating the cosine similarity between two unimodals includes:
[0027] Take the two unimodals as input data to obtain two unimodal vectors respectively, and the cosine similarity is expressed as:
[0028]
[0029] where: v 1 ·v 2 is the dot product of the unimodal vector v 1 and the unimodal vector v 2 ||v 1 || and ||v 2 || are the two-norms of the unimodal vector v 1 and the unimodal vector v 2 respectively.
[0030] According to another aspect of the present invention, there is provided a device, which includes:
[0031] One or more processors;
[0032] A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method for four-modal alignment data synthesis and processing as described in the first aspect.
[0033] Generally speaking, compared with the prior art, the above technical solution conceived by the present invention has the following beneficial effects:
[0034] Through the dual-modal intelligent alignment technology based on the general large model, by using fine-grained and concise prompt words design, it can generate more accurate alignment data for different scenario requirements. It is applicable to the detail extraction and core information capture of image, video, audio and text data, greatly improving the alignment accuracy and adaptability between multi-modal data.
[0035] The present invention not only covers the dual-modal alignment of pictures, texts, videos and audios, but also realizes the four-modal alignment technology centered on video. Through subtitle extraction, deep semantic analysis and multi-modal contrast learning, it ensures the semantic and temporal consistency between text, image, audio and video, thus expanding the application scope of multi-modal alignment, especially with remarkable effects in complex video content analysis.
[0036] Through the consistency evaluation technology based on the four-modal unified alignment model, the present invention can measure the alignment degree between each modality through cosine similarity calculation. This method can comprehensively evaluate the overall consistency of multi-modal generated content at the time, semantic and perceptual levels, ensuring that the finally output multi-modal data is highly coordinated and consistent semantically.
[0037] Through the prompt word automatic generation mechanism, combined with the multi-modal alignment model, the present invention can quickly realize the intelligent alignment of multi-modal data. Compared with the traditional manual adjustment method, it greatly improves the data processing efficiency, reduces the complexity, and has a higher automation level.
[0038] In summary, the present invention has significant improvements in the accuracy, diversity, consistency evaluation and automated processing of multi-modal data alignment, and has broad application prospects and technical advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. Obviously, the following described drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0040] Figure 1 It is a schematic flowchart of a method for synthesizing and processing four-modal alignment data provided in the first embodiment;
[0041] Figure 2 It is a schematic diagram of a method for realizing four-modal alignment centered on video provided in the first embodiment;
[0042] Figure 3FIG. 0 is a schematic process diagram of a method for achieving four-modal alignment with video as the core provided by the first embodiment;
[0043] Figure 4 FIG. 4 is a schematic system diagram of four-modal alignment data synthesis and processing provided by the second embodiment;
[0044] Figure 5 FIG. 8 is a schematic equipment diagram of four-modal alignment data synthesis and processing provided by the third embodiment. Detailed implementation manners
[0045] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0046] First embodiment
[0047] The first embodiment provides a method for four-modal alignment data synthesis and processing. The method includes the steps as Figure 1 shown:
[0048] S101: By separately designing a fine prompt and a concise prompt, the alignment of text-modal data with image, video and audio data is achieved.
[0049] Based on the multi-modal data intelligent alignment technology of a general large model, the first embodiment realizes the efficient alignment between multi-modal data through prompt engineering design. The specific technical solutions include: in image-text alignment, a fine prompt is designed to describe image details in a refined manner, such as shape, color, texture, and a concise prompt is used to summarize the main information of the image. For image-text alignment, the prompt can be refined to visual features, such as color, shape or scene context, such as action, background, so as to describe image details in a refined manner or summarize its main information; in video-text alignment, a fine prompt is used to analyze video details in detail, and a concise prompt obtains the core information of the video. Video-text alignment then guides the model to extract dynamic information in the video, such as action changes, spatio-temporal relationships or static information, such as the content of a specific frame, through the prompt. In audio-text alignment, a fine prompt analyzes the emotion, intonation, etc. of the audio, and a concise prompt quickly extracts the core content of the audio. The prompt can be used to analyze the semantic layer or perceptual layer of the audio; in video-image alignment, representative frames are selected by analyzing video dynamics, or the generation and alignment of images and videos are achieved through prompts. Video-image alignment selects representative frames by analyzing video dynamic changes, or designs precise prompts to achieve the generation and alignment of images and videos.
[0050] Flexibly design prompts through prompt engineering to achieve efficient alignment between image, video, audio, and text data. For image-text alignment, prompts can be refined to visual features such as color, shape, or scene context such as action, background, so as to precisely describe image details or summarize its main information. Video-text alignment extracts dynamic information in the video such as action changes, spatio-temporal relationships, or static information such as specific frame content through prompts. In audio-text alignment, prompts can be used to analyze the semantic layer or perceptual layer of audio. In addition, video-image alignment selects representative frames by analyzing the dynamic changes in the video or designs precise prompts to achieve the generation and alignment of images and videos. The flexibility of prompt design enables the system to automatically generate adapted alignment data, meet the multi-scenario and multi-modal alignment requirements, and improve the practicality and scalability of the system.
[0051] S102: Then extract text, audio, and image data from the video, align the text with the video, the image with the video, and the audio with the video, so as to achieve four-modal alignment with the video as the core.
[0052] As Figure 2 and Figure 3 shown, in order to achieve precise alignment between text and video, the subtitle extraction module provides three main solutions: external subtitle extraction, automatic speech recognition (ASR), and optical character recognition (OCR). These methods are flexibly applied to different video scenarios to ensure that complete text information can be extracted from the video and precise alignment operations can be performed.
[0053] For videos with external subtitles, the text in the subtitle file can be directly used for alignment analysis. Common subtitle formats such as SRT, ASS, VTT, etc. all contain timestamp information for synchronizing subtitles with video content. The system first identifies and parses the format of the subtitle file, automatically detects the character encoding of the subtitle file, and precisely aligns the subtitle text with the video frames according to the timestamp information. Through the external subtitle extraction solution, it can effectively ensure the perfect synchronization of subtitles with video content without additional processing of video audio or images.
[0054] For videos without subtitles, OCR or ASR technology can be used to extract subtitles. The OCR method realizes subtitle extraction through the following steps: first, process each frame of the video, extract the text area from it, then perform text detection and recognition, then perform subsequent processing, and finally align the recognized text with the time axis of the video to generate the final subtitle file.
[0055] The ASR method first extracts the audio track of the video for the next step of speech processing. Before audio analysis, noise reduction, enhancement and other processing are performed on the audio to improve the accuracy of speech recognition. Then, the speech is converted into text through a pre-trained ASR model, which has been trained on different languages and accents to handle various audio conditions. Each of these two technologies has its own advantages, and which method to choose depends on specific requirements and the language type of the video content. For example, if the video is in English, the Chinese subtitles can be obtained through the OCR solution to achieve better results.
[0056] In the process of aligning text with video, the subtitles contain rich semantic information and are a direct description of the video content. Based on this, by analyzing the high-energy points in the subtitles, we can capture the key information of the video. Here, the "high-energy points" refer to the text parts in the subtitles that best reflect the core content of the video. They are usually keywords or phrases rich in semantic information that can reflect the peak moments of the video plot, emotion or important information. These high-energy points usually correspond to the key plots or moments in the video, helping us identify the image frames that best represent the video content. To achieve this analysis, we fine-tuned a large language model to enable it to identify and extract high-energy points in several subtitles. Compared with the traditional frame selection based on visual features, the subtitle-based analysis method can better retain the context information in the video and ensure that the selected image frames are consistent with the core content of the video.
[0057] In the first step of aligning audio with video, the system first needs to extract the audio track corresponding to the video. The system extracts the original audio in the video, usually the accompanying audio track of the video, and directly uses this audio track for analysis and processing. This audio data can be human dialogue, background music or environmental sound effects. By extracting different types of audio, the system can ensure that the complete speech context in the video is covered.
[0058] After the audio track extraction is completed, the system needs to calculate the semantic consistency between the audio and the video to determine whether the audio meets the context requirements of the video. In this process, the LanguageBind model is used. The audio generates feature vectors through spectrogram representation, and the subtitles of the video generate corresponding text vectors by a pre-trained large language model. These vectors are input into the contrastive learning model to calculate the similarity between the two. To ensure the alignment effect between the audio and the video, the system also sets a similarity threshold to decide whether further processing is needed. The similarity threshold is set through a large number of experiments, considering the changes in different video scenarios and audio features. When the similarity between the audio and the video subtitles is lower than this threshold, the system will judge that the audio does not match the video content, and at this time, the EgoSonics model will be activated to generate an audio track that highly matches the video.
[0059] S103: After the four-modal alignment with the video as the core, calculate the cosine similarity between two single modalities, and evaluate the data alignment effect according to the cosine similarity.
[0060] Taking the two single modalities as input data, two single-modal vectors are obtained respectively, and the cosine similarity is expressed as:
[0061]
[0062] where: v 1 ·v 2 is the dot product of the single-modal vector v 1 and the single-modal vector v 2 ||v 1 || and ||v 2 || are the two-norms of the single-modal vector v 1 and the single-modal vector v 2 respectively.
[0063] Combined with the embodiments of the present invention, there is also a preferred implementation scheme. Specifically, the alignment of text-modal data with image, video, and audio data is achieved by separately designing fine-grained prompts and concise prompts. The method includes:
[0064] In image-text alignment, the fine-grained prompt describes the image details, and the concise prompt summarizes the main information of the image;
[0065] In video-text alignment, the fine-grained prompt analyzes the video details, and the concise prompt obtains the core information of the video;
[0066] In audio-text alignment, the fine-grained prompt analyzes the emotion and intonation of the audio, and the concise prompt extracts the core content of the audio.
[0067] In this first embodiment, efficient alignment between image, video, audio, and text data is achieved by flexibly designing prompts. For example, in image-text alignment, the prompt can be refined to visual features (such as color, shape) or scene context (such as action, background), so as to finely describe the image details or summarize its main information. The flexibility of prompt design enables the system to automatically generate alignment data adapted to different scenarios, meet the multi-modal alignment requirements, and thus improve the practicality and scalability of the system. The present invention solves the limitations of existing multi-modal large models in fine-grained and coarse-grained alignment requirements, and can effectively process the fine description of image details and the extraction of core information.
[0068] By adopting a text - video alignment scheme based on subtitle extraction, automatic speech recognition (ASR), and optical character recognition (OCR), combined with the in - depth semantic analysis of large language models, the context understanding ability of images and videos is enhanced. Through multi - modal contrast learning and the EgoSonics model, efficient alignment of audio and video in the semantic sharing space is achieved, and the alignment problem when audio and video are semantically inconsistent is solved.
[0069] Combined with the embodiments of the present invention, there is also a preferred implementation scheme. Specifically, the method of extracting text from a video and aligning the text with the video includes:
[0070] Perform frame - by - frame processing on the video, extract the text area therefrom, recognize the text, align the recognized text with the time axis of the video, generate a subtitle file, and achieve the bimodal alignment of text and video.
[0071] Combined with the embodiments of the present invention, there is also a preferred implementation scheme. Specifically, the method of extracting text from a video and aligning the text with the video further includes:
[0072] Extract the audio track of the video, convert the speech into text, and achieve the bimodal alignment of text and video.
[0073] Combined with the embodiments of the present invention, there is also a preferred implementation scheme. Specifically, the method of extracting image data from a video and aligning the image with the video includes:
[0074] Analyze high - energy points from the text corresponding to the video. The high - energy points correspond to key plots or moments in the video. Identify the image frames representing the video content through the high - energy points, thereby achieving the bimodal alignment of image and video.
[0075] Combined with the embodiments of the present invention, there is also a preferred implementation scheme. Specifically, the method of extracting audio data from a video and aligning the audio with the video includes:
[0076] Extract the original audio in the video, calculate the semantic consistency between the original audio and the video to determine whether the original audio meets the context requirements of the video, and calculate the similarity of the semantics between the original audio and the video;
[0077] If the similarity exceeds the first threshold, the original audio is consistent with the video content; otherwise, automatically generate an audio that matches the video, thereby achieving the bimodal alignment of audio and video.
[0078] Combined with the embodiments of the present invention, there is also a preferred implementation scheme. Specifically, the types of the original audio include at least two.
[0079] Combined with the embodiments of the present invention, there is also a preferred implementation solution. Specifically, the method for calculating the cosine similarity between two unimodals includes:
[0080] Taking the two unimodals as input data, two unimodal vectors are obtained respectively, and the cosine similarity is expressed as:
[0081]
[0082] where: v 1 ·v 2 is the dot product of the unimodal vector v 1 and the unimodal vector v 2 ||v 1 || and ||v 2 || are the two-norms of the unimodal vectors v 1 and v 2 respectively.
[0083] First, the data of each modality is preprocessed to ensure that it can be input into the CoDi model for encoding. After being processed by the CoDi model, the latent vector representations of each modality are obtained. Then, we measure their alignment degree by calculating the cosine similarity between different modalities. Specifically, the cosine similarity is used to evaluate the similarity of each pair of modalities, such as the semantic consistency between images and texts, or the temporal synchronization between videos and audios. By statistically analyzing the similarities of different modality combinations, we can obtain an overall consistency evaluation metric, which can comprehensively reflect the consistency of multi-modal generated content at multiple levels such as time, semantics, and perception.
[0084] In the first embodiment, by preprocessing the modality data such as images, texts, videos, and audios and inputting them into the CoDi model, the cosine similarity is used to calculate the alignment degree of different modalities, and a consistency evaluation metric is generated. For example, for a dual-modal alignment data of image-text, taking the image I and text T as the inputs of the CoDi model, two vectors v i and v t are obtained respectively, and their cosine similarity can be expressed as:
[0085]
[0086] where: v i ·v t is the dot product of the vectors, ||v i || and ||v t || are the two-norms of the vectors v i and v t respectively.
[0087] By statistically analyzing the similarity of a standard aligned dataset through the combination of recognized different modalities, we can obtain an overall consistency evaluation metric. If the cosine similarity of the test data pair is lower than this consistency evaluation metric, it indicates that there are significant inconsistencies between the modalities in some aspects. If it is higher than the consistency evaluation metric, it indicates that the modalities are already aligned. The present invention provides a comprehensive multi-modal data consistency evaluation method that can measure the alignment quality of multi-modal data in multiple dimensions such as time, semantics, and perception.
[0088] Embodiment 2:
[0089] This Embodiment 2 provides a system for synthesizing and processing four-modal aligned data, characterized in that, as Figures 2 to 4 shown, the system includes:
[0090] Subtitle extraction module: Extract text in the video;
[0091] Large language model analysis module: Perform semantic analysis to enhance the context understanding ability of images and videos;
[0092] Audio alignment module: Used to align audio with video;
[0093] Key frame selection module: Achieve alignment of images and videos through timestamp indexing;
[0094] Evaluation module: Used to evaluate the alignment effect of four-modal data;
[0095] Dual-modal alignment module: Design fine prompts and brief prompts respectively to achieve the alignment of text-modal data with image, video, and audio data.
[0096] Taking the video-centered four-modal unified alignment as the goal, this second embodiment designs and implements a four-modal alignment system. Through the dual-modal intelligent alignment technology based on a general large model and the design of fine-grained and concise prompt words, it can generate more accurate alignment data according to different scenario requirements. It is applicable to the detail extraction and core information capture of image, video, audio, and text data, greatly improving the alignment accuracy and adaptability between multi-modal data. This embodiment not only covers the dual-modal alignment of text and image, video, audio, etc., but also implements the video-centered four-modal alignment technology. Through subtitle extraction, deep semantic analysis, and multi-modal contrast learning, it ensures the semantic and temporal consistency between text, image, audio, and video, thus expanding the application scope of multi-modal alignment, especially with remarkable effects in complex video content analysis. Through the consistency evaluation technology based on the four-modal unified alignment model, this embodiment can measure the alignment degree between each modality through cosine similarity calculation. This method can comprehensively evaluate the overall consistency of multi-modal generated content at the time, semantic, and perceptual levels, ensuring that the finally output multi-modal data is highly coordinated and consistent semantically. This embodiment can quickly achieve the intelligent alignment of multi-modal data through the prompt word automatic generation mechanism combined with the multi-modal alignment model. Compared with the traditional manual adjustment method, it greatly improves the data processing efficiency, reduces the complexity, and has a higher automation level.
[0097] Embodiment Three:
[0098] This Embodiment Three provides a device for four-modal alignment data synthesis and processing, as Figure 5 shown, the device includes:
[0099] One or more processors;
[0100] A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method for four-modal alignment data synthesis and processing as described in any one of Embodiment One.
[0101] Figure 5 It is a schematic structural diagram of the device for four-modal alignment data synthesis and processing provided by this Embodiment Three. Figure 5 It shows a block diagram of an exemplary device for four-modal alignment data synthesis and processing suitable for implementing the embodiments of the present invention. Figure 5 The shown device for four-modal alignment data synthesis and processing is only an example and should not bring any limitations to the functions and usage scope of the embodiments of the present invention.
[0102] As Figure 5As shown, the device for four-modal alignment data synthesis and processing is presented in the form of a general device. The components of the device for four-modal alignment data synthesis and processing may include, but are not limited to: one or more processors or processing units, a memory, and a bus connecting different system components (including the memory and the processing unit).
[0103] The bus represents one or more of several types of bus architectures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus architectures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0104] The device for four-modal alignment data synthesis and processing typically includes a variety of computer system-readable media. These media can be any available media accessible by the device that can be modified by the intelligent logging interpretation model, including volatile and non-volatile media, removable and non-removable media.
[0105] The memory may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory. The device for four-modal alignment data synthesis and processing may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system can be used to read and write on non-removable, non-volatile magnetic media ( Figure 4 not shown, commonly referred to as a "hard disk drive"). Although Figure 4 not shown in the figure, a disk drive for reading and writing on a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing on a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical media) can be provided. In these cases, each drive can be connected to the bus through one or more data media interfaces. The memory may include at least one program product having a set (such as at least one) of program modules that are configured to perform the functions of the embodiments of the present invention.
[0106] A program / utility with a set (at least one) of program modules can be stored, for example, in the memory. Such program modules include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples. The program modules generally perform the functions and / or methods described in the embodiments of the present invention.
[0107] The device for four-modal alignment data synthesis and processing can also communicate with one or more external devices (such as keyboards, pointing devices, displays, etc.), and can also communicate with one or more devices that enable users to interact with the device for four-modal alignment data synthesis and processing, and / or communicate with any device (such as network cards, modems, etc.) that enables the device for four-modal alignment data synthesis and processing to communicate with one or more other devices. This communication can be carried out through an input / output (I / O) interface. Moreover, the device for intelligent well logging interpretation model correction can also communicate with one or more networks (such as local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) through a network adapter. As Figure 4 shown, the network adapter communicates with other modules of the device for four-modal alignment data synthesis and processing through a bus. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the device for four-modal alignment data synthesis and processing, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0108] The processing unit executes various functional applications and data processing by running programs stored in the memory, such as implementing the method for four-modal alignment data synthesis and processing provided in any embodiment of the present invention. That is: by separately designing fine prompts and brief prompts, the alignment of text modal data with image, video, and audio data is realized; then text, audio, and image data are extracted from the video, and the text is aligned with the video, the image is aligned with the video, and the audio is aligned with the video, so as to realize four-modal alignment with the video as the core.
[0109] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for synthesizing and processing quad-modal alignment data, characterized in that: The method comprises: By designing detailed prompt words and brief prompt words respectively, the alignment of text modality data with image, video and audio data is achieved; Then extract text, audio and image data from the video, align the text with the video, align the image with the video, and align the audio with the video, so as to achieve four-modal alignment with video as the core; After the four-modality alignment with video as the core, the cosine similarity between the two single modalities is calculated, and the data alignment effect is evaluated based on the cosine similarity.
2. The method for synthesizing and processing quad-modal alignment data according to claim 1, characterized in that: The method of realizing alignment of text modality data with image, video and audio data by respectively designing fine prompt words and brief prompt words includes: In image-text alignment, the fine prompt words describe the image details, and the concise prompt words summarize the main information of the image; In video-text alignment, the fine prompt words analyze the video details, and the concise prompt words obtain the core information of the video; In audio-text alignment, the detailed prompt words analyze the emotion and intonation of the audio, and the concise prompt words extract the core content of the audio.
3. The method for synthesizing and processing quad-modal alignment data according to claim 2, characterized in that: The method for extracting text from a video and aligning the text with the video comprises: The video is processed frame by frame, the text area is extracted from it, the text is recognized, the recognized text is aligned with the timeline of the video, and a subtitle file is generated to achieve bimodal alignment of text and video.
4. The method for synthesizing and processing quad-modal alignment data according to claim 3, characterized in that: The method of extracting text from a video and aligning the text with the video also includes: Extract the audio track of the video, convert the speech into text, and achieve bimodal alignment of text and video.
5. The method for synthesizing and processing quad-modal alignment data according to claim 2, characterized in that: The method of extracting image data from a video and aligning the image with the video comprises: High energy points are obtained by analyzing the text corresponding to the video, and the high energy points correspond to key plots or moments in the video. Image frames representing the video content are identified through the high energy points, thereby achieving bimodal alignment of the image and the video.
6. The method for synthesizing and processing quad-modal alignment data according to claim 2, characterized in that: The method for extracting audio data from a video and aligning the audio with the video comprises: Extracting the original audio in the video, calculating the semantic consistency between the original audio and the video to determine whether the original audio meets the contextual requirements of the video, and calculating the semantic similarity between the original audio and the video; If the similarity exceeds a first threshold, the original audio is consistent with the video content; otherwise, an audio matching the video is automatically generated, thereby achieving bimodal alignment of audio and video.
7. The method for synthesizing and processing quad-modal alignment data according to claim 6, characterized in that: The types of the original audio include at least two.
8. The method for synthesizing and processing quad-modal alignment data according to claim 1, characterized in that: The method for calculating the cosine similarity between two single modalities comprises: Taking the two unimodal states as input data, two unimodal vectors are obtained respectively, and the cosine similarity is expressed as: Where: v1·v2 is the dot product of the unimodal vector v1 and the unimodal vector v2, ||v1|| and ||v2|| are the bi-norms of the unimodal vector v1 and the unimodal vector v2 respectively.
9. A system for synthesizing and processing quad-modal alignment data, characterized in that: The system includes: Subtitle extraction module: extract text from the video; Large language model analysis module: performs semantic analysis to enhance the contextual understanding of images and videos; Audio alignment module: used to align audio with video; Key frame selection module: aligning images and videos through timestamp index; Evaluation module: used to evaluate the alignment effect of quad-modal data; Bimodal alignment module: Design detailed prompt words and concise prompt words respectively to realize the alignment of text modality data with image, video and audio data.
10. A device for synthesizing and processing quad-modal alignment data, characterized in that: Equipment includes: one or more processors; A storage device, used for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.
Citation Information
Cited By
Multi-language cross-modal information retrieval method, device and equipment
CN121030052A