Multi-round real-time multi-modal large model interaction method, related device and storage medium

Through audio to text and keyframe image segmentation processing, combined with multi-modal large models, the accuracy and efficiency of multi-round real-time multi-modal large model interaction is improved, the problem of low accuracy in the existing technology is solved, and more natural and efficient interaction is achieved.

CN120407854AActive Publication Date: 2025-08-01BEIJING REALAI TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510897925.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-08-01
Estimated Expiration
2045-07-01

AI Technical Summary

Technical Problem

The accuracy of existing large-modal interactions is low, especially in multiple rounds of real-time multimodal interactions.

Method used

By obtaining the audio and video input by the user, converting it into text and extracting keyframe images for segmentation processing, the target multimodal large model is used to fuse the audio and video information, and generate more accurate output text and convert it into audio.

Benefits of technology

It improves the accuracy and efficiency of multimodal large-modal interactions, enhances the naturalness and richness of multiple rounds of real-time interactions, supports complex interaction requirements, and reduces processing delays and the number of invalid keyframes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407854A_ABST
    Figure CN120407854A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the field of artificial intelligence, and provides a multi-round real-time multi-modal large model interaction method, a related device and a storage medium, and the multi-round real-time multi-modal large model interaction method comprises the steps: converting a first target audio into a text, and obtaining a first target task description text; determining a first key frame image set based on the first target video; performing image segmentation on each key frame image in the first key frame image set to obtain an image segmentation result of each key frame image in the first key frame image set, the image segmentation result comprising a plurality of image segmentation regions and corresponding region labels; processing the first target task description text and an image segmentation result of each key frame image in the first key frame image set based on a target multi-modal large model to obtain a first output text; and converting the first output text into voice to obtain a first output audio. According to the method and the device, the multi-modal large model interaction accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence, and more specifically to a multi-round real-time multi-modal large model interaction method, related devices, and storage media. Background Art

[0002] Large language models (LLMs) have been able to generate text with correct grammar, strong persuasion, and high similarity to human-written content through their excellent natural language generation capabilities. LLMs have been applied in multiple fields, bringing convenience and efficiency improvements and playing a crucial role, such as educational assistance, content creation, language translation, programming assistance, scientific research, information retrieval, entertainment, and games. However, the accuracy of existing large model interactions is relatively low. Summary of the Invention

[0003] Embodiments of the present application provide a multi-round real-time multi-modal large model interaction method, related devices, and storage media, which can improve the accuracy of multi-round real-time multi-modal large model interactions.

[0004] In a first aspect, embodiments of the present application provide a multi-round real-time multi-modal large model interaction method, which includes: Obtain a first target audio input by a user and a first target video corresponding to the first target audio; Convert the first target audio into text to obtain a first target task description text; Determine a first set of key frame images based on the first target video, where the first set of key frame images includes multiple key frame images in the first target video; Perform image segmentation on each of the key frame images in the first set of key frame images to obtain an image segmentation result for each of the key frame images in the first set of key frame images, where the image segmentation result includes multiple image segmentation regions and corresponding region labels; Process the first target task description text and the image segmentation results of each of the key frame images in the first set of key frame images based on a target multi-modal large model to obtain a first output text; Convert the first output text into speech to obtain a first output audio.

[0005] In one embodiment, the multi-round real-time multi-modal large model interaction method further includes: Obtain a second target audio input by the user and a second target video corresponding to the second target audio, where the input time of the second target audio is later than that of the first target audio; Convert the second target audio into text to obtain a second target task description text; Determine a second set of key-frame images based on the second target video, where the second set of key-frame images includes multiple key-frame images in the second target video; Perform image segmentation on each of the key-frame images in the second set of key-frame images to obtain the image segmentation results of each of the key-frame images in the second set of key-frame images, where the image segmentation results include multiple image segmentation regions and corresponding region labels; Based on the target multi-modal large model, process the second target task description text, the image segmentation results of each of the key-frame images in the second set of key-frame images, the first target task description text, the image segmentation results of each of the key-frame images in the first set of key-frame images, and the first output text to obtain a second output text; Convert the second output text into speech to obtain a second output audio.

[0006] In one embodiment, the determining the first set of key-frame images based on the first target video includes: Divide the first target video into multiple video segments; Based on the video segments, determine a third set of key-frame images to obtain multiple third sets of key-frame images corresponding to the multiple video segments; Merge the multiple third sets of key-frame images corresponding to the multiple video segments to obtain the first set of key-frame images.

[0007] In one embodiment, the determining the third set of key-frame images based on the video segment includes: Cluster multiple video frames in the video segment to obtain multiple frame clustering clusters; Based on the central video frame image corresponding to the clustering center of each frame clustering cluster, determine the third set of key-frame images.

[0008] In one embodiment, the determining the third set of key-frame images based on the video segment includes: Obtain the motion change parameter of the video segment, where the larger the motion change parameter, the faster the motion change of the video segment; Based on the motion change parameter of the video segment, determine the target number of images corresponding to the video segment, where the larger the motion change parameter, the larger the target number of images; Extract the video frame images of the target number of images from the video segment as the third set of key-frame images.

[0009] In one embodiment, the target multi-modal large model includes a plurality of different encoding modules, a modality fusion module, and a decoding module. Processing the first target task description text and the image segmentation results of each key frame image in the first key frame image set based on the target multi-modal large model to obtain a first output text, including: Using different encoding modules in the target multi-modal large model to encode the first target task description text and the image segmentation results of each key frame image in the first key frame image set respectively, to obtain a first feature vector corresponding to the first target task description text and second feature vectors corresponding to the image segmentation results of each key frame image in the first key frame image set; Using the modality fusion module in the target multi-modal large model to perform weighted fusion on the first feature vector and the second feature vectors to obtain a fused feature vector; Using the decoding module in the target multi-modal large model to process the fused feature vector to obtain a first output text.

[0010] In a second aspect, an embodiment of the present application provides a multi-round real-time multi-modal large model interaction device, which has the function of implementing the multi-round real-time multi-modal large model interaction method provided in the above first aspect. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware.

[0011] In one embodiment, the multi-round real-time multi-modal large model interaction device includes: An acquisition module, configured to acquire a first target audio input by a user and a first target video corresponding to the first target audio; A first conversion module, configured to convert the first target audio into text to obtain a first target task description text; A determination module, configured to determine a first key frame image set based on the first target video, where the first key frame image set includes a plurality of key frame images in the first target video; A segmentation module, configured to perform image segmentation on each key frame image in the first key frame image set to obtain the image segmentation results of each key frame image in the first key frame image set, where the image segmentation results include a plurality of image segmentation regions and corresponding region labels; A processing module, configured to process the first target task description text and the image segmentation results of each key frame image in the first key frame image set based on a target multi-modal large model to obtain a first output text; A second conversion module, configured to convert the first output text into speech to obtain a first output audio.

[0012] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which includes instructions that, when running on a computer, cause the computer to execute the multi-round real-time multi-modal large model interaction method as described in the first aspect.

[0013] In a fourth aspect, an embodiment of the present application provides a computing device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the computer program, it implements the multi-round real-time multi-modal large model interaction method as described in the first aspect.

[0014] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor coupled to a transceiver of a terminal device and is used to execute the technical solution provided in the first aspect of the embodiments of the present application.

[0015] In a sixth aspect, an embodiment of the present application provides a chip system, which includes a processor for supporting a terminal device to implement the functions involved in the above first aspect. For example, generating or processing the information involved in the image processing method provided in the above first aspect.

[0016] In a possible design, the above chip system further includes a memory, which is used to store program instructions and data necessary for the terminal. The chip system may be composed of chips or may include chips and other discrete devices.

[0017] In a seventh aspect, an embodiment of the present application provides a computer program product containing instructions that, when the computer program product runs on a computer, cause the computer to execute the multi-round real-time multi-modal large model interaction method provided in the above first aspect.

[0018] In the embodiments of the present application, compared with the prior art, a first target audio input by a user and a first target video corresponding to the first target audio are obtained; the first target audio is converted into text to obtain a first target task description text; a first set of key frame images is determined based on the first target video, and the first set of key frame images includes multiple key frame images in the first target video; image segmentation is performed on each key frame image in the first set of key frame images to obtain an image segmentation result of each key frame image in the first set of key frame images, and the image segmentation result includes multiple image segmentation regions and corresponding region labels; the first target task description text and the image segmentation results of each key frame image in the first set of key frame images are processed based on a target multimodal large model to obtain a first output text; the first output text is converted into speech to obtain a first output audio. After obtaining the first target audio and the first target video corresponding to the first target audio in the present application, the audio is converted into text, and the key frame images of the first target video are extracted for segmentation processing, so that the text converted from the audio and the key frame images can be used to predict the output text and convert it into audio, which can improve the accuracy of multimodal large model interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The objectives, features, and advantages of the embodiments of the present application will become easy to understand by reading the detailed description of the embodiments of the present application with reference to the accompanying drawings. Among them: Figure 1 FIG. [X] is a schematic diagram of a multi-round real-time multimodal large model interaction system for the multi-round real-time multimodal large model interaction method in the embodiments of the present application; Figure 2 FIG. [X] is a schematic flow diagram of a multi-round real-time multimodal large model interaction method in the embodiments of the present application; Figure 3 FIG. [X] is a schematic diagram of information interaction of a multi-round real-time multimodal large model interaction method in the embodiments of the present application; Figure 4 FIG. [X] is a schematic structural diagram of a multi-round real-time multimodal large model interaction device in the embodiments of the present application; Figure 5 FIG. [X] is a schematic structural diagram of a computing device in the embodiments of the present application; Figure 6 FIG. [X] is a schematic structural diagram of a mobile phone in the embodiments of the present application; Figure 7 FIG. [X] is a schematic structural diagram of a server in the embodiments of the present application.

[0020] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] In the description and claims of the embodiments of the present application and the above-mentioned drawings, terms such as "first" and "second" are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules does not necessarily have to be limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices. The division of modules in the embodiments of the present application is only a logical division, and there may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling, direct coupling, or communication connection to each other may be through some interfaces, and the indirect coupling and communication connection between modules may be in an electrical or other similar form, which are not limited in the embodiments of the present application. Also, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed to multiple circuit modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present application.

[0022] General large models are currently widely used in application scenarios such as chat conversations, text editing, artistic creation, code writing, mathematical reasoning, and bioinformatics. Although they have created many new business models and are very powerful, after general large models are launched for users to use, there are mainly algorithm risks, data risks, and application risks in these three types of applications: translation, chat, and collaboration.

[0023] The embodiments of the present application also provide a multi-round real-time multi-modal large model interaction method, related device, and storage medium, which can be applied to a multi-round real-time multi-modal large model interaction system. The multi-round real-time multi-modal large model interaction system may include a multi-round real-time multi-modal large model interaction device, and the multi-round real-time multi-modal large model interaction device can be integrally deployed or separately deployed. The multi-round real-time multi-modal large model interaction device is at least used to obtain user input data and a plurality of first attack action types; randomly select a first attack action type from the plurality of first attack action types as the target attack action type; select an attack action from the attack action set corresponding to the target attack action type as the target attack action; process the user input data based on the target attack action to obtain target input data; and attack the model under test based on the target input data to obtain an attack result.

[0024] The solution provided by the embodiments of this application relates to technologies such as Artificial Intelligence (AI) and Machine Learning (ML). The specific description is as follows through the following embodiments: Among them, AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0025] AI technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0026] The accuracy of interaction of existing large models is relatively low.

[0027] Compared with the existing technology, in the embodiments of this application, the first target audio input by the user and the first target video corresponding to the first target audio are obtained; the first target audio is converted into text to obtain the first target task description text; a first key frame image set is determined based on the first target video, and the first key frame image set includes multiple key frame images in the first target video; image segmentation is performed on each key frame image in the first key frame image set to obtain the image segmentation results of each key frame image in the first key frame image set, and the image segmentation results include multiple image segmentation regions and corresponding region labels; the first target task description text and the image segmentation results of each key frame image in the first key frame image set are processed based on the target multi-modal large model to obtain the first output text; the first output text is converted into speech to obtain the first output audio. After obtaining the first target audio and the first target video corresponding to the first target audio in this application, the audio is converted into text, and the key frame images of the first target video are extracted for segmentation processing, so that the text converted from the audio and the key frame images can be used to predict the output text and convert it into audio, which can improve the accuracy of interaction of the multi-modal large model.

[0028] In some embodiments, referring to Figure 1 , the multi-round real-time multi-modal large model interaction method provided by the embodiments of this application can be based on Figure 1Implementation of a multi-round real-time multi-modal large model interaction system is shown. The multi-round real-time multi-modal large model interaction system may include an electronic device 100 and a memory 200. The electronic device 100 may be a server or a terminal device.

[0029] It should be noted that the server involved in the embodiments of this application may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0030] The terminal device involved in the embodiments of this application may be a device that provides voice and / or data connectivity to users, a handheld device with wireless connection capabilities, or other processing devices connected to a wireless modem. For example, a mobile phone (or "cellular" phone) and a computer with a mobile terminal. For example, it may be a portable, pocket-sized, handheld, computer-integrated, or vehicle-mounted mobile device that exchanges voice and / or data with a wireless access network. For example, a Personal Communication Service (PCS) phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA), etc.

[0031] Refer to Figure 2 , Figure 2 is a schematic flowchart of a multi-round real-time multi-modal large model interaction method provided by the embodiments of this application. This method can be executed by a multi-round real-time multi-modal large model interaction device. The method includes steps 101-106: Step 101, obtain a first target audio input by the user and a first target video corresponding to the first target audio.

[0032] In the implementation of this application, the first target audio is the voice information of the user recorded by a recording device such as a microphone. Among them, the first target audio and the first target video are played synchronously.

[0033] Specifically, during the playback of the first target video, obtain the first target audio input by the user to obtain the first target audio and the first target video corresponding to the first target audio. Among them, the recording start time of the first target audio is the same as the playback start time of the first target video, and the recording end time of the first target audio is the same as the playback end time of the first target video.

[0034] In a specific embodiment, an input video stream is obtained. During the process of playing the video stream, a first target audio input by the user is obtained, and a first target video corresponding to the first target audio is intercepted from the video stream to obtain the first target audio and the first target video corresponding to the first target audio.

[0035] Step 102: Convert the first target audio into text to obtain a first target task description text.

[0036] In the embodiment of the present application, the first target audio is input into a speech-to-text model to be converted into text, and a first target task description text is obtained. The speech-to-text (STT) model is the core implementation of speech recognition, and its goal is to automatically convert an audio signal into a text sequence. This technology is the basis for scenarios such as human-computer interaction, intelligent assistants, and speech translation. The technical framework covers multiple fields such as signal processing, machine learning, and natural language understanding. The speech-to-text model can be a Kaldi model, a SpeechRecognition model, and a Transformer-based model, which can be set according to specific situations.

[0037] Step 103: Determine a first set of key frame images based on the first target video. The first set of key frame images includes multiple key frame images in the first target video.

[0038] In a specific embodiment, the color histogram difference between two adjacent video frame images in the first target video is calculated. When the color histogram difference exceeds a preset color threshold, the two video frame images are determined as key frame images. This method can effectively capture scenes with large color changes in the video. The preset color threshold can be set according to specific situations.

[0039] In another specific embodiment, the structural similarity index (SSIM) between two adjacent video frame images in the first target video is calculated, and two video frame images with a structural similarity index lower than a first preset similarity index threshold are determined as key frame images, and frames with rich structural information are extracted as key frame images. This method is more sensitive to the structural changes of video content. The first preset similarity index threshold can be set according to specific situations.

[0040] In yet another specific embodiment, the motion intensity between two adjacent video frame images in the first target video is calculated, and two video frame images with a motion intensity higher than a preset intensity threshold are determined as key frame images, and frames with a larger motion amplitude are extracted as key frame images. This method is suitable for capturing scenes with more object movements in the video. The preset intensity threshold can be set according to specific situations.

[0041] In yet another specific embodiment, a key-frame deep learning model is constructed. The input is a sequence of video frames, and the output is the importance score of each frame. Then, the frames with high importance scores are selected as key frames. The sequence of video frames corresponding to the first target video is input into the key-frame deep learning model to obtain the importance scores of the individual video frame images in the first target video. The video frame images with importance scores higher than a preset score are determined as key-frame images. The preset score can be set according to specific circumstances.

[0042] In the embodiments of the present application, determining the first set of key-frame images based on the first target video includes: (1) Divide the first target video into multiple video segments.

[0043] In the embodiments of the present application, a sliding window of a preset length slides on the first target video at a preset step size to obtain multiple video segments of the preset length.

[0044] (2) Determine the third set of key-frame images based on the video segments to obtain multiple third sets of key-frame images corresponding to the multiple video segments.

[0045] In a specific embodiment of the present application, calculate the color histogram difference between two adjacent video frame images in the video segment. When the color histogram difference exceeds a preset color threshold, the two video frame images are determined as key-frame images.

[0046] In another specific embodiment, calculate the structural similarity index between two adjacent video frame images in the video segment, and determine the two video frame images with the structural similarity index lower than the first preset similarity index threshold as key-frame images.

[0047] In yet another specific embodiment, calculate the motion intensity between two adjacent video frame images in the video segment, and determine the two video frame images with the motion intensity higher than the preset intensity threshold as key-frame images.

[0048] In yet another specific embodiment, a key-frame deep learning model is constructed. The input is a sequence of video frames, and the output is the importance score of each frame. Then, the frames with high importance scores are selected as key frames. The sequence of video frames corresponding to the video segment is input into the key-frame deep learning model to obtain the importance scores of the individual video frame images in the video segment. The video frame images with importance scores higher than a preset score are determined as key-frame images.

[0049] In yet another specific embodiment, determining the third set of key-frame images based on the video segment includes: 1-1: Cluster the multiple video frames in the video segment to obtain multiple frame clusters.

[0050] 1-2: Determine a third key frame image set based on the central video frame image corresponding to the cluster center of each frame cluster.

[0051] In a specific embodiment, central video frame images corresponding to cluster centers of multiple frame clusters are determined as the third key frame image set.

[0052] In another specific embodiment, a structural similarity index is calculated between two adjacent video frame images in a video clip, and two video frame images whose structural similarity index is lower than a first preset similarity index threshold are determined as key frame images and placed in a fifth key frame image set to obtain a fifth key frame image set. The central video frame images corresponding to the cluster centers of the plurality of frame clusters are determined as a fourth key frame image set. The structural similarity index is calculated between each central video frame image in the fourth key frame image set and each key frame image in the fifth key frame image set, and the maximum value of the structural similarity index between each key frame image and the central video frame image is determined as the maximum similarity index corresponding to the central video frame image to obtain the maximum similarity index corresponding to each central video frame image. The central video frame images of the target number whose maximum similarity index is ranked from large to small are determined as a third key frame image set.

[0053] In another specific embodiment, the motion intensity between two adjacent video frame images in the video clip is calculated, and two video frame images with motion intensities higher than a preset intensity threshold are determined as key frame images and put into a fifth key frame image set.

[0054] In another specific embodiment, a keyframe deep learning model is constructed, with a video frame sequence as input and an importance score for each frame as output. Frames with high importance scores are then selected as keyframes. The video frame sequence corresponding to a video clip is input into the keyframe deep learning model, and the importance score of each video frame image in the video clip is obtained. Video frames with importance scores higher than a preset score are identified as keyframe images and placed into a fifth keyframe image set.

[0055] (3) Merge multiple third key frame image sets corresponding to multiple video clips to obtain a first key frame image set.

[0056] Furthermore, determining a third key frame image set based on the video clip includes: (1) Obtaining the motion change parameter of the video clip, wherein the larger the motion change parameter, the faster the motion change of the video clip.

[0057] In a specific embodiment, the motion intensity between two adjacent video frame images in a video segment is calculated to obtain multiple motion intensities in the video segment, and an average value of the multiple motion intensities is determined as the motion change parameter of the video segment.

[0058] (2) Determine the number of target images corresponding to the video segment based on the motion change parameter of the video segment.

[0059] Among them, the larger the motion change parameter, the larger the number of target images. According to the characteristics of the video content, the present application dynamically adjusts the threshold for key frame extraction. For example, for a video with fast motion changes, the threshold is lowered to extract more key frames.

[0060] (3) Extract video frame images with the number of target images from the video segment as the third key frame image set.

[0061] In a specific embodiment, cluster multiple video frames in the video segment to obtain frame cluster clusters with the number of target images, and determine the third key frame image set based on the central video frame images corresponding to the cluster centers of each frame cluster cluster. For example, determine the number of target images as the number of clusters K, and use K-means clustering to cluster multiple video frames in the video segment to obtain frame cluster clusters with the number of target images. Specifically, randomly initialize K cluster centers; assign each video frame to the nearest center to form K clusters; recalculate the mean of each cluster as the new center, and iterate until the center no longer changes significantly or reaches the maximum number of iterations.

[0062] In another specific embodiment, cluster multiple video frames in the video segment based on a preset number of images to obtain frame cluster clusters with the preset number of images, and the preset number of images is greater than the number of target images. Obtain the central video frame images with the number of target images from the central video frame images with the preset number of images to obtain the third key frame image set.

[0063] Specifically, the preset number of images is N, and multiple video frames are randomly and equally divided into Q video frame sets. Cluster the Q video frame sets respectively, and cluster each video frame set into N second video frame clusters. Specifically, use the K-means algorithm to cluster the Q video frame sets respectively, and cluster each video frame set into N second video frame clusters. Determine each second video frame cluster as a target video frame cluster, and calculate the video frame cluster similarity between the target video frame cluster and the second video frame clusters in the Q video frame sets respectively. Merge the second video frame clusters with the largest video frame cluster similarity to the target video frame cluster in each video frame set to obtain the first video frame cluster corresponding to the target video frame cluster, obtain the first video frame clusters corresponding to each second video frame cluster, and obtain N first video frame clusters corresponding to N second video frame clusters. Determine the first video frame cluster as a frame cluster cluster, and obtain N frame cluster clusters corresponding to N first video frame clusters.

[0064] Specifically, the structural similarity index between each central video frame image in the fourth key frame image set and each key frame image in the fifth key frame image set is calculated, and two video frame images whose structural similarity index is lower than a first preset similarity index threshold are determined as key frame images and added to the fifth key frame image set to obtain the fifth key frame image set. The central video frame images corresponding to the cluster centers of the multiple frame clusters are determined as the fourth key frame image set. The structural similarity index between each central video frame image in the fourth key frame image set and each key frame image in the fifth key frame image set is calculated, and the maximum value of the structural similarity index between each key frame image and the central video frame image is determined as the maximum similarity index corresponding to the central video frame image, thereby obtaining the maximum similarity index corresponding to each central video frame image. The central video frame images with the target number of images ranked highest in descending order by maximum similarity index are determined as the third key frame image set.

[0065] Furthermore, multiple threads are started to synchronously determine multiple third key frame image sets corresponding to the multiple video clips based on the multiple video clips. Multi-threading technology is used to process video frames in parallel to increase the speed of key frame extraction.

[0066] Step 104 : performing image segmentation on each key frame image in the first key frame image set to obtain an image segmentation result for each key frame image in the first key frame image set. The image segmentation result includes a plurality of image segmentation regions and corresponding region labels.

[0067] In the embodiment of the present application, each key frame image in the first key frame image set is input into an image segmentation model for image segmentation to obtain an image segmentation result for each key frame image in the first key frame image set.

[0068] Image segmentation is the process of dividing an image into multiple meaningful regions. Based on the segmentation granularity and task objectives, it can be categorized into the following: Semantic Segmentation: Classifies each pixel into a predefined category (e.g., "person," "car," or "background"), without distinguishing between different instances of the same category. Instance Segmentation: Distinguishes not only between pixel categories but also between different instances of the same category (e.g., different people or different vehicles). Panoptic Segmentation: Fusion of semantic and instance segmentation, processing both semantic categories (e.g., sky, road) and instance objects (e.g., people, cars).

[0069] The image segmentation model can perform entity segmentation and entity recognition to obtain multiple image segmentation regions and corresponding region labels. The image segmentation regions are the regions where different entities (such as people, objects, background, etc.) are located, and the region labels are the labels of the image segmentation regions. For example, the region labels are people, objects, background, etc. Among them, the image segmentation model can be MaskR-CNN, DeepLab, etc., which can be set according to specific situations.

[0070] Step 105: Based on the target multi-modal large model, process the first target task description text and the image segmentation results of each key frame image in the first key frame image set to obtain the first output text.

[0071] In the embodiments of the present application, the target multi-modal large model includes multiple different encoding modules, a modality fusion module, and a decoding module. Processing the first target task description text and the image segmentation results of each key frame image in the first key frame image set based on the target multi-modal large model to obtain the first output text includes: (1) Use different encoding modules in the target multi-modal large model to encode the first target task description text and the image segmentation results of each key frame image in the first key frame image set respectively, to obtain the first feature vector corresponding to the first target task description text and the second feature vectors corresponding to the image segmentation results of each key frame image in the first key frame image set.

[0072] In the embodiments of the present application, different encoding modules in the target multi-modal large model can be CNN, Transformer, etc. Encode the key frame pictures, segmentation maps and labels, text descriptions of tasks, and historical data respectively to obtain their feature vector representations.

[0073] (2) Use the modality fusion module in the target multi-modal large model to perform weighted fusion on the first feature vector and the second feature vector to obtain a fused feature vector.

[0074] Specifically, the modality fusion module in the target multi-modal large model can perform weighted fusion on the first feature vector and the second feature vector through early fusion, mid-term fusion, or late fusion.

[0075] Early fusion: Perform fusion before inputting into the large model, and concatenate the first feature vector and the second feature vector. Mid-term fusion: Perform fusion in the middle layer of the large model, and use the attention mechanism to perform weighted fusion on the first feature vector and the second feature vector. Late fusion: Perform fusion in the output layer of the large model, obtain the outputs of speech and video respectively, and then perform weighted summation.

[0076] (3) Use the decoding module in the target multi-modal large model to process the fused feature vector to obtain the first output text.

[0077] Generate a first output text, i.e., a text description of the task result, based on the fused feature vector. An autoregressive decoder and a non-autoregressive decoder can be used.

[0078] Furthermore, in the embodiment of the present application, the multi-round real-time multi-modal large model interaction method includes: obtaining an audio sample, a video sample corresponding to the audio sample, and an annotated output text corresponding to the audio sample; converting the audio sample into text to obtain a task description text sample; determining a set of key frame image samples based on the video sample, performing image segmentation on each key frame image in the set of key frame image samples to obtain an image segmentation result of each key frame image in the set of key frame image samples, where the image segmentation result includes a plurality of image segmentation regions and corresponding region labels; inputting the task description text sample and the image segmentation results of each key frame image in the set of key frame image samples into a target multi-modal large model to obtain a predicted output text; calculating a cross-entropy loss, a multi-task loss, and a contrast loss based on the predicted output text and the annotated output text, determining a total loss based on the cross-entropy loss, the multi-task loss, and the contrast loss, and iteratively training the target multi-modal large model based on the total loss until the total loss is less than a preset loss value. The cross-entropy loss: is used to generate a text description of the task result. The multi-task loss: if multiple tasks need to be completed simultaneously (such as object detection, image description, etc.), a multi-task loss function is used. The contrast loss: learns the corresponding relationship between different modality data.

[0079] Step 106, convert the first output text into speech to obtain a first output audio.

[0080] In the embodiment of the present application, a Text-to-Speech (TTS) model is used to convert the first output text into speech to obtain a first output audio. Play the first output audio.

[0081] Furthermore, the multi-round real-time multi-modal large model interaction method further includes: (1) Obtain a second target audio input by the user and a second target video corresponding to the second target audio, where the input time of the second target audio is later than that of the first target audio.

[0082] Specifically, obtain an input video stream. During the process of playing the video stream, obtain the first target audio input by the user, and intercept the first target video corresponding to the first target audio from the video stream according to the first target audio to obtain the first target audio and the first target video corresponding to the first target audio; then obtain the second target audio input by the user, and intercept the second target video corresponding to the second target audio from the video stream according to the second target audio to obtain the second target audio and the second target video corresponding to the second target audio.

[0083] (2) Convert the second target audio into text to obtain the second target task description text.

[0084] (3) Determine the second set of key frame images based on the second target video, where the second set of key frame images includes multiple key frame images in the second target video.

[0085] (4) Perform image segmentation on each key frame image in the second set of key frame images to obtain the image segmentation results of each key frame image in the second set of key frame images. The image segmentation results include multiple image segmentation regions and corresponding region labels.

[0086] (5) Process the second target task description text, the image segmentation results of each key frame image in the second set of key frame images, the first target task description text, the image segmentation results of each key frame image in the first set of key frame images, and the first output text based on the target multi-modal large model to obtain the second output text.

[0087] Specifically, use different encoding modules in the target multi-modal large model to encode the second target task description text, the image segmentation results of each key frame image in the second set of key frame images, the first target task description text, the image segmentation results of each key frame image in the first set of key frame images, and the first output text respectively, to obtain the first feature vector corresponding to the second target task description text, the second feature vector corresponding to the image segmentation results of each key frame image in the second set of key frame images, and the historical feature vector corresponding to the encoding of the first target task description text, the image segmentation results of each key frame image in the first set of key frame images, and the first output text. Use the modality fusion module in the target multi-modal large model to perform weighted fusion on the first feature vector, the second feature vector, and the historical feature vector to obtain the fused feature vector. Use the decoding module in the target multi-modal large model to process the fused feature vector to obtain the second output text.

[0088] (6) Convert the second output text into speech to obtain the second output audio.

[0089] This application realizes multi-modal large model interaction, can process and understand various types of data, and improves the naturalness and richness of interaction. It supports multi-round real-time interaction, can better meet the complex interaction needs of users, and improves the efficiency and practicality of interaction. Through componentized orchestration, different functional modules can be flexibly combined, facilitating the extension and customization of the system. A key frame extraction method based on speech correlation is proposed, which improves the accuracy and efficiency of key frame extraction. A processing method for multi-modal data is proposed, which can effectively fuse various types of data and improve the interaction performance of the large model.

[0090] Improved key frame extraction efficiency: Experimental data shows that compared with fixed frame extraction, the number of invalid key frames is reduced by 43%, and the processing delay is reduced to 120 ms / frame. Instruction correlation accuracy: In the COCO dataset test, the multi-modal matching strategy enables the accuracy of relevant image selection to reach 92.7% (baseline 78.2%). Multi-round dialogue consistency: The introduction of the memory network increases the task completion rate of 5 consecutive rounds of dialogue from 65% to 89%.

[0091] Further, referring to Figure 3 , Figure 3 is a schematic flowchart of a multi-round real-time multi-modal large model interaction method provided by an embodiment of the present application. This method can be executed by a multi-round real-time multi-modal large model interaction device. The method includes the steps: (1) Obtain a first target audio input by the user and a first target video corresponding to the first target audio.

[0092] In a specific embodiment, an input video stream is obtained. During the playback of the video stream, the first target audio input by the user is obtained, and the first target video corresponding to the first target audio is intercepted from the video stream according to the first target audio, so as to obtain the first target audio and the first target video corresponding to the first target audio.

[0093] (2) Convert the first target audio into text to obtain a first target task description text.

[0094] (3) Determine a first set of key frame images based on the first target video. The first set of key frame images includes multiple key frame images in the first target video.

[0095] (4) Perform image segmentation on each key frame image in the first set of key frame images to obtain the image segmentation results of each key frame image in the first set of key frame images. The image segmentation results include multiple image segmentation regions and corresponding region labels.

[0096] (5) Process the first target task description text and the image segmentation results of each key frame image in the first set of key frame images based on a target multi-modal large model to obtain a first output text.

[0097] (6) Convert the first output text into speech to obtain a first output audio.

[0098] (7) Obtain a second target audio input by the user and a second target video corresponding to the second target audio, where the input time of the second target audio is later than that of the first target audio.

[0099] (8) Convert the second target audio into text to obtain a second target task description text.

[0100] (9) Determine a second set of key frame images based on the second target video, where the second set of key frame images includes multiple key frame images in the second target video.

[0101] (10) Perform image segmentation on each key frame image in the second set of key frame images to obtain the image segmentation results of each key frame image in the second set of key frame images. The image segmentation results include multiple image segmentation regions and corresponding region labels.

[0102] (11) Based on the target multi-modal large model, process the second target task description text, the image segmentation results of each key frame image in the second set of key frame images, the first target task description text, the image segmentation results of each key frame image in the first set of key frame images, and the first output text to obtain a second output text.

[0103] (12) Convert the second output text into speech to obtain a second output audio.

[0104] Compared with the prior art, in the embodiments of the present application, a first target audio input by a user and a first target video corresponding to the first target audio are obtained; the first target audio is converted into text to obtain a first target task description text; a first set of key frame images is determined based on the first target video, where the first set of key frame images includes multiple key frame images in the first target video; image segmentation is performed on each key frame image in the first set of key frame images to obtain the image segmentation results of each key frame image in the first set of key frame images. The image segmentation results include multiple image segmentation regions and corresponding region labels; based on the target multi-modal large model, the first target task description text and the image segmentation results of each key frame image in the first set of key frame images are processed to obtain a first output text; the first output text is converted into speech to obtain a first output audio. After the present application obtains the first target audio and the first target video corresponding to the first target audio, the audio is converted into text, and the key frame images of the first target video are extracted for segmentation processing, so that the text converted from the audio and the key frame images can be used to predict the output text and convert it into audio, which can improve the accuracy of multi-modal large model interaction.

[0105] Refer to Figure 4 As Figure 4 shown, it is a schematic structural diagram of a multi-round real-time multi-modal large model interaction device. The multi-round real-time multi-modal large model interaction device in the embodiments of the present application can implement corresponding to the above Figure 2Steps of the multi-round real-time multi-modal large model interaction method executed in the corresponding embodiment. The functions implemented by the multi-round real-time multi-modal large model interaction device can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware. The multi-round real-time multi-modal large model interaction device 60 may include an acquisition module 601, a first conversion module 602, a determination module 603, a segmentation module 604, a processing module 605, and the function implementation of the second conversion module 606 can refer to Figure 2 The operations executed in the corresponding embodiment will not be elaborated here.

[0106] The multi-round real-time multi-modal large model interaction device includes: An acquisition module 601, configured to acquire a first target audio input by a user and a first target video corresponding to the first target audio; A first conversion module 602, configured to convert the first target audio into text to obtain a first target task description text; A determination module 603, configured to determine a first set of key frame images based on the first target video, where the first set of key frame images includes multiple key frame images in the first target video; A segmentation module 604, configured to perform image segmentation on each of the key frame images in the first set of key frame images to obtain an image segmentation result of each of the key frame images in the first set of key frame images, where the image segmentation result includes multiple image segmentation regions and corresponding region labels; A processing module 605, configured to process the first target task description text and the image segmentation results of each of the key frame images in the first set of key frame images based on a target multi-modal large model to obtain a first output text; A second conversion module 606, configured to convert the first output text into speech to obtain a first output audio.

[0107] In one embodiment, the multi-round real-time multi-modal large model interaction method further includes: Acquiring a second target audio input by a user and a second target video corresponding to the second target audio, where the input time of the second target audio is later than that of the first target audio; Converting the second target audio into text to obtain a second target task description text; Determining a second set of key frame images based on the second target video, where the second set of key frame images includes multiple key frame images in the second target video; Perform image segmentation on each of the key frame images in the second set of key frame images to obtain the image segmentation results of each of the key frame images in the second set of key frame images, where the image segmentation results include multiple image segmentation regions and corresponding region labels; Based on the target multi-modal large model, process the second target task description text, the image segmentation results of each of the key frame images in the second set of key frame images, the first target task description text, the image segmentation results of each of the key frame images in the first set of key frame images, and the first output text to obtain a second output text; Convert the second output text into speech to obtain a second output audio.

[0108] In one embodiment, the determining the first set of key frame images based on the first target video includes: Divide the first target video into multiple video segments; Based on the video segments, determine a third set of key frame images to obtain multiple third sets of key frame images corresponding to the multiple video segments; Merge the multiple third sets of key frame images corresponding to the multiple video segments to obtain the first set of key frame images.

[0109] In one embodiment, the determining the third set of key frame images based on the video segment includes: Cluster multiple video frames in the video segment to obtain multiple frame clustering clusters; Based on the central video frame images corresponding to the clustering centers of each frame clustering cluster, determine the third set of key frame images.

[0110] In one embodiment, the determining the third set of key frame images based on the video segment includes: Obtain the motion change parameter of the video segment, where the larger the motion change parameter, the faster the motion change of the video segment; Based on the motion change parameter of the video segment, determine the target image number corresponding to the video segment, where the larger the motion change parameter, the larger the target image number; Extract the video frame images of the target image number from the video segment as the third set of key frame images.

[0111] In one embodiment, the target multi-modal large model includes multiple different encoding modules, a modality fusion module, and a decoding module. The processing the first target task description text and the image segmentation results of each of the key frame images in the first set of key frame images based on the target multi-modal large model to obtain the first output text includes: Use different encoding modules in the target multimodal large model to encode the first target task description text and the image segmentation results of each key frame image in the first key frame image set, respectively, to obtain a first feature vector corresponding to the first target task description text and a second feature vector corresponding to the image segmentation results of each key frame image in the first key frame image set; Use the modality fusion module in the target multimodal large model to perform weighted fusion on the first feature vector and the second feature vector to obtain a fused feature vector; Use the decoding module in the target multimodal large model to process the fused feature vector to obtain a first output text.

[0112] The multi-round real-time multimodal large model interaction device 60 in the embodiments of the present application has been described above from the perspective of modular functional entities. Next, the multi-round real-time multimodal large model interaction device in the embodiments of the present application will be described from the perspective of hardware processing.

[0113] Figure 4 The devices shown can all have the structure as Figure 5 shown. When Figure 4 the multi-round real-time multimodal large model interaction device 60 shown has the structure as Figure 5 shown, Figure 5 the processor and transceiver in it can implement the same or similar functions as the modules provided in the corresponding device embodiments of the device, Figure 5 and the memory in it stores the computer program that the processor needs to call when executing the above multi-round real-time multimodal large model interaction method.

[0114] The embodiments of the present application also provide a terminal device, as Figure 6 shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal device can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS) device, an in-vehicle computer, etc. Taking the terminal device as a mobile phone as an example: Figure 6 The block diagram of a part of the structure of the mobile phone related to the terminal device provided by the embodiments of the present application is shown. Refer to Figure 6, the mobile phone includes components such as a Radio Frequency (RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, sensors 1050, an audio circuit 1060, a wireless fidelity (WiFi) module 1070, a processor 1080, and a power supply 1090. Those skilled in the art can understand that Figure 6 the mobile phone structure shown in

[0115] does not limit the mobile phone and may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. Figure 6 The following specifically introduces each component of the mobile phone: The RF circuit 1010 can be used for receiving and transmitting signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 1080 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit 1010 includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a Low Noise Amplifier (LNA), a duplexer, etc. In addition, the RF circuit 1010 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0116] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1020. The memory 1020 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 1020 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0117] The input unit 1030 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 1030 may include a touch panel 1031 and other input devices 1032. The touch panel 1031, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 1031), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 1031 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1080, and can receive and execute commands sent by the processor 1080. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 1031. In addition to the touch panel 1031, the input unit 1030 may further include other input devices 1032. Specifically, the other input devices 1032 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0118] The display unit 1040 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 1040 may include a display panel 1041. Optionally, the display panel 1041 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 1031 can cover the display panel 1041. When the touch panel 1031 detects a touch operation on or near it, it is transmitted to the processor 1080 to determine the type of touch event. Subsequently, the processor 1080 provides a corresponding visual output on the display panel 1041 according to the type of touch event. Although in Figure 6 , the touch panel 1031 and the display panel 1041 are implemented as two independent components to realize the input and input functions of the mobile phone, but in some embodiments, the touch panel 1031 and the display panel 1041 can be integrated to realize the input and output functions of the mobile phone.

[0119] The mobile phone may further include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1041 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1041 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors that the mobile phone can also be configured with, they will not be elaborated here.

[0120] The audio circuit 1060, the speaker 1061, and the microphone 1062 can provide an audio interface between the user and the mobile phone. The audio circuit 1060 can transmit the electrical signal converted from the received audio data to the speaker 1061, and the speaker 1061 converts it into a sound signal for output; on the other hand, the microphone 1062 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1060 and then converted into audio data. After the audio data is output to the processor 1080 for processing, it is sent to another mobile phone, for example, via the RF circuit 1010, or the audio data is output to the memory 1020 for further processing.

[0121] Wi-Fi belongs to short-range wireless transmission technology. Through the Wi-Fi module 1070, a mobile phone can help users send and receive emails, browse the web, and access streaming media, etc. It provides users with wireless broadband Internet access. Although Figure 6 the Wi-Fi module 1070 is shown, it can be understood that it does not belong to the essential components of the mobile phone and can be completely omitted within the scope of not changing the essence of the invention according to needs.

[0122] The processor 1080 is the control center of the mobile phone. It connects various parts of the entire mobile phone using various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 1020, and by invoking data stored in the memory 1020, it executes various functions of the mobile phone and processes data, thereby monitoring the mobile phone as a whole. Optionally, the processor 1080 may include one or more processing units; optionally, the processor 1080 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1080 either.

[0123] The mobile phone also includes a power source 1090 (such as a battery) that supplies power to each component. Optionally, the power source can be logically connected to the processor 1080 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.

[0124] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0125] In the embodiment of the present application, the processor 1080 included in the mobile phone also has the function of controlling the execution of the multi-round real-time multi-modal large model interaction method process executed by the above multi-round real-time multi-modal large model interaction device.

[0126] The embodiment of the present application also provides a server. Please refer to Figure 6 , Figure 6FIG. 0 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1100 may vary significantly due to different configurations or performances, and may include one or more central processing units (CPU) 1122 (for example, one or more processors) and a memory 1132, and one or more storage media 1130 (for example, one or more mass storage devices) for storing application programs 1142 or data 1144. Among them, the memory 1132 and the storage media 1130 may be transient storage or persistent storage. The programs stored in the storage media 1130 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1122 may be configured to communicate with the storage media 1130 and execute a series of instruction operations in the storage media 1130 on the server 1100.

[0127] The server 1100 may further include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1158, and / or one or more operating systems 1141, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, and so on.

[0128] The steps performed by the server in the above embodiments may be based on the Figure 7 structure of the server 1100 shown. For example, for example, the steps performed by the Figure 7 multi-round real-time multi-modal large model interaction device 60 shown in the above embodiments may be based on the Figure 7 server structure shown. For example, the central processing unit 1122 executes the following operations by calling instructions in the memory 1132: Obtain a first target audio input by a user and a first target video corresponding to the first target audio; Convert the first target audio into text to obtain a first target task description text; Determine a first set of key frame images based on the first target video, where the first set of key frame images includes multiple key frame images in the first target video; Perform image segmentation on each of the key frame images in the first set of key frame images to obtain an image segmentation result for each of the key frame images in the first set of key frame images, where the image segmentation result includes multiple image segmentation regions and corresponding region labels; Process the text description of the first target task and the image segmentation results of each key frame image in the first set of key frame images based on the target multi-modal large model to obtain a first output text; Convert the first output text into speech to obtain a first output audio.

[0129] In one embodiment, the multi-round real-time multi-modal large model interaction method further includes: Obtain a second target audio input by the user and a second target video corresponding to the second target audio, where the input time of the second target audio is later than that of the first target audio; Convert the second target audio into text to obtain a second target task description text; Determine a second set of key frame images based on the second target video, where the second set of key frame images includes multiple key frame images in the second target video; Perform image segmentation on each key frame image in the second set of key frame images to obtain the image segmentation results of each key frame image in the second set of key frame images, where the image segmentation results include multiple image segmentation regions and corresponding region labels; Process the second target task description text, the image segmentation results of each key frame image in the second set of key frame images, the first target task description text, the image segmentation results of each key frame image in the first set of key frame images, and the first output text based on the target multi-modal large model to obtain a second output text; Convert the second output text into speech to obtain a second output audio.

[0130] In one embodiment, the determining the first set of key frame images based on the first target video includes: Divide the first target video into multiple video segments; Determine a third set of key frame images based on the video segments to obtain multiple third sets of key frame images corresponding to the multiple video segments; Merge the multiple third sets of key frame images corresponding to the multiple video segments to obtain the first set of key frame images.

[0131] In one embodiment, the determining the third set of key frame images based on the video segments includes: Cluster multiple video frames in the video segment to obtain multiple frame clustering clusters; Determine the third set of key frame images based on the central video frame image corresponding to the clustering center of each frame clustering cluster.

[0132] In one embodiment, determining the third set of key frame images based on the video segment includes: Obtain the motion change parameters of the video segment, where the larger the motion change parameters, the faster the motion change of the video segment; Determine the number of target images corresponding to the video segment based on the motion change parameters of the video segment, where the larger the motion change parameters, the larger the number of target images; Extract the video frame images of the number of target images from the video segment as the third set of key frame images.

[0133] In one embodiment, the target multi-modal large model includes multiple different encoding modules, a modal fusion module, and a decoding module. Processing the first target task description text and the image segmentation results of each key frame image in the first set of key frame images based on the target multi-modal large model to obtain a first output text includes: Use different encoding modules in the target multi-modal large model to encode the first target task description text and the image segmentation results of each key frame image in the first set of key frame images respectively, to obtain a first feature vector corresponding to the first target task description text and a second feature vector corresponding to the image segmentation results of each key frame image in the first set of key frame images; Use the modal fusion module in the target multi-modal large model to perform weighted fusion on the first feature vector and the second feature vector to obtain a fused feature vector; Use the decoding module in the target multi-modal large model to process the fused feature vector to obtain a first output text.

[0134] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0135] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0136] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or modules, and can be in electrical, mechanical, or other forms.

[0137] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0138] In addition, in each embodiment of the embodiments of the present application, each functional module can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0139] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various alternative implementation manners.

[0140] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0141] A computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0142] The technical solutions provided by the embodiments of the present application have been introduced in detail above. In the embodiments of the present application, specific examples are used to illustrate the principles and implementation manners of the embodiments of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the embodiments of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the embodiments of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the embodiments of the present application.

Claims

1. A multi-round real-time multi-modal large model interaction method, characterized in that, The multi-round real-time multi-modal large model interaction method includes: Obtain a first target audio input by a user and a first target video corresponding to the first target audio; Convert the first target audio into text to obtain a first target task description text; Determine a first set of key frame images based on the first target video, where the first set of key frame images includes multiple key frame images in the first target video; Perform image segmentation on each of the key frame images in the first set of key frame images to obtain an image segmentation result for each of the key frame images in the first set of key frame images, where the image segmentation result includes multiple image segmentation regions and corresponding region labels; Process the first target task description text and the image segmentation results of each of the key frame images in the first set of key frame images based on a target multi-modal large model to obtain a first output text; Convert the first output text into speech to obtain a first output audio.

2. The multi-round real-time multi-modal large model interaction method according to claim 1, wherein The multi-round real-time multi-modal large model interaction method further includes: Obtain a second target audio input by a user and a second target video corresponding to the second target audio, where the input time of the second target audio is later than that of the first target audio; Convert the second target audio into text to obtain a second target task description text; Determine a second set of key frame images based on the second target video, where the second set of key frame images includes multiple key frame images in the second target video; Perform image segmentation on each of the key frame images in the second set of key frame images to obtain an image segmentation result for each of the key frame images in the second set of key frame images, where the image segmentation result includes multiple image segmentation regions and corresponding region labels; Process the second target task description text, the image segmentation results of each of the key frame images in the second set of key frame images, the first target task description text, the image segmentation results of each of the key frame images in the first set of key frame images, and the first output text based on a target multi-modal large model to obtain a second output text; Convert the second output text into speech to obtain a second output audio.

3. The multi-round real-time multi-modal large model interaction method according to claim 1, characterized in that, The determining the first set of key frame images based on the first target video includes: Divide the first target video into multiple video segments; Determine a third set of key frame images based on the video segments to obtain multiple third sets of key frame images corresponding to the multiple video segments; Merge the multiple third sets of key frame images corresponding to the multiple video segments to obtain the first set of key frame images.

4. The multi-round real-time multi-modal large model interaction method according to claim 3, characterized in that, The determining the third set of key frame images based on the video segments includes: Cluster multiple video frames in the video segment to obtain multiple frame clustering clusters; Determine the third set of key frame images based on the central video frame images corresponding to the clustering centers of each frame clustering cluster.

5. The multi-round real-time multi-modal large model interaction method according to claim 3, wherein The determining the third set of key frame images based on the video segments includes: Obtain a motion change parameter of the video segment, where the larger the motion change parameter, the faster the video segment changes in motion; Determine the number of target images corresponding to the video segment based on the motion change parameter of the video segment, where the larger the motion change parameter, the larger the number of target images; Extract video frame images with the number of target images from the video segment as the third set of key frame images.

6. The multi-round real-time multi-modal large model interaction method according to claim 1, wherein The target multi-modal large model includes multiple different encoding modules, a modality fusion module, and a decoding module. Processing the first target task description text and the image segmentation results of each key frame image in the first set of key frame images based on the target multi-modal large model to obtain a first output text includes: Use different encoding modules in the target multi-modal large model to encode the first target task description text and the image segmentation results of each key frame image in the first set of key frame images respectively, to obtain a first feature vector corresponding to the first target task description text and a second feature vector corresponding to the image segmentation results of each key frame image in the first set of key frame images; Use the modality fusion module in the target multi-modal large model to perform weighted fusion on the first feature vector and the second feature vector to obtain a fused feature vector; Use the decoding module in the target multi-modal large model to process the fused feature vector to obtain a first output text.

7. A multi-round real-time multi-modal large model interaction device, characterized in that, This multi-round real-time multi-modal large model interaction device includes: An acquisition module configured to acquire a first target audio input by a user and a first target video corresponding to the first target audio; A first conversion module configured to convert the first target audio into text to obtain a first target task description text; A determination module configured to determine a first set of key frame images based on the first target video, where the first set of key frame images includes multiple key frame images in the first target video; A segmentation module configured to perform image segmentation on each key frame image in the first set of key frame images to obtain the image segmentation results of each key frame image in the first set of key frame images, where the image segmentation results include multiple image segmentation regions and corresponding region labels; A processing module configured to process the first target task description text and the image segmentation results of each key frame image in the first set of key frame images based on the target multi-modal large model to obtain a first output text; A second conversion module configured to convert the first output text into speech to obtain a first output audio.

8. A computing device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the computer program, it implements the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, It includes instructions that, when running on a computer, cause the computer to execute the method according to any one of claims 1-6.

10. A computer program product containing instructions, the computer program product includes program instructions that, when the program instructions run on a computer or a processor, cause the computer or the processor to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Video and text mutual inspection method and device, equipment, storage medium and terminal

    CN115495615A

  • Audio generation method and device, electronic equipment and storage medium

    CN119255028A

  • Animation video generation method and device based on key frame, equipment and storage medium

    CN119342307A

  • Data detection method, device and equipment and computer storage medium

    CN119540824A

  • Video processing method and device based on multi-modal information fusion, equipment and medium

    CN119580738A