Video picture search split-screen interaction method and terminal based on multi-mode large model

By using a large multimodal model to understand video images in real time and interact with voice, the problem of video search interruption in the existing technology is solved, and seamless information acquisition and interactive experience are achieved during video playback.

CN120744174APending Publication Date: 2025-10-03NANJING KUKAI SMART SCREEN TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510793868.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing video content search technology for smart terminals cannot understand dynamic video streams in real time, causing users to interrupt playback to search for information while watching videos, affecting the user experience.

Method used

A multimodal large model is used to understand video images in real time, and split-screen interaction is achieved through voice interaction, allowing users to search and obtain relevant information while watching videos.

Benefits of technology

It enables real-time search without interruption during video playback, improves the user interaction experience, and provides instant screen information answers through voice and split-screen interaction systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744174A_ABST
    Figure CN120744174A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode large model-based video picture search split-screen interaction method and a terminal, and belongs to the technical field of intelligent terminal video interaction, and the method comprises the following steps: when a question search function is triggered and started, controlling to obtain a question search instruction, and simultaneously controlling to intercept video picture information of a front and back predetermined frame when the question search function is triggered and started; identifying the question search instruction, and identifying the intention of a user to search for video picture information; performing element identification on the intercepted predetermined frame of video picture information through a multi-mode large model, and finding out elements matched with the intention of a user to search for the video picture information; according to the found elements matched with the intention of the user and the identified intention of the user for searching the video picture information, automatically searching a search result through a multi-mode large model; and displaying through a preset split-screen interaction interface. According to the method and the device, a video picture can be deeply understood, generation type dialogue interaction can be performed on a user, and convenience is provided for use of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent terminal video interaction technology, and in particular to a video picture search and split-screen interaction method, device, intelligent terminal and storage medium based on a multimodal large model. Background Art

[0002] With the development of science and technology and the continuous improvement of people's living standards, the use of various smart terminals such as smart TVs is becoming more and more popular, and the smart terminals in the existing technology have more and more functions.

[0003] Existing video content search technologies for smart terminals primarily rely on static image matching or fixed tag retrieval, failing to provide real-time semantic understanding of dynamic video streams. Traditional systems require users to pause playback and manually capture footage, disrupting the viewing experience. When users discover a character or scene of interest while watching a film or television show, existing methods are unable to instantly access relevant information through natural language interaction. Instead, users must exit the playback interface and conduct a separate search, resulting in inefficient information acquisition and a disconnected user experience.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a video screen search split-screen interaction method, device, intelligent terminal and storage medium based on a multimodal large model. The present invention has the advantages of realizing real-time search of screen information during video playback without interrupting playback, thereby improving user interaction experience; through the knowledge coverage capability of the multimodal large model, it can understand the actual scene in the video screen, and realize the search for characters in the video screen, the title of the movie from, and the screen commentary; through split-screen interaction, while watching the video, it is possible to communicate with the TV through voice, realizing an interactive system in which the right side does not affect the viewing of the video, and the left split-screen area answers the screen information search results; the present invention can realize in-depth understanding of a video screen and generative dialogue interaction with the user, which provides convenience for user use.

[0006] The present application provides a video screen search split-screen interaction method based on a multimodal large model, and the technical solution is as follows: detecting and judging whether the smart terminal is currently in a video screen playing state; when it is detected that the smart terminal is currently in a video screen playing state, controlling to allow the start of a question search function for the video screen; when it is detected that the question search function for the video screen is triggered to start, controlling to obtain a question search instruction, and at the same time controlling to intercept a predetermined frame of video screen information before and after the question search function is triggered to start; identifying the question search instruction, and identifying the user's intention to search for video screen information; according to the identified user's intention to search for video screen information, performing element identification on the intercepted predetermined frame of video screen information through the multimodal large model, and finding elements that match the user's intention to search for video screen information from the identified elements; according to finding the elements that match the user's intention and the identified user's intention to search for video screen information, automatically searching for search results that match the user's intention to search for video screen information through the multimodal large model; and displaying the search results through a preset split-screen interactive interface.

[0007] Furthermore, the present application also proposes to pre-set the video playback screen of the smart terminal: a function for searching for information in the video screen through voice commands or text commands, and setting the searched information to be displayed through a split-screen interactive function.

[0008] Furthermore, the present application also proposes that when the smart terminal is in the state of playing a video screen, and detects the voice button pressing event of the remote control or the voice wake-up command, the question search function for the video screen is controlled to be triggered and started; when it is detected that the question search function for the video screen is triggered and started, the question search command is controlled to be obtained, and at the same time, the predetermined frames of video screen information before and after the question search function is triggered and started are captured, and saved in the current question search dialogue state.

[0009] Furthermore, the present application also proposes to perform voice recognition on the obtained question search instructions through the cloud to identify the text information and the user's intention to search for video screen information.

[0010] Furthermore, the present application also proposes to obtain the predetermined frames of video screen information before and after the question search function is triggered and started in the current conversation state, and identify the text information and the user's intention to search for video screen information, and send them to the cloud video image understanding service; judge the identified text information and the user's intention to search for video screen information to determine whether it is a search intention for the intercepted predetermined frame of video screen information; input the identified text information and the identified user's intention to search for video screen information into the multimodal large model, perform element recognition on the intercepted predetermined frame of video screen information through the multimodal large model, and find out the elements that match the user's intention to search for video screen information from the identified elements.

[0011] Furthermore, the present application also proposes to obtain the elements that match the user's intention and the identified intention of the user to search for video screen information; based on the elements that match the user's intention and the identified intention of the user to search for video screen information, automatically search for search results that match the user's intention to search for video screen information through a multimodal large model; and control the multimodal large model to send the search results in text form.

[0012] Furthermore, the present application also proposes that the intelligent terminal receives search results sent in text form, controls the reduction of the currently playing video screen, and controls the display of a pre-set split-screen search result streaming result display area; controls the split-screen search result streaming result display area on one side of the mobile terminal to display the search results, and the video area on the other side continues to play the video screen, thereby realizing a multimodal large model image search split-screen effect.

[0013] Furthermore, the present application also proposes a video screen search split-screen interaction device based on a multimodal large model, the device comprising: a playback state detection module for detecting whether the intelligent terminal is currently in the video screen playback state; a question search function trigger module for controlling the start of the question search function for the video screen when it is detected that the intelligent terminal is currently in the video screen playback state; a question search trigger module for controlling the acquisition of the question search instruction when it is detected that the question search function for the video screen is triggered and at the same time controlling the capture of the predetermined frame video screen information before and after the question search function is triggered and started; an intention recognition module for controlling the question search instruction Identify and recognize the user's intention to search for video screen information; a matching module is used to identify elements of the captured predetermined frame video screen information through a multimodal large model based on the identified user's intention to search for video screen information, and find elements that match the user's intention to search for video screen information from the identified elements; an automatic search module is used to automatically search for search results that match the user's intention to search for video screen information through a multimodal large model based on the elements that match the user's intention and the identified user's intention to search for video screen information; a split-screen display module is used to display the search results through a preset split-screen interactive interface.

[0014] Furthermore, the present application also proposes an intelligent terminal comprising a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, wherein the one or more programs include those for executing the above method.

[0015] Furthermore, the present application also proposes a computer-readable storage medium, which enables the electronic device to perform the above method when instructions in the storage medium are executed by a processor of the electronic device.

[0016] From the above, it can be seen that the present application provides a video screen search split-screen interaction method, device, intelligent terminal and storage medium based on a multimodal large model. Through real-time detection of video playback status, intelligent capture of key frames, multimodal large model analysis and split-screen display technology, accurate information search can be achieved while maintaining the continuity of video playback. It has the advantages of realizing real-time search of screen information during video playback without interrupting playback, thereby improving user interaction experience.

[0017] This technology leverages a multimodal large model to understand and search for visual information within videos, allowing users to describe their search queries in natural language. Whether searching for the source film, character information, or objects or buildings within a video, users can express their search queries based on their conversations. Through interactive left-right split-screen interaction, the large model on the left provides real-time answers to questions about the video content without exiting the current video. This makes video search more intuitive and convenient, significantly improving the user experience.

[0018] The comprehension capabilities of a multimodal large model allow users to search for videos using everyday language, without having to strictly adhere to specific formats, keywords, or fields, reducing the user's cognitive burden. Compared to existing technologies, this application not only solves the problem of users' complex video search needs, but also provides a more intelligent, natural, and efficient search experience, thus bringing users a brand new way to search for videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 It is a flow chart of the video picture search split-screen interaction method based on the multimodal large model provided in Example 1 of the present invention.

[0021] Figure 2 It is a flow chart of the video picture search split-screen interaction method based on the multimodal large model provided in Example 2 of the present invention.

[0022] Figure 3 A principle block diagram of an embodiment of a video picture search split-screen interaction device based on a multimodal large model provided by the present invention.

[0023] Figure 4 This is a block diagram of the internal structure of the smart terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0025] It should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0026] With the prevalence of smart devices and the explosive growth of video content, users often need to search for content while watching videos. Existing traditional video search technologies have significant limitations: First, existing systems only support simple searches based on fixed images or preset keywords, failing to understand dynamic scenes and complex content within videos. Second, the search process often requires interrupting video playback, severely impacting the viewing experience. Taking smart TVs as an example, existing technology can only search for celebrity information within a fixed scene library by matching static images. It lacks a deep understanding of continuous video footage and cannot support interactive searches using natural language. These technical limitations make it difficult for users to quickly access relevant information when encountering content of interest while watching, severely limiting the interactive experience of smart devices. Furthermore, existing solutions lack an effective split-screen interaction mechanism, forcing users to pause video playback to retrieve search results. This operating mode neither meets the modern user's demand for a smooth experience nor accommodates multitasking scenarios.

[0027] To address these issues, existing technologies suffer from interaction lag and insufficient semantic understanding when processing dynamic video content searches. Analysis revealed that dynamic image information exhibits spatiotemporal continuity, making it difficult for a single frame to convey complete semantics. Research has shown that large multimodal models possess cross-modal association capabilities, enabling simultaneous analysis of visual elements and natural language commands. Further exploration revealed that combining keyframe extraction from video streams with voice interaction enables real-time semantic capture. Furthermore, a split-screen display mechanism was designed to maintain video playback continuity, forming a complete instant search solution.

[0028] Therefore, the present application proposes a split-screen interactive method for searching video images based on a multimodal large model. According to the interactive characteristics of existing smart TVs, the present application needs to meet the needs of users' more complex video image search and analysis scenarios to realize a complex ability to search for information in the video image through voice dialogue, and to respond to the dialogue of the video image through a multimodal large model, so as to achieve an intelligent experience of searching, analyzing and answering the information in the video while watching the video. The key technology is to solve the problem that in the TV system, when watching the video, the TV can search for information in the video image through voice expression, and answer the information through the search information results, without affecting the watching of the film and television, thereby realizing the experience of split-screen interactive video search.

[0029] For example, take all current smart TVs with voice capabilities as an example, as shown in the following scenario: Existing smart TVs all have a voice input feature for remote control dialogue. However, when watching a video, users typically see the video content appearing on the screen, but cannot search the dynamic video screen through text input or voice interaction, thus failing to meet user video search scenarios. For example, scenario one: a user is watching a short video that only recommends a highlight clip. After watching this highlight clip, the user wants to know which movie it is from. With the present invention, the user can interact with the system through voice, saying, "This video is from this movie!" Scenario two: a user sees a person in a video but can't remember who they are. With the present invention, the user can interact with the system through voice, saying, "Who is the man with the gun in this video?" Scenario three: a user sees a beautiful car in a video but doesn't know the brand. With the present invention, the user can ask, "What brand is the car that just drove by in the video?" Scenario four: a user sees a magnificent landmark in a video and wants to know where it was filmed. With the present invention, the user can ask, "Where is the tallest building in the video?" The present invention can realize the split-screen interaction technology and method of voice search for the video screen of the smart TV based on the multimodal large model. As long as the user wants to search for information seen through the video, it can be achieved through voice dialogue. Moreover, through the split-screen interaction mode, the current video viewing mode can be exited without interruption. A screen area is divided on the left to display the multimodal large model search result information, and the video window on the right is appropriately reduced to continue playing the video.

[0030] Returning to the user video search scenario mentioned above, using the present invention's application method, scenario one involves the following: "Which movie is this video from?" Suppose the user is watching a compelling video from the movie "Kung Fu." Using a multimodal large model, the video can be searched for. Through split-screen interaction, the left-hand side of the split-screen area displays the search results for the user's video using text: "This video is from the classic movie "Kung Fu," starring Stephen Chow." Scenario two involves the following: "Who's the man with the gun in this video?" Suppose the user is watching an action shootout video featuring Andy Lau. The left-hand side of the split-screen area can then display relevant results: "The man holding the gun in the video is Andy Lau," along with Andy Lau's biography and other information. ....; Scenario three: "What brand is the car that just drove past in the video?", assuming that a video of an Audi car drifting is just being played, the split-screen area on the left can search out relevant results: This car is an Audi model xxx, and the car's brief introduction information...; Scenario four: "Where is the tallest building in the video?", assuming that a scene video of the Oriental Pearl Tower in Shanghai is just being played, the split-screen area on the left can search out relevant results: The tallest building in the picture is the Oriental Pearl Tower, a landmark building in Shanghai, and the introduction information of the Oriental Pearl Tower; thus satisfying the user's intelligent search for video content without interrupting the split-screen interactive experience of the current video being watched.

[0031] Example 1 like Figure 1 As shown, a video image search and split-screen interaction method based on a multimodal large model in embodiment 1 of the present invention includes the following steps: Step S100: Detect and determine whether the smart terminal is currently in a video playback state; In the embodiment of the present invention, the smart terminal may be a television for example, and the played video screen may be a television playing various video screens.

[0032] Step S200: When it is detected that the smart terminal is currently playing a video, the control allows the start of a question search function for the video; In an embodiment of the present invention, when it is detected that a smart terminal such as a smart TV is currently playing a video, a question search function for the video can be started at any time through a voice button on a remote control or a separate voice start command.

[0033] Step S300: When it is detected that the question search function for the video screen is triggered and started, the question search instruction is obtained and the video screen information of the predetermined frames before and after the question search function is triggered and started is captured; In the embodiment of this step, for example, when it is detected that a user operates and presses the voice button of the remote control, or it is detected that a user issues a separate voice start command "Xiaoman Xiaoman, help me check which movie this video is from", it is determined that the question search function for the video screen is triggered and started; at this time, the predetermined frames of video screen information before and after the question search function is triggered and started will be controlled to be intercepted, for example, the 5 frames of video screen information before and after the current video screen when the question search function is triggered and started will be controlled to be intercepted.

[0034] Step S400: Identify the question search instruction and identify the user's intention to search for video information; In the embodiment of this step, the question search instruction spoken by the user is recognized. For example, when the user's question search instruction is a voice instruction, the voice question search instruction is recognized by voice-to-text conversion, and then the intention is recognized through the text. For example: the user's voice text result is "Which movie is this video clip from?", the intention recognition can identify that the user intends to search for video information.

[0035] Step S500: Based on the identified user's intention to search for video image information, perform element recognition on the intercepted predetermined frame of video image information using the multimodal large model, and find elements that match the user's intention to search for video image information from the identified elements; In the embodiment of this step, based on the identified intention of the user to search for video screen information, the multimodal large model is used to perform element recognition on the intercepted predetermined frame video screen information, and elements that match the user's intention to search for video screen information are found from the identified elements. For example, when the user's voice command recognizes that the user's intention to search for video screen information is "I want to identify the brand of the car that just drove past in the video", assuming that a video of an Audi car drifting is just being played, then in the embodiment of the present invention, the multimodal large model is used to perform element recognition on the intercepted predetermined frame video screen information, and elements that match the user's intention to search for video screen information are found from the identified elements: elements related to the car that just drove past in the intercepted video screen.

[0036] Step S600: Based on the found elements matching the user's intention and the identified user's intention to search for video image information, automatically searching for search results matching the user's intention to search for video image information using the multimodal large model; Continuing from the above, in the embodiment of this step, based on finding the elements that match the user's intention, "elements related to the car that just drove past in the captured video footage," and identifying the user's intention to search for video footage information, "to identify the brand of the car that just drove past in the video," the multimodal large model is controlled to automatically search for the brand of the car that is "elements related to the car that just drove past in the captured video footage," and find that the brand of the car is an Audi model xxx.

[0037] Step S700: Display the search results through a preset split-screen interactive interface.

[0038] This step embodiment is as described above, and controls the display of the searched related results in the split-screen area of ​​the mobile terminal display screen, such as the left area: this car is an Audi xxx model car, and the brief introduction information of this car.

[0039] The above-mentioned embodiment of the present application proposes a technical solution for activating the search function after detecting the video playback status of the smart terminal, synchronously capturing related video frames when obtaining user question instructions, identifying screen elements and search intentions through a multimodal large model, and finally presenting search results on a split-screen interface.

[0040] Among them, video screen status detection refers to determining whether the device is in an active playback state by parsing the output signal of the video decoder or monitoring the screen rendering process. This can be achieved by polling the operating system layer API interface. This detection mechanism ensures that the function is only activated in valid usage scenarios. Question search command interception includes voice waveform acquisition and text conversion processing. Specifically, the endpoint detection algorithm can be used to identify the valid command interval. This process accurately captures the user's intention. The capture of predetermined frame video screen information covers multiple key frames before and after the trigger moment. Specifically, it can be implemented by video stream buffer access technology. This design ensures the integrity of spatiotemporal correlation information. Multimodal large model element recognition refers to the joint embedding analysis of visual features and text semantics. Specifically, it can be implemented by visual language pre-training model. This technology supports cross-modal semantic matching. Split-screen interactive interface display involves dynamic division of screen space. Specifically, the graphical interface rendering engine can be used to adjust the layout parameters. This method maintains the visible area of ​​the main playback screen.

[0041] Specifically, when the system detects the video playback state, it continuously monitors voice wake-up commands or physical button signals. After the user triggers the search function, the device simultaneously performs three operations: collects voice commands and converts them into text, intercepts the video frame sequence containing action coherence before and after the user triggers the search, and maintains the video playback process without interruption. The multimodal large model classifies the intent of the text command, and at the same time parses the visual elements in the video frame sequence to establish a semantic association map. For example, when the user asks "the historical background of the building in the current picture", the model simultaneously identifies the architectural features and the time and space qualifiers in the voice command, and presents the relevant historical data in the split-screen area without interrupting the video playback.

[0042] Compared to existing technologies, traditional systems are limited to static image retrieval and fixed keyword matching, and are unable to process the spatiotemporal correlations in dynamic video streams. This solution overcomes the information limitations of a single frame through multimodal feature fusion and real-time frame sequence analysis. Existing split-screen technology is mostly used for parallel application display and fails to fully integrate with semantic search. This solution creatively combines split-screen interaction with multimodal understanding to create a seamless viewing and search experience.

[0043] Through the above technical solution, this application implements a natural language interactive search function during video playback, accurately capturing the spatiotemporal information of dynamic images, maintaining viewing continuity while providing real-time knowledge services. This solution solves the problems of operational interruption and one-sided semantic understanding existing in traditional technologies, improving the efficiency of information acquisition and the naturalness of interaction with video content.

[0044] The present application further proposes to pre-set the function of searching for information in the video screen through voice commands or text commands in the video playback screen of the smart terminal, and to display the searched information through a split-screen interactive function.

[0045] The voice command or text command search function refers to triggering the retrieval of video content through natural language input. This can be achieved by using the built-in microphone of the smart terminal or the voice button on the smart TV remote control to collect voice signals, and then converting them into text commands through a cloud-based voice recognition interface, or by entering text commands through a virtual keyboard. The split-screen interactive display function divides the screen into two independent areas. Specifically, graphical interface rendering technology can be used to dynamically scale the video playback window, and a fixed-ratio display area is opened on one side of the screen. Parallel content display is achieved through the operating system-level multi-tasking mechanism.

[0046] Specifically, during the initialization phase of the video playback application, the system calls the interface to expand the functionality of the user interface and embeds a voice recognition component and a text input entry in the playback control bar. When it is detected that the user has activated the search function, the preset event handler is triggered to logically isolate the current playback screen from the interactive interface. The split-screen display module automatically calculates the display area ratio based on the device resolution. For example, it scales the main video screen to 70% of the right side of the screen, while creating a semi-transparent overlay in the 30% area on the left to present the search result list. This process is implemented through real-time rendering by the graphics processing unit to ensure that the video decoding and interface drawing threads do not interfere with each other.

[0047] Compared to existing technologies, traditional video terminals only support keyword-based static image retrieval and are unable to dynamically capture video frames for analysis during playback. This solution natively integrates a multimodal search portal into the playback interface, eliminating the need for users to switch applications. Furthermore, the split-screen mechanism breaks the limitations of a single full-screen display, allowing for simultaneous information acquisition and content consumption.

[0048] Through the above technical solution, this application realizes the integration of video playback and information retrieval functions, allowing users to obtain extended information related to the screen without interrupting the viewing process. The split-screen layout design effectively solves the problem of traditional pop-up windows blocking the screen, ensuring the complete visibility of the core video content, while providing an interactive information display area, significantly improving the efficiency of human-computer interaction.

[0049] The present application further proposes that when the smart terminal is in the state of playing a video screen and detects the voice button pressing event of the remote control or the voice wake-up instruction, the question search function for the video screen is controlled to be triggered and started; when the question search function for the video screen is detected to be triggered and started, the question search instruction is controlled to be obtained, and at the same time, the predetermined frames of video screen information before and after the question search function is triggered and started are captured and saved in the current question search dialogue state.

[0050] Among them, the remote control voice key press event refers to the event of triggering the voice input function through a physical button. Specifically, it can be implemented by an infrared signal receiving module or a Bluetooth communication module to ensure that the user can quickly activate the search function through the hardware device. Among them, the voice wake-up instruction refers to the instruction to activate the voice recognition function through a specific keyword. Specifically, it can be implemented by an offline voice recognition engine or a cloud-based voice recognition service to achieve contactless operation. Among them, intercepting the video screen information of the predetermined frames before and after the question search function is triggered (for example, 5-15 frames before and after) refers to obtaining a fixed number of video frames before and after the trigger moment. Specifically, it can be implemented by the frame data extraction technology of the video buffer to retain the context screen when the search intention is generated. Among them, saving in the current question search dialogue state refers to associating the intercepted video frame with the user's question instruction and storing it. Specifically, it can be implemented by memory cache or temporary database technology to maintain data consistency during the search process.

[0051] Specifically, when the user triggers voice input through the physical button of the remote control or speaks the preset wake-up word, the system will immediately start the video screen search function. At this time, the video player will automatically extract the screen data of several frames before and after the trigger moment, for example, it can be a video image of 5 frames before the trigger and 15 frames after the trigger. These video frames and the voice command issued by the user are synchronously stored in a temporary session container to form a search context data packet containing time correlation. As a result, the system can accurately restore the visual scene when the search is triggered when processing the user's question.

[0052] Compared to existing technologies, traditional video search functions typically only capture the current playback screen as search basis, failing to capture changes in the screen before and after the user's question. This solution, by dynamically capturing multiple frames before and after the trigger moment, effectively avoids missing screen information due to video playback progress and ensures the integrity of the search basis.

[0053] Through the above technical solution, this application solves the problem of incomplete search information caused by capturing images at a single point in time, ensuring that the multimodal large model can accurately understand the user's intent based on a continuous sequence of images when performing element recognition. For example, when a user asks about a moving object in a dynamic scene, the continuous images of the previous and next frames can provide the model with the key data required for motion trajectory analysis, thereby improving the relevance and accuracy of search results.

[0054] The present application further proposes to perform voice recognition on the obtained question search instructions through the cloud to identify the text information and the user's intention to search for video screen information.

[0055] Among them, the question search instruction refers to the query request about the content of the video screen input by the user through voice. Specifically, it can be achieved by using a microphone to collect audio signals, converting them into digital signals, and uploading them to the cloud server. For example, the user's voice data is collected in real time through the voice input module of the smart terminal. Voice recognition in the cloud refers to the conversion of voice signals into processable text information. Specifically, a pre-trained voice recognition model can be used to decode the audio data, such as an end-to-end voice recognition system built based on a deep neural network. Text information refers to the text content obtained after voice recognition. Specifically, natural language processing technology can be used to perform semantic analysis on the recognition results, such as extracting key entities through part-of-speech tagging and syntactic analysis. The user's intention to search for video screen information refers to the core query target hidden in the question instruction. Specifically, an intention classification model can be used to perform multi-label classification on the text, such as a recurrent neural network structure based on an attention mechanism to perform intent recognition on user questions.

[0056] Specifically, when a user triggers the question-and-search function via voice, the smart terminal uploads the collected audio data to a cloud server. The cloud server then uses a speech recognition service to decode the audio and generate the corresponding text information. The system then performs semantic analysis on the text information and uses an intent recognition model to determine the type of video content the user is searching for, such as character identification, scene source location, or object information query. For example, when a user asks, "What other movies has this actor acted in?" the speech recognition module converts the query into text, and the intent classification module identifies the query as information about the actor's films.

[0057] Compared to existing technologies, traditional video search systems typically rely on local speech recognition modules for command processing, which is limited by device computing power and makes it difficult to achieve high-precision intent analysis. This solution, however, employs a cloud-based collaborative processing mechanism, leveraging server-level computing resources for speech-to-text conversion and deep semantic understanding, supporting more complex natural language expressions and diverse query intent recognition. For example, while existing technologies may only recognize simple keywords and fail to process complex sentences, this solution uses multi-level cloud-based processing to accurately parse long sentences containing multiple modifiers.

[0058] Through the above technical solutions, this application can effectively improve the accuracy of user command recognition in video screen search scenarios, and solve the problem of misjudgment of intent caused by insufficient local processing capabilities in the existing technology. Through cloud-coordinated speech recognition and semantic analysis, the system can adapt to different accents, speaking speeds and complex sentence structures, ensuring that the user's search intent is fully captured and converted into executable query conditions, providing accurate input for subsequent multimodal matching. For example, for compound query instructions containing time adverbs and place adverbs, this solution can accurately separate the core search target and limiting conditions, avoiding search result deviations caused by local recognition errors.

[0059] The present application further proposes obtaining the predetermined frames of video screen information before and after the question search function is triggered and started in the current conversation state, and identifying the text information and identifying the user's intention to search for video screen information, and sending them to the cloud video image understanding service; judging the identified text information and identifying the user's intention to search for video screen information to determine whether it is a search intention for the intercepted predetermined frame of video screen information; inputting the identified text information and identifying the user's intention to search for video screen information into a multimodal large model, performing element recognition on the intercepted predetermined frame of video screen information through the multimodal large model, and finding elements that match the user's intention to search for video screen information from the identified elements.

[0060] Among them, the current dialogue state refers to the contextual information saved by the system when the user triggers the question search function. Specifically, it can be implemented by temporary storage or caching technology to associate user intentions with corresponding video screen segments to ensure the continuity of the search processing process. Among them, the cloud-based video image understanding service refers to the computing resources deployed on the remote server. Specifically, it can be implemented by a distributed architecture to process the correlation analysis between video screen information and user intentions, and solve the problem of insufficient computing power of local devices. Among them, the multimodal large model refers to a deep learning model that can process text, image and video data at the same time. Specifically, it can be implemented by a pre-trained visual-language joint model. By identifying elements in the video screen through cross-modal feature fusion, the matching accuracy of search intent and screen content is improved.

[0061] Specifically, when the user triggers the video screen search function through voice or text, the system will upload the captured video screen clip and the identified user intention to the cloud at the same time. The cloud service first determines whether the user's intention is for the currently captured video screen. For example, it can analyze whether the user's question contains keywords related to the content of the screen. Subsequently, the multimodal large model recognizes elements of the video frame. For example, it can extract the characters, objects or scene features in the picture and match them with the semantic information in the user's intention. For example, if the user asks "Who is the actor in the picture", the model will prioritize identifying the facial features in the video frame and compare them with the actor information in the database, and finally output the matching results.

[0062] Compared to existing technologies, traditional video search techniques typically only support retrieval based on fixed keywords or static images and are unable to dynamically analyze contextual information within video clips. However, this solution, by combining user intent with a large multimodal model, can identify semantic elements based on real-time video content. For example, even if the user doesn't explicitly describe the details of the scene, accurate matching can still be achieved through model inference, avoiding the search failures often caused by missing keywords in traditional technologies.

[0063] Through the above technical solution, this application achieves a deep understanding of dynamic video content and intent matching, overcoming the limitations of existing video searches that rely on fixed keywords or static images. Through the collaborative processing of cloud services and multimodal models, users can obtain accurate search results without having to precisely describe the details of the image. For example, while watching a movie, users can directly ask the name of a building in the image, and the system will automatically identify and return relevant information, significantly reducing the complexity of search operations.

[0064] The present application further proposes a split-screen interaction method for video screen search based on a multimodal large model, including: obtaining elements that match the user's intention and identifying the user's intention to search for video screen information; automatically searching for matching search results through the multimodal large model based on the above elements and intention; and controlling the multimodal large model to send the search results in text form.

[0065] Among them, the elements that match the user's intention refer to the key entities extracted after semantic analysis of the video screen through a multimodal large model, such as objects or scenes such as people, vehicles, and bags. Specifically, this can be achieved by combining a visual feature extraction algorithm with a natural language processing model to establish an association between search conditions and video content. Multimodal large model automatic search refers to the process of cross-modal matching of text intent with visual elements. Specifically, it can be achieved by using a pre-trained graphic and text joint embedding model to generate structured information related to user needs. The distribution of search results in text form refers to converting the search results into a text data stream, which can be encapsulated in JSON format and transmitted to the terminal device to achieve low-latency interactive response.

[0066] Specifically, when a query search command is received, video footage from predetermined frames before and after the trigger is captured and fed into a large multimodal model. This model performs object detection and semantic segmentation on the video frames, extracting the physical elements within them. Simultaneously, the natural language understanding module converts user intent into structured query conditions. The large multimodal model calculates similarity between visual elements and textual conditions, filtering out candidate results that meet threshold requirements. Finally, the search results are converted into plain text and transmitted to the terminal device via a pre-defined interface.

[0067] Compared to existing technologies, traditional video search methods only support static image matching based on fixed keywords and are unable to process contextual associations within dynamic video frames. This solution, however, leverages a large multimodal model to achieve cross-modal semantic understanding, capturing temporal information changes within video footage. For example, when a user asks for the name of a building that just appeared, the system automatically correlates the consecutive frames before and after the search trigger, identifies the building's exterior features, and outputs accurate results.

[0068] Through the above technical solution, this application solves the problem of semantic matching between dynamic video content and user natural language queries, avoiding the search failures caused by missing keywords in traditional methods. Transmitting search results in text format reduces data volume, allowing the split-screen interactive interface to update information in real time while maintaining smooth video playback. Users can obtain relevant answers to the screen without interrupting viewing, shortening the operation path by at least two interaction steps.

[0069] The present application further proposes the steps of displaying the search results through a preset split-screen interactive interface, including: the smart terminal receives the search results sent in text form, and controls the reduction of the currently playing video screen, and controls the display of the pre-set split-screen search results streaming result display area; controls the split-screen search results streaming result display area on one side of the mobile terminal to display the search results, and the video area on the other side continues to play the video screen, thereby realizing a multi-modal large model image search split-screen effect.

[0070] A split-screen interface is a layout that divides the screen into two (or more) independent display areas. This can be achieved by dynamically adjusting the window size ratio, for example by calling the split-screen API provided by the operating system to scale the video playback window. This interface design allows search results and video images to be presented in parallel, preventing users from interrupting their video viewing to view search results.

[0071] The split-screen search results streaming display area is a visualization module used to display the output of a multimodal large model in real time. This can be implemented as a scrolling text box or a card-style information panel. This area uses asynchronous rendering technology to load search results one by one, ensuring that users receive instant information feedback.

[0072] Specifically, when the smart terminal receives the text-based search results sent from the cloud, the video player will reduce the current full-screen playback image to 50%-70% of the original size, for example, by calling the graphics processing unit for real-time resolution adjustment. At the same time, a vertical information display area with a width of 30%-50% of the screen is generated on the other side of the screen. This area runs independently from the video playback process through an event-driven mechanism. The video decoder continuously outputs a compressed low-resolution video stream to the reduced playback window, while the search result display module parses the text data in an asynchronous thread manner and renders it into interactive mixed text and graphics content. The two display areas achieve synchronous screen refresh through the hardware acceleration layer to ensure that users do not perceive any freezes or delays in video playback when browsing search results.

[0073] Compared with existing technologies, traditional video search solutions require completely pausing video playback and jumping to an independent information page, interrupting the user's viewing experience. This solution uses dynamic split-screen technology to achieve parallel presentation of video images and search results, allowing users to obtain extended information related to the screen content in real time while maintaining video playback continuity. The fixed-ratio split-screen mode in existing technologies cannot adapt to terminal devices of different sizes. This solution uses an adaptive layout algorithm to automatically optimize the split-screen ratio based on screen orientation and resolution.

[0074] Through the above technical solution, this application solves the technical problem of screen interruptions caused by users viewing search results while watching a video, and achieves seamless collaborative display of video content and search information. This solution allows users to maintain the continuity of video playback in scenarios where they need to obtain screen-related information, such as simultaneously querying actor information while watching a film or TV series without affecting the plot. The interactive nature of the search results display area further supports users in performing operations such as information filtering and content collection, forming a complete video-enhanced interactive experience.

[0075] The present invention is further described in detail below through another specific application embodiment.

[0076] Example 2 like Figure 2 As shown, the second embodiment provides a video image search and split-screen interaction method based on a multimodal large model, including: S10, determining whether the smart terminal, such as a smart TV, is currently playing a video, if not, proceeding to S11, if yes and detecting a remote control voice press event in S12, proceeding to S13; S11: Do not perform video processing if the video is not in the scene; S12: Detecting a remote control voice press event and proceeding to S13 and S14; S13. The smart terminal device captures 5 frames of video information before and after the voice button is pressed, and saves them in the current conversation state of the terminal; S14: Perform voice recognition and intent recognition through the cloud, and then enter S15; S15: The cloud generates intent recognition results and proceeds to S16; S16: The cloud sends the intent and the user's voice recognition text results to the smart terminal, and then proceeds to S17; S17, determining whether the user's speaking intention is to search for video images, if not, proceed to S18, if yes, proceed to S19; S18. Execute other intention result processes normally; S19, obtaining the five frames of video information before and after the current conversation state and the text of the speech recognition result, and then proceeding to S20; S20, upload the image information and user voice result text to the cloud video image understanding service and enter S21; S21: The user's voice-expressed text search information and the video image are input into the multimodal large model, and the process proceeds to S22; S22, the multimodal large model video search results are sent in text streaming format and enter S23; S23, the smart terminal starts receiving video search text results and enters S24; S24, zooming out the currently playing video screen and entering S25; S25, displaying the split-screen search results streaming result display area on the left, and proceeding to S26; S26. The video search results are displayed on the left side of the smart terminal, and the video screen continues to be played in the video area on the right side, achieving a multi-modal large model image search split-screen effect.

[0077] Specifically, in the embodiment of the present invention, first, when the user presses the voice button on the remote control to prepare to speak, the intelligent terminal device determines whether a video is currently playing. If not, no video processing is performed. If a video is currently playing, the intelligent terminal device begins to capture information about the five frames of video before and after the voice button is pressed and stores it in the terminal's current conversation state. The reason why the video image capture must begin when the voice button is pressed in the embodiment of the present invention is that the image the user is searching for is the image they are viewing when they begin to speak. If the user waits until they have finished speaking before capturing the video image information, the video image will have already played and is no longer the video scene the user is searching for.

[0078] At the same time, the user voice recognition process is carried out at the same time. The smart terminal will transmit the user's voice to the cloud for recognition in real time. The cloud will perform voice-to-text recognition, and then the cloud will perform intent recognition through text. For example: the user's voice-to-text result is "Which movie is this video clip from?", and the intent recognition can recognize that the user wants to search for video information, and send the intent and the user's voice-to-text result to the device terminal. The device terminal performs logical distribution according to the received intent. If it is not a video screen search intent, it will go through other intent processes normally. If it is a video screen search intent, the device terminal obtains the previously saved 5 frames of video screen information and the user's expression The cloud receives the request information and calls the multimodal large model with the video screen information and the search text information. The multimodal large model begins to understand the search based on the video screen information and the text of the search expression, and streams the text results to the device terminal. The terminal begins to receive the search results and at the same time, it shrinks the currently playing video screen and divides the current device terminal interaction into left and right display areas. The left side dynamically displays the video screen search results of the multimodal large model in real time, while the right side continues to play the video watched by the user, thus realizing the split-screen interactive effect of the multimodal large model for video screen search.

[0079] Exemplary devices like Figure 3 As shown, an embodiment of the present invention provides a video picture search split-screen interaction device based on a multimodal large model, the device comprising: The playback state detection module 310 is used to detect whether the smart terminal is currently playing a video image; The question search function triggering module 320 is used to control and allow the start of the question search function for the video screen when it is detected that the smart terminal is currently in the video screen playing state; The question search trigger module 330 is used to control the acquisition of the question search instruction when detecting that the question search function for the video screen is triggered and to control the interception of the video screen information of the predetermined frames before and after the question search function is triggered; Intention recognition module 340, for recognizing the question search instruction and identifying the user's intention to search for video image information; Matching module 350 is used to identify elements of the captured predetermined frame of video image information based on the identified user's intention to search for video image information using a multimodal large model, and find elements that match the user's intention to search for video image information from the identified elements; An automatic search module 360 ​​is configured to automatically search for search results that match the user's intent to search for video image information based on the elements found and the identified user's intent to search for video image information using a multimodal large model; The split-screen display module 370 is used to display the search results through a preset split-screen interactive interface.

[0080] The playback status template refers to a functional unit that detects playback status by monitoring the operating status of the smart terminal player or calling system interfaces. This can be implemented using a system event listener or a player status polling mechanism. It determines the video playback status in real time to trigger subsequent functional modules. The question-and-search trigger module activates the search function by capturing hardware events or parsing voice wake-up commands. This can be implemented using a remote control key event listener or a voice command recognition engine. It simultaneously captures commands and collects video data when the user triggers a search. The predetermined frame of video information can be a 5-second video clip before and after the trigger moment, captured using video buffer capture technology to preserve the contextual content required for the search. A multimodal large model refers to a deep learning model with image understanding and natural language processing capabilities. This can be implemented using a Transformer architecture that integrates a visual encoder and a text decoder to simultaneously interpret user command semantics and video elements. A split-screen interactive interface divides the screen into a video playback area and a search results area. This can be achieved through the dynamic layout adjustment function of the graphical interface framework to display search feedback without interrupting video playback.

[0081] Specifically, when the playback status template detects video playback behavior, the question search function trigger module activates the voice or text command receiving function. After the user triggers the question search function through the remote control button or voice wake-up word, the question search trigger module immediately captures the video screen of a predetermined number of frames before and after the current moment, and obtains the search instruction through the microphone or input box. The intention recognition module converts the voice instruction into text and extracts the search intention, such as recognizing that the user needs to query the identity information of the person in the picture. The matching module associates the text intention with the video screen elements, such as comparing the facial features of the character with the database records in the multimodal large model. The automatic search module generates structured data based on the matching results, such as outputting the character's name, list of works and introduction. The split-screen display module zooms the original video screen to one side of the screen and displays the search results in the form of a scrolling list on the other side. For example, while maintaining video playback on the left, the actor's resume information is displayed on the right.

[0082] Compared to existing technologies, traditional video terminals only support single-mode searches based on static images, and displaying search results requires interrupting the playback interface. This device, through a multimodal large model, achieves the combined parsing of dynamic video content and natural language commands, solving the problem of real-time image understanding and interactive response. The split-screen display mechanism overcomes the operational limitations of full-screen switching, allowing information acquisition and content viewing to proceed in parallel.

[0083] Through the above technical solutions, this application realizes the dynamic analysis of real-time screen elements and multi-round dialogue interaction during video playback, avoiding the viewing interruption problem caused by screen switching in traditional solutions. Through the split-screen layout design, users can obtain extended information without blocking the main screen, improving the efficiency of information retrieval. The application of multimodal large models means that video content understanding is no longer limited to the preset tag library, and can handle open semantic query requirements.

[0084] Based on the above embodiment, the present invention also provides an intelligent terminal, whose principle block diagram can be shown as follows: Figure 4 The intelligent terminal includes a processor, a memory, a network interface, a display screen, and a database connected via a system bus.

[0085] An intelligent terminal of the present application further includes one or more programs, wherein one or more programs are stored in a memory and configured to be executed by one or more processors, including a method for executing a video screen search split-screen interaction method based on a multimodal large model.

[0086] Among them, the memory refers to a hardware device used to store program code and data, which can be implemented as a solid-state drive or flash memory chip. Its function is to provide storage space for the processor to execute instructions. The processor refers to an arithmetic unit used to execute program instructions. It can be implemented as a central processing unit or a graphics processing unit. Its function is to control the process of video screen search and split-screen interaction by running program code. The method steps included in the program include detecting the video playback status, triggering the question search function, intercepting video frames, identifying user intent, matching screen elements, generating search results and split-screen display. Its function is to integrate the multimodal large model and split-screen interaction technology into the terminal device to achieve synchronous processing of video playback and information search.

[0087] Specifically, when the program is executed by the processor, it first continuously detects whether the terminal is in the video playback state. If the playback state is detected, the search function entrance of the voice or text command is activated. When the user triggers the search function through the remote control button or voice wake-up, the program automatically captures several frames of video before and after the trigger moment (for example, the frame from 5 seconds before the trigger to 2 seconds after the trigger), and sends the captured content together with the user command to the cloud service. The multimodal large model recognizes elements of the video frames, such as extracting character features, object categories or scene information in the picture, and analyzes the semantic intent of the user command. By comparing the matching degree of the identified elements with the user's intention, a search result containing the movie name, character introduction or scene explanation is generated, and the results are finally presented in a split-screen format on one side of the video playback interface.

[0088] In some embodiments, the video capture range can be dynamically adjusted based on the network transmission rate, for example, only capturing key frames when the network bandwidth is low. The layout ratio of the split-screen display area can be set to an adjustable mode, for example, allowing the user to adjust the size of the search results display area and the video playback area by sliding a gesture.

[0089] Compared to existing technologies, existing smart terminals typically use a fixed image library matching approach, supporting only single information retrieval based on static images and unable to perform real-time analysis of dynamic video content. This solution uses a multimodal large model to achieve semantic understanding of continuous video frames. Combined with split-screen interactive technology, it resolves the interface conflict between video playback and information retrieval, allowing users to obtain in-depth content interpretation without pausing playback.

[0090] Through the above technical solution, this application achieves dynamic parsing and semantic matching of real-time screen elements during video playback, allowing users to trigger in-depth searches of the currently playing content through natural language commands. The split-screen display mechanism ensures the continuity of video viewing while providing immediately relevant extended information, effectively addressing the technical limitations of traditional video playback devices, which cannot simultaneously handle content understanding and interactive querying.

[0091] The present application further proposes a computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform a video picture search and split-screen interaction method based on a multimodal large model.

[0092] The term "computer-readable storage medium" refers to a physical medium used to store executable program code. Specifically, this can be implemented using a storage device such as a solid-state drive, flash memory chip, or optical disk. Its function is to carry program instructions for implementing the video search and split-screen interaction method. Processor execution instructions refer to the reading and operation of program code from the storage medium by a central processing unit. Specifically, this can be implemented using a multi-core processor or distributed computing architecture. Its function is to drive electronic devices to execute operational processes such as video status detection, intent recognition, and split-screen display.

[0093] Specifically, when the instructions in the storage medium are executed by the processor, the electronic device first detects the current video playback status, and activates the voice or text search function after confirming that it is in the playback state. When a user question instruction is received, several frames of video before and after the trigger moment are automatically captured (for example, 5 frames before and 15 frames after the trigger), the voice instruction is converted into text and the search intent is identified. The video frame is identified by the multimodal large model, and the identified objects, people or scene features are matched with the user's intent, and finally the search results are displayed in the split-screen interface. The video playback area and the search result display area are presented simultaneously in the form of a left and right split screen, and the split-screen ratio can be dynamically adjusted according to the device screen size.

[0094] Compared to existing technologies, existing storage media only support fixed-scene image search and are unable to handle dynamic video content understanding and real-time interaction requirements. This solution, by integrating multimodal processing instructions into the storage medium, enables electronic devices to analyze dynamic video frames. It also enables the parallel display of video playback and search results through split-screen interactive instructions.

[0095] Through the above technical solution, this application solves the technical problem that existing storage media cannot support deep understanding of video images and real-time interaction, allowing users to quickly obtain detailed information about people and scenes in the picture through natural language commands without interrupting video playback, effectively improving the convenience of information acquisition and interactive experience during video viewing.

[0096] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A video image search and split-screen interaction method based on a multimodal large model, characterized in that: include: Detect and determine whether the smart terminal is currently playing a video; When it is detected that the smart terminal is currently playing a video, the control allows the question search function to be started for the video; When it is detected that the question search function for the video screen is triggered and started, the control obtains the question search instruction and simultaneously controls the interception of the video screen information of the predetermined frames before and after the question search function is triggered and started; Identifying the question search instruction and identifying the user's intention to search for video image information; Based on the identified user's intention to search for video image information, the multimodal large model is used to perform element recognition on the intercepted predetermined frame of video image information, and elements that match the user's intention to search for video image information are found from the recognized elements; Based on the elements that match the user's intention and the identified user's intention to search for video image information, the multimodal large model is used to automatically search for search results that match the user's intention to search for video image information; The search results are displayed through a preset split-screen interactive interface.

2. The video image search and split-screen interaction method based on a multimodal large model according to claim 1 is characterized in that: The step of detecting and determining whether the smart terminal is currently in the state of playing a video screen includes: pre-adding a setting to the video playback screen of the smart terminal: a function of searching for information in the video screen through voice commands or text commands, and setting the searched information to be displayed through a split-screen interactive function.

3. The video image search and split-screen interaction method based on a multimodal large model according to claim 1 is characterized in that: When the question search function for the video screen is detected to be triggered and started, the steps of obtaining the question search instruction and simultaneously controlling the interception of the video screen information of predetermined frames before and after the question search function is triggered and started include: When the smart terminal is in the video playback state, if it detects the remote control voice button press event or the voice wake-up command, it will control the question search function for the video screen to start; When it is detected that the question search function for the video screen is triggered and started, the control obtains the question search instruction, and at the same time controls the interception of the predetermined frame video screen information before and after the question search function is triggered and started, and saves it in the current question search dialogue state.

4. The video image search and split-screen interaction method based on a multimodal large model according to claim 1 is characterized in that: The step of identifying the question search instruction and identifying the user's intention to search for video image information includes: The obtained question search instruction is subjected to voice recognition through the cloud to identify the text information and the user's intention to search for video image information.

5. The video image search and split-screen interaction method based on a multimodal large model according to claim 1 is characterized in that: The step of performing element recognition on the intercepted predetermined frame of video image information using the multimodal large model based on the identified user's intention to search for video image information, and finding elements matching the user's intention to search for video image information from the identified elements includes: Obtain the video information of the predetermined frames before and after the question search function is triggered in the current conversation state, recognize the text information and the user's intention to search for video information, and send it to the cloud video image understanding service; The identified text information and the identified intention of the user to search for video screen information are judged to determine whether the search intention is for the captured predetermined frame video screen information; the identified text information and the identified intention of the user to search for video screen information are input into the multimodal large model, and the elements of the captured predetermined frame video screen information are identified through the multimodal large model, and elements that match the user's intention to search for video screen information are found from the identified elements.

6. The video image search and split-screen interaction method based on a multimodal large model according to claim 1 is characterized in that: The step of automatically searching for search results that match the user's intention to search for video image information based on finding elements that match the user's intention and identifying the user's intention to search for video image information through a multimodal large model includes: Obtaining the found elements that match the user's intent and the identified user's intent to search for video image information; Based on the elements that match the user's intention and the identified user's intention to search for video image information, the multimodal large model is used to automatically search for search results that match the user's intention to search for video image information; The multimodal large model is controlled to send the search results in text form.

7. The video image search and split-screen interaction method based on a multimodal large model according to claim 1 is characterized in that: The step of displaying the search results through a preset split-screen interactive interface includes: The intelligent terminal receives the search results sent in text form, controls the zooming out of the currently playing video screen, and controls the display of the pre-set split-screen search results streaming result display area; The split-screen search result streaming result display area on one side of the mobile terminal is controlled to display the search results, while the video area on the other side continues to play the video screen, thereby achieving a multi-modal large model image search split-screen effect.

8. A video image search and split-screen interaction device based on a multimodal large model, characterized in that: The device comprises: The playback status detection module is used to detect whether the smart terminal is currently playing a video; The question search function trigger module is used to control the start of the question search function for the video screen when it is detected that the smart terminal is currently in the video screen playback state; The question search trigger module is used to control the acquisition of the question search instruction when detecting that the question search function for the video screen is triggered and to control the interception of the video screen information of the predetermined frames before and after the question search function is triggered; An intention recognition module is used to recognize the question search instruction and identify the user's intention to search for video image information; A matching module is used to identify elements of the captured predetermined frames of video information based on the identified user's intention to search for video information using a multimodal large model, and to find elements that match the user's intention to search for video information from the identified elements; An automatic search module is used to automatically search for search results that match the user's intention to search for video image information based on finding elements that match the user's intention and identifying the user's intention to search for video image information through a multimodal large model; The split-screen display module is used to display the search results through a preset split-screen interactive interface.

9. An intelligent terminal, characterized in that: The device comprises a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, and the one or more programs include being used to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Multi-window interaction method and device based on intelligent screen service agent assistant, terminal and medium

    CN121326154A