Video interaction method, device and product based on large model

By using a large-model-based video interaction method, which combines spatial directional actions and information input, the target object is identified and data is processed. This solves the problem of the limitation of AI systems in understanding user intent in video calls and screen sharing scenarios, and achieves efficient and accurate human-computer interaction.

CN120897083AActive Publication Date: 2025-11-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511038178.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-04
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

Existing AI systems have limitations in understanding user intent in video calls and screen-sharing scenarios, resulting in low efficiency and insufficient accuracy in human-computer interaction.

Method used

By using a large-model-based video interaction method, the target object of the spatially directional action associated with the video frame is determined, and data processing instructions are generated based on the input information. The large model is used for data processing, combining a multimodal data interaction method of spatially directional action and information input.

Benefits of technology

It reduces the communication cost of human-computer interaction, improves the efficiency of understanding and processing user intentions, and provides an intuitive way of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120897083A_ABST
    Figure CN120897083A_ABST
Patent Text Reader

Abstract

The invention provides a video interaction method and device based on a large model, electronic equipment, a storage medium and a computer program product, relates to the technical field of artificial intelligence, in particular to the technical fields of large models, natural language understanding, video understanding and the like, and can be applied to video call and screen sharing scenes. The specific implementation scheme is as follows: in a video interaction process with a large model, determining a target object targeted by a spatial directivity action associated with a video picture in the video interaction process; determining a data processing instruction for the target object according to the input information associated with the spatial directivity action; and performing data processing on the target object according to the data processing instruction by adopting the large model to obtain a data processing result. According to the method and the device, the user is allowed to express the intention in a visual mode of combining spatial directional actions and information input, such as'finger 'and'speaking', so that the communication cost in the man-machine interaction process is reduced, and the understanding efficiency and the processing accuracy of the user intention in the man-machine interaction process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical fields of large models, natural language understanding, video understanding, and the like, and more particularly to a video interaction method and device based on a large model, an electronic device, a storage medium, and a computer program product, which can be applied to video calls and shared screen scenarios. BACKGROUND

[0002] Existing AI (Artificial Intelligence) systems can achieve a certain degree of human-computer interaction in video call, screen sharing, and the like. However, the AI system has limitations in understanding the user's intention in the human-computer interaction process. SUMMARY

[0003] The present disclosure provides a video interaction method and device based on a large model, an electronic device, a storage medium, and a computer program product.

[0004] According to a first aspect, a video interaction method based on a large model is provided, including: determining a target object to which a spatially directional action associated with a video picture in a video interaction process is directed in the video interaction process with a large model; determining a data processing instruction for the target object according to input information associated with the spatially directional action; and performing data processing on the target object according to the data processing instruction by using the large model to obtain a data processing result.

[0005] According to a second aspect, a video interaction device based on a large model is provided, including: an object determination unit configured to determine a target object to which a spatially directional action associated with a video picture in a video interaction process is directed in the video interaction process with a large model; an instruction determination unit configured to determine a data processing instruction for the target object according to input information associated with the spatially directional action; and a data processing unit configured to perform data processing on the target object according to the data processing instruction by using the large model to obtain a data processing result.

[0006] According to a third aspect, an electronic device is provided, including: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in any implementation manner of the first aspect.

[0007] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions are used to enable a computer to perform the method described in any implementation manner of the first aspect.

[0008] According to a fifth aspect, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method as described in any implementation of the first aspect.

[0009] According to the technology of the present disclosure, a large model-based video interaction method and device are provided. In the video interaction process with the large model, the target object to which the spatially directional action associated with the video frame in the video interaction process is directed is determined. According to the input information associated with the spatially directional action, a data processing instruction for the target object is determined. The large model is used to perform data processing on the target object according to the data processing instruction to obtain a data processing result. Thus, a human-computer interaction mode combining spatially directional action, video frame and input information multi-modal data is provided, allowing the user to express the intention in an intuitive way combining spatially directional action and information input, such as "pointing" and "saying", reducing the communication cost in the human-computer interaction process and improving the understanding efficiency and processing accuracy of the user's intention in the human-computer interaction process.

[0010] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0011] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them: Figure 1 is an exemplary system architecture diagram to which an embodiment according to the present disclosure can be applied; Figure 2 is a flowchart of one embodiment of the large model-based video interaction method according to the present disclosure; Figure 3 is a schematic diagram of an application scenario of the large model-based video interaction method according to the present embodiment; Figure 4 is a schematic diagram of the data processing instruction combining text and graphics according to the present embodiment; Figures 5A-5G is a schematic diagram of the data processing process based on multiple spatially directional actions according to the present embodiment; Figure 6 is a flowchart of another embodiment of the large model-based video interaction method according to the present disclosure; Figure 7 is a flowchart of another embodiment of the large model-based video interaction method according to the present disclosure; Figure 8 is a structural diagram of one embodiment of the large model-based video interaction device according to the present disclosure; Figure 9is a structural schematic diagram of a computer system suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION

[0012] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, and should be considered as merely exemplary. Thus, those skilled in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope and spirit of the present disclosure. Also, in the following description, descriptions of well-known functions and structures are omitted for clarity and conciseness.

[0013] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solutions comply with relevant laws and regulations and do not violate public order and good customs.

[0014] Figure 1 An exemplary architecture 100 of a large model-based video interaction method and device to which the present disclosure can be applied is shown.

[0015] As shown in Figure 1 The system architecture 100 can include terminal devices 101, 102, 103, a network 104 and a server 105. The terminal devices 101, 102, 103 are communicatively connected to form a topological network, and the network 104 is used as a medium to provide communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0016] The terminal devices 101, 102, 103 can be hardware devices or software that support network connection for data interaction and data processing. When the terminal devices 101, 102, 103 are hardware, they can be various electronic devices that support network connection, information acquisition, interaction, display, processing, etc., including but not limited to smartphones, tablet computers, e-book readers, laptop computers and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they can be installed in the above-mentioned electronic devices. They can be implemented as multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made herein.

[0017] The server 105 can be a server that provides various services, such as a background processing server that acquires multi-modal data such as video frames displayed by the terminal devices 101, 102, 103, input information and spatially directional actions of a user with respect to the video frames, and performs human-computer interaction based on the multi-modal data. As an example, the server 105 can be a cloud server.

[0018] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, software or software modules used to provide distributed services), or as a single software or software module. This is not specifically limited here.

[0019] It should also be noted that the large model-based video interaction method provided by the embodiments of the present disclosure is generally executed by a server, but the possibility of being executed by a terminal device or by a server and a terminal device in cooperation with each other cannot be excluded. Accordingly, the large model-based video interaction apparatus includes various parts (for example, various units) that can all be disposed in a server, all be disposed in a terminal device, or be disposed in a server and a terminal device, respectively.

[0020] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the system architecture is only illustrative. According to the needs of implementation, there can be any number of terminal devices, networks, and servers. When the electronic device on which the large model-based video interaction method runs does not need to perform data transmission with other electronic devices, the system architecture can only include the electronic device (for example, a terminal device or a server) on which the large model-based video interaction method runs.

[0021] For reference Figure 2 , Figure 2 A flowchart of a large model-based video interaction method provided by the embodiments of the present disclosure is shown in FIG. 2. In flowchart 200, the following steps are included: Step 201, in the process of video interaction with a large model, determining a target object to which a spatially directional action associated with a video frame in the process of video interaction is directed.

[0022] In this embodiment, the execution subject (for example, the server in Figure 1 ) of the large model-based video interaction method can obtain, from the terminal device of the user, a video frame in the process of video interaction between the user and the large model, a spatially directional action associated with the video frame, and determine, in the video frame, a target object to which the spatially directional action associated with the video frame is directed.

[0023] The video frame is a video frame displayed in real time by the terminal device of the user, which can be a video frame in a video call or a video frame in a shared screen. It can be obtained in real time by screen recording, based on a camera or a player directly obtaining a video stream, and the like. It should be noted that any acquisition method needs to obtain the authorization of the user before performing the real-time acquisition operation of the video frame.

[0024] The spatially directed action associated with the video picture can be an action made by the user on a display screen of the video picture, such as a spatially directed action based on a mouse pointer, a touch screen, or the like, such as a click, a touch, a circle selection, a line drawing, a drag, or the like, on a target object in the display screen, or an eye movement action representing a gaze point or a gaze track of the user in the video picture (in this case, an eye movement tracking technology such as an eye movement tracker is required to obtain the spatially directed action); or can be a spatially directed action expressed by the user in the video picture based on a gesture or the like, such as a pointing, a circle selection, a grabbing, or the like, of a target object in a real environment by the user, a real environment including the physical gesture being captured by a camera, and a video picture being obtained in real time.

[0025] The target object can be any object in the video picture, which can be a real entity object such as a flower or an article, or a non-entity object presented in the form of a word, a figure, a symbol, or the like, or a virtual object generated by a computer with the development of virtual reality, augmented reality, or the like.

[0026] The user can perform video interaction such as a video call or a screen sharing with the large model. The large model (artificial intelligence large model) refers to a class of artificial intelligence models with a large number of parameters constructed by an artificial neural network, such as a large language model, a visual large model, a multi-modal large model, and a basic scientific large model, or the like. Taking the large language model as an example, it is a large-scale language model constructed based on a deep learning technology, mainly used for processing natural language processing tasks. It learns the patterns and structures of language by training on large-scale data, and can generate natural language text or understand natural language input. The embodiment can specifically adopt a multi-modal large language model, which usually includes the following modules: An input module: receives multi-modal data input by the user, such as text, images, videos, input information, or the like, such as a text generation request of a question, an instruction, or a dialogue content.

[0027] A preprocessing module: pre-processes the input multi-modal data, such as text data pre-processing including word segmentation, stop word removal, text cleaning, or the like, to convert the text into a form that can be processed by the model.

[0028] An encoding module: encodes the pre-processed multi-modal data into a vector form so that the model can understand and process it. Common encoding methods include word embedding and an encoder in the Transformer architecture. Among them, word embedding includes Word2Vec (Words to Vector), GloVe (Global Vectors for Word Representation), or the like.

[0029] Model Module: The core part, usually based on deep learning architectures like Transformer, is responsible for processing the encoded data vectors and performing language understanding and generation. The model learns the complex patterns and semantic relationships of language through multiple layers of neural network structures.

[0030] Decoding Module: Decodes the output vectors from the model into natural language text, images, videos, input information, generating responses to user inputs or processing results. Decoding methods can include greedy decoding, beam search, etc.

[0031] Output Module: Outputs the decoded text in a user-readable form, such as text, images, videos, input information displayed on the screen.

[0032] As an example, first, according to the implementation type of the spatially directed action, determine the target object to which the spatially directed action is directed, for example, in response to the spatially directed action being an action made by the user on the display screen of the video picture, determine the location to which the spatially directed action is directed based on the display screen; in response to the spatially directed action being a spatially directed action expressed by the user based on gestures and other body movements in the video picture, determine the target object to which the spatially directed action is directed based on the recognition of the spatially directed action in the video picture.

[0033] As another example, input the implementation type of the spatially directed action and the video picture into a pre-trained target object recognition model, and determine the target object to which the spatially directed action is directed through the target object recognition model. The target object recognition model is used to represent the object relationship between the implementation type of the spatially directed action, the video picture, and the target object to which the spatially directed action of the video picture is directed. It can be obtained by using machine learning algorithms such as supervised learning and training neural networks such as recurrent neural networks and convolutional neural networks. The target object recognition model can be implemented through a large model.

[0034] In the process of identifying the target object, at least one of the following data processing methods can be used: Object recognition: Use a pre-trained visual model to recognize visual objects covered or adjacent to the mouse pointer or touch point, such as user interface elements (buttons, text boxes, scroll bars), specific text content (titles, paragraphs, keywords), specific areas in images, specific objects or characters in videos, etc.

[0035] Semantic segmentation or instance segmentation: Based on the context of the user dialogue, perform fine semantic segmentation or instance segmentation on the video picture to determine the semantic category (e.g., "chart area", "text area") or specific instance object (e.g., "third quarter column chart in sales trend chart") of the area where the mouse pointer or touch point is located.

[0036] Context understanding: combine the surrounding visual elements and overall picture layout of the spatially directional action to determine the target object pointed by the spatially directional action. For example, if the spatially directional action points to an interactive button, the target is identified as "button"; if the spatially directional action points to a specific text paragraph, the target object is identified as "text paragraph"; if the spatially directional action points to a person's face, it is identified as "person's face".

[0037] Step 202, according to the input information associated with the spatially directional action, determine the data processing instruction for the target object.

[0038] In this embodiment, the above execution subject can determine the data processing instruction for the target object according to the input information associated with the spatially directional action.

[0039] The input information associated with the spatially directional action refers to the input information issued by the user when making the spatially directional action. The input information is represented by at least one of text, voice, image, etc. Taking text as an example, the user can input text data to the above execution subject while making the spatially directional action. For example, the above execution subject displays a text input box in the terminal screen in response to detecting the user's spatially directional action, and the text information input by the user through the text input box is the input information associated with the spatially directional action. Taking voice as an example, the user issues voice while making the spatially directional action. In this case, the user can intuitively express his intention by combining "pointing" and "saying", "pointing" specifically represents the spatially directional action, and "saying" specifically represents the voice. For example, the user points to "flowers" in the video picture and says "what kind of flowers are these". Taking image as an example, the user can input image data to the above execution subject while making the spatially directional action. For example, the above execution subject displays an image upload box in the terminal screen in response to detecting the user's spatially directional action, and the image data uploaded by the user through the image upload box is the input information associated with the spatially directional action.

[0040] Taking the input information represented by voice and image as an example, the user issues the voice "process this into the style shown in the uploaded image" while making the spatially directional action, and uploads image data to the above execution subject through the image upload box.

[0041] As an example, the semantics represented by the input information can be recognized, for example, the recognized text of the voice is obtained by using voice-to-text technology, and the recognized text is directly taken as the data processing instruction for the target object.

[0042] As a further example, the recognized text can be deeply understood according to the video picture, the spatially directional action, to determine a real operation intention of the user for the target object; and a data processing instruction for the target object is generated according to the real operation intention. For example, the video picture, the spatially directional action and the recognized text are input into a large model, the real operation intention of the user for the target object is determined by the large model, and the data processing instruction is generated.

[0043] The data processing instruction can be various data processing instructions in office, learning, entertainment and other scenarios, including but not limited to: The data processing instruction aims to obtain the static attributes of the target object in the geometric, physical and semantic layers, including but not limited to size, material, color, texture, quality, transparency, reflectivity and category, etc. The data processing instruction is used to call prior knowledge or external knowledge base information related to the target object, covering historical background, cultural connotation, functional use, structural principle, ecological significance, related events and multi-language interpretation, etc. The data processing instruction allows real-time modification of the spatial pose, appearance features and dynamic behavior of the target object in the picture, such as displacement, rotation, scaling, deformation, color mapping, transparency adjustment, lighting condition change, texture replacement, and redefinition of interaction relationships such as collision, occlusion, adsorption, linkage with other objects.

[0044] In step 203, a large model is used to perform data processing on the target object according to the data processing instruction, to obtain a data processing result.

[0045] In this embodiment, the above execution subject can use a large model to perform data processing on the target object according to the data processing instruction, to obtain a data processing result.

[0046] In this embodiment, the data processing instruction is input into the large model, the large model deeply understands the data processing instruction based on its powerful natural language understanding and logical analysis capability, and performs data processing on the target object based on the understanding result, and displays the data processing result through the user's terminal device. The data processing result can be represented in the form of text, input information, image, video and other data, or in the form of a combination of multiple data.

[0047] Continuing to refer to the above data processing instruction for analyzing the health status of the leaf, the large model extracts features from the image of the target object (leaf), and determines the health status according to the feature extraction result, and provides targeted suggestions according to the health status.

[0048] As another example, the data processing instruction is "change the yellow leaves of this plant into green leaves", the large model extracts features of the target object (plant), and determines the yellow leaves on the plant based on the feature extraction result, and adjusts the yellow leaves on the plant in the video picture to green leaves.

[0049] With reference to the foregoing Figure 3 , Figure 3 is one of the application scenarios of the large model-based video interaction method according to the embodiment. The user 301 is processing a presentation through the terminal device 302, and wants to get help from the AI system set in the server 303 in the presentation processing process through the shared screen. During the presentation processing, the user finds that he does not understand the meaning of a technical term "classic computing bottleneck" in the presentation, and points to the technical term 304 in the video picture currently displayed by the terminal device through the mouse pointer, and issues an inquiry voice 305 "what is the meaning of this term". The AI system obtains the video picture, and determines the technical term (target object) to which the spatially directional action associated with the video picture is directed; according to the inquiry voice associated with the spatially directional action, the data processing instruction "what is the meaning of the term "classic computing bottleneck" is determined; the large model is used to perform data processing on the target object according to the data processing instruction, and the specific meaning of the technical term (data processing result) is obtained and displayed through the terminal device.

[0050] In the embodiment, a large model-based video interaction method is provided. In the video interaction process with the large model, the target object to which the spatially directional action associated with the video picture in the video interaction process is directed is determined; according to the input information associated with the spatially directional action, the data processing instruction for the target object is determined; the large model is used to perform data processing on the target object according to the data processing instruction, and the data processing result is obtained, thereby providing a human-computer interaction mode combining spatially directional action, video picture and input information multi-modal data, allowing the user to express the intention in an intuitive way combining spatially directional action and information input, such as "pointing" and "saying", reducing the communication cost in the human-computer interaction process, and improving the understanding efficiency and processing accuracy of the user's intention in the human-computer interaction process.

[0051] In some optional implementation manners of the embodiment, the execution subject can execute the step 202 in the following manner. In a first step, a semantic description text of the target object is generated in response to the description part of the target object in the recognized text of the input information being implicitly referred to.

[0052] Firstly, the semantics of the input information is recognized to obtain the recognized text. For example, based on the voice-to-text technology, the recognized text of the voice is obtained, and the recognized text follows the timestamp of the input information. According to the timestamp of the recognized text and the timestamp of the spatially directional action, the description part of the target object in the recognized text is determined, for example, the part of the text in the recognized text that is the same as the timestamp of the spatially directional action is taken as the description part of the target object. In order to improve the determination accuracy of the description part of the target object in the recognized text, the semantics understanding of the recognized text and the determination of the description part of the target object in the recognized text can also be combined.

[0053] Then, based on the semantic understanding of the description part of the target object, it is determined whether it is an implicit reference description. The implicit reference description refers to the language expression that does not directly say the specific person, thing, object and the like target object, but indirectly implies the object to be pointed out through the context, the information known by both parties or other associated data, such as “this” and “which” belong to the implicit reference description.

[0054] Finally, in response to the description part of the target object in the recognized text being an implicit reference description, a semantic description text capable of representing the target object is generated. For example, the semantic description text represents what the target object is, such as the category, name and the like of the target object.

[0055] In the second step, the semantic description text and the recognized text are combined to determine the data processing instruction.

[0056] In the present implementation, the semantic description text is used to replace the description part of the target object in the recognized text to obtain the fused text in which the implicit reference description is adjusted to the explicit reference description, and the fused text is taken as the data processing instruction. For example, the recognized text “this is what kind” in which the implicit reference description of the target object is “this” and the semantic description text of the target object is “a pot of flowers”, and the fused text is “this pot of flowers is what kind”.

[0057] In some implementations, in order to improve the semantic integrity and accuracy of the data processing instruction, a semantic description text with more information can also be generated. The semantic description text includes basic attribute information such as morphological features, appearance attributes, spatial position relationship such as position and interaction relationship with other objects, state and action description information, function and attribute association information (for a target with a specific function, the function description will be involved, such as “the device is a portable charger for powering electronic devices”; it can also contain attribute association with other objects, such as “the key matches the door lock on the right side of the picture” and scene and context information (such as “the book in the picture is placed on the bookshelf in the library, surrounded by similar books”).

[0058] In this implementation, the large model can deeply understand the recognized text, and determine the target information required for performing the data processing task represented by the recognized text in combination with the deep understanding result, and then perform video understanding or image understanding on the target object to determine the target information of the target object, so as to generate the semantic description text.

[0059] In this implementation, when the description part of the target object in the recognized text is implicit reference description, the data processing instruction is determined in combination with the semantic description text of the target object and the recognized text, thereby improving the accuracy and integrity of the data processing instruction.

[0060] In some optional implementations of the embodiment, the execution subject can execute the second step in the following manner: first, determine the fused text in combination with the semantic description text and the recognized text; and then, combine the visual data of the target object into the description part of the target object in the fused text to determine the data processing instruction.

[0061] The visual data can be image data or video data of the target object. Taking the image data as an example, the image part corresponding to the target object in the video frame can be cropped to obtain the visual data in the form of the target object image. Taking the video data as an example, the target video frame including the target object in the plurality of video frames displayed on the screen of the terminal device can be screened out as the visual data in the form of the target object video. In order to further improve the pertinence of the visual data in the form of the target object video, the target object in each target video frame including the target object can also be cropped according to a fixed size to obtain a cropped frame, and the plurality of cropped frames can be combined according to the time sequence relationship between the target video frames to obtain the visual data in the form of the video including only the target object.

[0062] For example, for a static target object, image-form visual data can be used; for a dynamic target object, video-form visual data can be used.

[0063] In this implementation, first, the semantic description text is used to replace the description part of the target object in the recognized text to obtain the fused text in which the implicit reference description is adjusted to the explicit reference description; and then, the visual data of the target object is combined into the description part of the target object in the fused text to generate the data processing instruction in the form of the combination of images and texts or the combination of videos and texts. Continue to refer to Figure 4 , which shows a schematic diagram of the data processing instruction in the form of the combination of images and texts. The data processing instruction 400 includes an image part 401, a pot of peony, and a text part 402, “What kind of flower is this?”.

[0064] In the implementation, after the implicit reference description in the recognized text is made explicit, the visual data of the target object is further combined to obtain the data processing instruction of the image-text combination or video-text combination, the data richness and semantic expression of the data processing instruction are further improved, and the intention understanding accuracy and data processing efficiency of the large model based on the data processing instruction are further improved.

[0065] In some optional implementations of the embodiment, the execution subject can execute the step 202 by combining the visual data of the target object to the description part of the target object in the recognized text in response to the description part of the target object in the recognized text of the input information being explicit reference description, and determining the data processing instruction.

[0066] The explicit reference description refers to a description manner in which a specific object is directly and clearly pointed out by explicit words or phrases in language expression, so that the identity, range or characteristics of the referred object can be directly recognized and understood. For example, for the target object of peony, the description part is "this pot of peony".

[0067] In the implementation, after the visual data of the target object is combined to the description part of the target object in the fused text, the data processing instruction of the image-text combination or video-text combination is generated.

[0068] In the implementation, after the description part of the target object in the recognized text is determined to be explicit reference description, the visual data of the target object is further combined to obtain the data processing instruction of the image-text combination or video-text combination, the data richness and semantic expression of the data processing instruction are further improved, and the intention understanding accuracy and data processing efficiency of the large model based on the data processing instruction are further improved.

[0069] In some optional implementations of the embodiment, one input information of the user is associated with multiple spatially directional actions. For example, the user makes multiple spatially directional actions in the process of issuing one voice. In the action type dimension, the multiple spatially directional actions can be actions of the same type, such as two spatially directional actions both representing that the user clicks the target object in the video picture through the mouse pointer; or can be actions of different types, such as one spatially directional action representing that the user clicks the target object in the video picture through the mouse pointer, and the other spatially directional action representing that the user selects the target object in the video picture through eye movement. Generally, the multiple spatially directional actions are directed to different target objects, but the case that the multiple spatially directional actions are directed to the same target object is not excluded.

[0070] In the time sequence dimension, the multiple spatially directional actions can be actions simultaneously issued by the user, such as touching two target objects on a video picture with two fingers respectively; or actions issued at different times, such as a first spatially directional action representing that the user touches a first target object in the video picture by a finger, and a second spatially directional action representing that the user touches a second target object in the video picture by the finger.

[0071] In the association relationship dimension, the multiple spatially directional actions can be actions having an association relationship, such as interchanging the positions of target objects to which the two spatially directional actions are directed; or actions not having an association relationship, such as the two spatially directional actions being directed to two independent target objects that are irrelevant to each other, and the subsequent data processing processes corresponding to the two target objects being irrelevant to each other.

[0072] In the present implementation, the execution subject can execute the step 202 in the following manner: In a first step, for the multiple spatially directional actions, a semantic description text of a target object is generated in response to the description part of the target object to which the spatially directional action is directed in the recognized text of the input information being an implicit reference description.

[0073] First, the semantics of the input information is recognized to obtain a recognized text thereof. For example, based on a voice-to-text technology, a recognized text of the voice is obtained, and the recognized text follows the time stamp of the input information. Then, for each spatially directional action in the multiple spatially directional actions, the description part of the target object to which the spatially directional action is directed in the recognized text is determined according to the time stamp of the recognized text and the time stamp of the spatially directional action, such as the part of the text in the recognized text that is the same as the time stamp of the spatially directional action being taken as the description part of the target object. In order to improve the determination accuracy of the description part of the target object in the recognized text, the semantic understanding of the recognized text and the time stamp determination of the description part of the target object in the recognized text can also be combined. Based on the semantic understanding of the description part of the target object to which the spatially directional action is directed, it is determined whether the description part is an implicit reference description. In response to the description part of the target object to which the spatially directional action is directed in the recognized text being an implicit reference description, a semantic description text capable of representing the target object is generated.

[0074] In a second step, the recognized text of the input information and the multiple spatially directional actions are time sequence aligned to determine the time sequence correspondence relationship between the multiple spatially directional actions and the recognized text.

[0075] The recognized text and the plurality of spatially-oriented actions are time-aligned based on the timestamp of the recognized text and the timestamp of each of the plurality of spatially-oriented actions, and a description part in the recognized text corresponding to the same timestamp as each of the plurality of spatially-oriented actions is determined, to obtain a time correspondence between the plurality of spatially-oriented actions and the recognized text.

[0076] For example, the recognized text is "exchange this object with this object", in which the first "this object" corresponds to the timestamp of the first spatially-oriented action (pointing to the peony), and the second "this object" corresponds to the timestamp of the second spatially-oriented action (pointing to the Chinese rose). It can be understood that the description part corresponding to the time correspondence of one spatially-oriented action is the description part of the target object of the spatially-oriented action.

[0077] In the third step, a data processing instruction is determined based on the time correspondence, and the semantic description text and the recognized text are combined.

[0078] For example, for each spatially-oriented action with implicit reference description of the description part of the target object, the input information description text of the target object of the spatially-oriented action is replaced with the description part of the recognized text corresponding to the time correspondence of the spatially-oriented action, to obtain the data processing instruction. That is, for each implicit reference description in the at least one implicit reference description associated with the input information, the semantic description text of the target object corresponding to the implicit reference description is replaced with the implicit reference description in the recognized text.

[0079] In the present embodiment, a data processing instruction generation method under the condition of a plurality of spatially-oriented actions is provided. Based on the support of the plurality of spatially-oriented actions, the flexibility and convenience of the user in expressing the intention of "pointing" and "saying" by combining spatially-oriented actions and information input are further improved, and the data processing method for the target object is also enriched.

[0080] In some optional implementations of the present embodiment, the execution subject can execute the third step in the following manner: first, a fused text is determined based on the time correspondence and the combination of the semantic description text and the recognized text; and then, for the plurality of spatially-oriented actions, the visual data of the target object of the spatially-oriented action is combined into the description part of the target object in the fused text, to obtain the data processing instruction.

[0081] In the implementation mode, first, for each spatially directional action of the description part of the target object, the semantic description text of the target object of the spatially directional action is replaced with the description part of the recognized text having a time sequence corresponding relationship with the spatially directional action, to obtain a fused text after adjusting the implicit reference description to explicit reference description; then, for each target object, the visual data of the target object is combined into the description part of the target object in the fused text, to generate a combination of image and text, video and text, or a combination of image, video and text data processing instructions. In the same data processing instruction, the visual data of the target object of each spatially directional action can adopt the same data form, such as image; or different data forms, for example, the visual data of some target objects is image, and the visual data of some target objects is video.

[0082] In the implementation mode, after the implicit reference description in the recognized text is made explicit, the visual data of the target object of each spatially directional action is further combined to obtain a data processing instruction, which improves the data richness of the data processing instruction and enriches the data processing mode for the target object.

[0083] With reference to Figures 5A-5G , a data processing process diagram based on multiple spatially directional actions is shown.

[0084] The user's voice is "I want to put this picture on the window sill at this position, can you generate an effect picture for me to see?", which corresponds to the video screen shown in Figure 5A , the user's first spatially directional action is to click the position corresponding to the "picture" 501 through the mouse pointer; which corresponds to the video screen shown in Figure 5B , the user's second spatially directional action is to select the target position range 502 on the "window sill" through the mouse pointer, and finally the mouse pointer stops at the position shown in Figure 5C .

[0085] For the first spatially directional action, the above execution subject combines the position pointed by the first spatially directional action and the part of the speech recognition text "this picture" corresponding to the first spatially directional action to recognize the target object it points to, and the recognition effect is as shown in Figure 5D .

[0086] For the second spatially directional action, the above execution subject combines the action track of the second spatially directional action and the part of the speech recognition text "the position of the window sill" corresponding to the second spatially directional action to recognize the target position range, and the recognition effect is as shown in Figure 5E .

[0087] For each target object, the visual data of the target object is combined into the description part of the target object in the post-fusion text, and a picture-text combined data processing instruction is generated, as shown in Figure 5F .

[0088] The large model performs data processing on the target object based on the data processing instruction to obtain a data processing result, as shown in Figure 5G .

[0089] In some optional implementations of the embodiment, the execution subject can execute the step 203 by using a large model to perform data processing on the target object according to the data processing instruction, the context of the input information and the video picture, and obtain a data processing result.

[0090] As an example, the execution subject can use a multi-modal large model to deeply analyze the data processing instruction that fuses the visual information and the text information of the target object, and understand the complex intent contained therein. The large model can accurately distinguish the instruction types such as asking questions, requesting to perform specific operations, seeking explanations, emotional interaction, etc. Further, the execution subject can input the data processing instruction, the context of the input information and the video picture into the large model. The large model combines the context of the historical dialogue, the overall semantics of the current video picture and the precise target object pointed by the user, and performs deeper reasoning, and then performs data processing on the target object based on the deep understanding result.

[0091] In the implementation, the large model performs data processing on the target object based on the data processing instruction, the context of the input information and the video picture, which further improves the accuracy of the data processing result.

[0092] In some optional implementations of the embodiment, the execution subject can execute the step 201 by the following way: First, in the video interaction process, the position of the spatially directional action on the video picture is determined.

[0093] In the video interaction process between the user and the large model, the execution subject collects the input information of the user, the video picture of the terminal device of the user and the spatially directional action of the user on the video picture in real time, and performs time sequence alignment on the multi-modal data such as the input information, the video picture and the spatially directional action. It should be noted that the collection of the multi-modal data is performed under the authorization of the user.

[0094] For the spatially directed action made by the user through the mouse, first, the absolute coordinates of the mouse pointer on the screen are captured by the terminal device, and then the absolute coordinates are mapped to the pixel coordinates relative to the original resolution of the video picture through the horizontal and vertical scaling ratios according to the actual display area of the current video picture on the screen, so as to determine the position of the mouse pointer in the video picture.

[0095] For the spatially directed action made by the user through the touch screen, first, the absolute coordinates of the user touch on the screen are captured by the terminal device, and then the absolute coordinates are back calculated to the original coordinate system of the video picture according to the transformation matrix generated by the stretching, cropping or rotating of the video picture, combined with the scaling relationship between the original size of the video and the display area, so as to determine the accurate position of the touch point in the video picture.

[0096] The second step is to determine the target object at the position in the video picture.

[0097] As an example, the above execution subject can input the video picture and the position into a target object recognition model, and the target object recognition model outputs the target object at the position in the video picture.

[0098] In the present implementation, a target object recognition method combined with a spatially directed action is provided, and the recognition efficiency and accuracy of the target object are improved based on the explicit pointing of the action.

[0099] In some optional implementations of the present embodiment, the above execution subject can perform the above second step by determining the target object at the position in the video picture according to the description part at the same time as the spatially directed action in the recognized text of the input information.

[0100] As an example, the recognized text and the spatially directed action can be time-aligned to determine the description part at the same time as the spatially directed action in the recognized text; the semantic information of the target object is determined according to the semantic understanding of the description part; and the target object at the position in the video picture is determined according to the semantic information.

[0101] In the present implementation, based on the spatially directed action, the target object in the video picture is further determined based on the description part at the same time as the spatially directed action in the recognized text of the input information, which helps to further improve the recognition efficiency and accuracy of the target object.

[0102] In some optional implementations of the present embodiment, the above execution subject can perform the above first step by determining the target position range according to the position points of the action trajectory represented by the spatially directed action on multiple video pictures in the video interaction process.

[0103] In the implementation, the execution subject can determine whether the position of the same spatial directional action in the video screen moves, and in response to determining that the position moves, determine that the spatial directional action is represented by an action trajectory, and determine position points on the plurality of video screens to determine a target position range represented by the plurality of position points.

[0104] For example, for each video screen (video frame) touched by the spatial directional action of the user, a position point of the spatial directional action in the video screen is determined, the plurality of position points are connected according to the time sequence relationship between the video screens, so as to determine an action trajectory of the spatial directional action in the video, and a target position range indicated by the spatial directional action such as circle selection and line drawing is determined.

[0105] In the implementation, the execution subject can perform the second step by performing object recognition in the target position range in the video screen to determine the target object.

[0106] In response to recognizing one object in the target position range in the video screen, the object is determined as the target object; in response to recognizing a plurality of objects in the target position range in the video screen, a description part in the recognized text at the same time as the spatial directional action is subjected to semantic understanding, semantic information of the target object is determined, and the target object is determined from the plurality of objects according to the semantic information.

[0107] When the semantic information of the target object is relatively ambiguous, a determination request can be sent to the user for the plurality of recognized objects, so as to filter the target object based on the indication of the user.

[0108] In the implementation, an implementation of recognizing a target object in a target position range is provided, the user can indicate the target object through an action sequence, the action flexibility of the user is improved, and the experience in the human-computer interaction process is improved.

[0109] In some optional implementations of the embodiment, the execution subject can perform the determination process of the target position range in the following manner: First, in the video interaction process, a hot area to which the spatial directional action is directed is determined according to the position points of the action trajectory represented by the spatial directional action on the plurality of video screens and the attribute information of the action trajectory.

[0110] The attribute information of the action trajectory includes, but is not limited to, shape, speed, dwell time, and the like. The action trajectory data represented by the captured position points is input into a pre-trained directional intention judgment model. The model can understand the directional intention represented by different trajectory patterns. For example, the model can identify that the user is performing a circle selection, a line drawing, or the like. Further, the model analyzes the shape, speed, dwell time, and the like of the pointer motion trajectory, and combines the video screen to predict the hot area to which the user intends to point.

[0111] Then, according to the hot area, the target position range is determined.

[0112] As an example, the position range represented by the hot area can be taken as the target position range.

[0113] In the present implementation, the target position range is determined according to the position points of the action trajectory on the plurality of video screens and the attribute information of the action trajectory, which improves the accuracy of the target position range.

[0114] In some optional implementations of the present embodiment, the above-mentioned execution subject can also perform the above-mentioned first step by determining the position points of the instantaneous action represented by the spatially directional action on the video screen in the video interaction process.

[0115] The instantaneous action is, for example, a click based on a mouse, a touch screen on a screen, or the like.

[0116] In the present implementation, the above-mentioned execution subject can determine whether the position of the spatially directional action in the video screen moves, and in response to determining that no movement occurs, determine that the spatially directional action is an instantaneous action and determine the position points thereof on the video screen.

[0117] In the present implementation, the above-mentioned execution subject can perform the above-mentioned first step by performing object recognition at the position points in the video screen to determine the target object.

[0118] The position points are points on or adjacent to the target object, and based on the position points, the object corresponding thereto can be recognized to obtain the target object.

[0119] In the present implementation, an implementation of recognizing a target object based on position points is provided, which supports the user to indicate a target object through an instantaneous action, improves the action flexibility of the user, and improves the experience degree in the human-computer interaction process.

[0120] In some optional implementations of the embodiment, the video interaction process includes a video call process, a shared screen process, and a video dialogue process, i.e., the video picture can be a video picture displayed in a video call interface, can be a video picture displayed in a shared screen interface, or can be a video picture in a dialogue between a user and a large model through video.

[0121] Taking a video picture displayed in a video call interface as an example, a user can make a video call with another user or with a large model, and the execution subject collects a video picture displayed in a video call interface, spatially directional actions of the user, and input information in real time.

[0122] In the application of AI video call or video stream analysis, a user can directly point to an object, a person, or a scene in the real world as if communicating with a real person. The large model can understand the target object pointed to by the user and conduct accurate dialogue.

[0123] In an example, when a user is showing a plant, the user points to a leaf of the plant with a finger on a touch screen or a mouse pointer and asks, "Is this leaf sick?" The execution subject can identify that the user is pointing to the "leaf" as the target object, rather than the entire plant, and determine a data processing instruction for analyzing the health of the leaf.

[0124] Taking a video picture displayed in a shared screen interface as an example, a user can share a screen of a terminal with the execution subject, so that the execution subject collects a video picture displayed in a video call interface, spatially directional actions of the user, and input information in real time.

[0125] In some AI screen sharing collaboration tools, a user shares screen content of a terminal device (computer or mobile phone) with a large model. The large model can understand any UI (User Interface) element, text, picture, chart, etc. pointed to by the user on the screen and determine a data processing instruction for the target object on the shared screen.

[0126] In an example, a user stops a mouse pointer on a specific word on a page in a shared presentation and asks, "What is the definition of this word?" through voice, and the execution subject can generate a data processing instruction for determining the definition of the word.

[0127] Taking a video picture in a dialogue exchange as an example, a user can send a pre-recorded video to a large model. During the recording of the video, the user makes a spatially directional action and inputs information. For example, during the recording of the video, the user points at a target object in a real scene by a finger and utters corresponding speech, and a terminal device of the user collects a picture of the real scene and speech of the user to obtain the recorded video. For another example, the user records content displayed on a screen of the terminal device by a screen recording manner, which includes a spatially directional action made by a mouse pointer, and collects speech uttered by the user during the screen recording to obtain the recorded video.

[0128] In the video call, screen sharing, and video dialogue scenarios, the user is allowed to express an intention in an intuitive manner combining a spatially directional action and information input, for example, "pointing" and "speaking", which reduces the communication cost in human-computer interaction and improves the understanding efficiency of user intention and the data processing accuracy in human-computer interaction.

[0129] With reference to Figure 6 , another embodiment of a large model-based video interaction method according to the present disclosure is shown in a schematic flow 600. In the flow 600, the following steps are included: Step 601, in a video interaction process with a large model, a type of action of a spatially directional action associated with a video picture in the video interaction process is determined.

[0130] The type of action includes a sequence action and an instantaneous action.

[0131] Step 602, in response to the type of action being the sequence action, a target position range is determined according to position points of an action trajectory represented by the spatially directional action on a plurality of video pictures.

[0132] Step 603, an object recognition is performed in the target position range in the video picture to determine a target object.

[0133] Step 604, in response to the type of action being the instantaneous action, a position point of the instantaneous action represented by the spatially directional action on the video picture is determined.

[0134] Step 605, an object recognition is performed at the position point in the video picture to determine a target object.

[0135] Step 606, a description type of a description part of the target object in a recognized text of input information associated with the spatially directional action is determined.

[0136] The description type includes an implicit reference description and an explicit reference description.

[0137] Step 607, in response to the description part of the target object in the recognized text of the input information being an implicit reference description, generating a semantic description text of the target object.

[0138] Step 608, combining the semantic description text and the recognized text, determining a fused text.

[0139] Step 609, combining the visual data of the target object into the description part of the target object in the fused text, determining a data processing instruction.

[0140] Step 610, in response to the description part of the target object in the recognized text of the input information being an explicit reference description, combining the visual data of the target object into the description part of the target object in the recognized text, determining a data processing instruction.

[0141] Step 611, using a large model to perform data processing on the target object according to the data processing instruction, obtaining a data processing result.

[0142] Compared with the above-mentioned flow 200, the flow 600 of the large model-based video interaction method in the embodiment specifically illustrates the object recognition process based on the spatially directional action, the generation process of the data processing instruction, and allows the user to express the intention in an intuitive way by combining the spatially directional action and the information input, such as "pointing" and "speaking". This further reduces the communication cost in the human-computer interaction process and improves the understanding efficiency and processing accuracy of the user's intention in the human-computer interaction process.

[0143] Referring back to Figure 7 , a schematic flow 700 of yet another embodiment of a large model-based video interaction method according to the present disclosure is shown. In the flow 700, the following steps are included: Step 701, in the video interaction process with the large model, for a plurality of spatially directional actions associated with the received input information, the following operations are performed: Step 7011, determining the action type of the spatially directional action.

[0144] Step 7012, in response to the action type being a sequence action, determining a target position range according to the position points of the action trajectory represented by the spatially directional action on the plurality of video frames.

[0145] Step 7013, performing object recognition in the target position range in the video frame, determining a target object.

[0146] Step 7014, in response to the action type being a transient action, determining a position point of the transient action represented by the spatially directional action on the video frame.

[0147] Step 7015, performing object recognition at the position point in the video frame, determining a target object.

[0148] Step 7016, determining the description type of the description part of the target object in the recognized text of the input information.

[0149] Step 7017, in response to the description part of the target object in the recognized text being an implicit reference description, generating a semantic description text of the target object.

[0150] Step 702, performing time alignment on the recognized text and the plurality of spatially directional actions, and determining a time correspondence relationship between the plurality of spatially directional actions and the recognized text.

[0151] Step 703, according to the time correspondence relationship, combining the semantic description text and the recognized text to determine a fused text.

[0152] Step 704, for the plurality of spatially directional actions, combining the visual data of the target object to which the spatially directional action is directed into the description part of the target object in the fused text to obtain a data processing instruction.

[0153] Step 705, using a large model to perform data processing on the target object according to the data processing instruction to obtain a data processing result.

[0154] The flow 700 of the video interaction method based on a large model in this embodiment, compared with the flow 200 described above, specifically illustrates the object recognition process of the plurality of spatially directional actions, the generation process of the data processing instruction, and allows the user to express the intention in an intuitive way by combining the spatially directional action and the information input, such as "pointing" and "saying", further reduces the communication cost in the human-computer interaction process, and improves the understanding efficiency and processing accuracy of the user's intention in the human-computer interaction process.

[0155] Continuing to refer to Figure 8 , as an implementation of the method shown in the above figures, the present disclosure provides an embodiment of a video interaction device based on a large model, the system embodiment corresponds to the method embodiment shown in Figure 2 , and the system can be specifically applied to various electronic devices.

[0156] As shown in Figure 8 , the video interaction device 800 based on a large model includes: an object determination unit 801 configured to determine, in a video interaction process with a large model, a target object to which a spatially directional action associated with a video picture in the video interaction process is directed; an instruction determination unit 802 configured to determine a data processing instruction for the target object according to input information associated with the spatially directional action; and a data processing unit 803 configured to use a large model to perform data processing on the target object according to the data processing instruction to obtain a data processing result.

[0157] In some optional implementation manners of the present embodiment, the instruction determining unit 802 is further configured to: in response to the description part of the target object in the recognized text of the input information being the implicit reference description, generate a semantic description text of the target object; and determine the data processing instruction in combination with the semantic description text and the recognized text.

[0158] In some optional implementation manners of the present embodiment, the instruction determining unit 802 is further configured to: determine a fused text in combination with the semantic description text and the recognized text; and determine the data processing instruction by combining the visual data of the target object into the description part of the target object in the fused text.

[0159] In some optional implementation manners of the present embodiment, the instruction determining unit 802 is further configured to: in response to the description part of the target object in the recognized text of the input information being the explicit reference description, combine the visual data of the target object into the description part of the target object in the recognized text to determine the data processing instruction.

[0160] In some optional implementation manners of the present embodiment, the input information is associated with a plurality of spatially directional actions, and the instruction determining unit 802 is further configured to: for the plurality of spatially directional actions, in response to the description part of the target object to which the spatially directional action is directed in the recognized text of the input information being the implicit reference description, generate a semantic description text of the target object; perform time sequence alignment on the recognized text of the input information and the plurality of spatially directional actions to determine a time sequence correspondence relationship between the plurality of spatially directional actions and the recognized text; and determine the data processing instruction in combination with the semantic description text and the recognized text according to the time sequence correspondence relationship.

[0161] In some optional implementation manners of the present embodiment, the instruction determining unit 802 is further configured to: determine a fused text in combination with the semantic description text and the recognized text according to the time sequence correspondence relationship; and for the plurality of spatially directional actions, combine the visual data of the target object to which the spatially directional action is directed into the description part of the target object in the fused text to obtain the data processing instruction.

[0162] In some optional implementation manners of the present embodiment, the data processing unit 803 is further configured to: adopt a large model to perform data processing on the target object according to the data processing instruction, the context of the input information and the video picture, to obtain a data processing result.

[0163] In some optional implementation manners of the present embodiment, the object determining unit 801 is further configured to: in the video interaction process, determine the position of the spatially directional action on the video picture; and determine the target object at the position in the video picture.

[0164] In some optional implementations of the present embodiment, the object determining unit 801 is further configured to determine the target object of the position in the video picture according to a description part in the recognized text of the video that is at the same time as the spatially directional action.

[0165] In some optional implementations of the present embodiment, the object determining unit 801 is further configured to, in the video interaction process, determine a target position range according to position points of an action trajectory represented by the spatially directional action on a plurality of video pictures; and determine the target object by performing object recognition in the target position range in the video picture.

[0166] In some optional implementations of the present embodiment, the object determining unit 801 is further configured to, in the video interaction process, determine a hot area to which the spatially directional action is directed according to the position points of the action trajectory represented by the spatially directional action on the plurality of video pictures and attribute information of the action trajectory; and determine the target position range according to the hot area.

[0167] In some optional implementations of the present embodiment, the object determining unit 801 is further configured to, in the video interaction process, determine a position point of a momentary action represented by the spatially directional action on the video picture; and determine the target object by performing object recognition at the position point in the video picture.

[0168] In some optional implementations of the present embodiment, the video interaction process includes a video call process, a shared screen process, and a video conversation process.

[0169] In the present embodiment, a large model-based video interaction apparatus is provided. In the video interaction process with the large model, the target object to which the spatially directional action associated with the video picture in the video interaction process is directed is determined. According to the input information associated with the spatially directional action, a data processing instruction for the target object is determined. The large model is used to perform data processing on the target object according to the data processing instruction, and a data processing result is obtained. Thus, a human-computer interaction mode combining multi-modal data of spatially directional actions, video pictures, and input information is provided. The user can express the intention in an intuitive way by combining the spatially directional action and the information input, such as “pointing” and “speaking”. The communication cost in the human-computer interaction process is reduced, and the understanding efficiency and processing accuracy of the user intention in the human-computer interaction process are improved.

[0170] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, which comprises at least one processor and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to implement the large model-based video interaction method described in any of the above embodiments.

[0171] According to an embodiment of the present disclosure, the present disclosure also provides a readable storage medium storing computer instructions for enabling a computer to implement the large model-based video interaction method described in any of the above embodiments when executed.

[0172] The present disclosure provides a computer program product, which, when executed by a processor, can implement the large model-based video interaction method described in any of the above embodiments.

[0173] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0174] As shown in Figure 9 The device 900 includes a computing unit 901 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0175] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; the storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0176] The computing unit 901 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 performs various methods and processes described above, such as the large model based video interaction method. For example, in some embodiments, the large model based video interaction method can be implemented as a computer software program tangibly embodied in a machine readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded onto the RAM 903 and executed by the computing unit 901, one or more steps of the large model based video interaction method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the large model based video interaction method by any other appropriate means, such as by means of firmware.

[0177] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0178] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, implements the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0179] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0180] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0181] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0182] The computer system can include clients and servers. This relationship can be. The servers are generally remote from the users and can be accessed via the Internet using a communication network. The relationship can be a client-server relationship over a communications network, and as such, various methods known in the art for establishing and maintaining network connections, such as by way of internet accessible servers, can be employed. The term "communication network" is used expansively to include any interoperation arrangement between two or more electronic devices, and this includes a local area network, a wide area network, the Internet, etc.

[0183] According to the technical scheme of the embodiment of the present disclosure, a video interaction method and device based on a large model are provided. In the video interaction process with the large model, the target object to which the spatially directional action associated with the video frame in the video interaction process is directed is determined. According to the input information associated with the spatially directional action, a data processing instruction for the target object is determined. The large model is used to perform data processing on the target object according to the data processing instruction, and a data processing result is obtained. Thus, a human-computer interaction mode combining spatially directional action, video frame and input information multi-modal data is provided, allowing the user to express the intention in an intuitive way by combining spatially directional action and information input, such as "pointing" and "saying", reducing the communication cost in the human-computer interaction process, and improving the understanding efficiency and processing accuracy of the user's intention in the human-computer interaction process.

[0184] It should be understood that the various forms of flow shown above can be reordered, added or deleted steps. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical scheme provided by the present disclosure can be achieved, which is not limited herein.

[0185] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A video interaction method based on a large model, comprising: During video interaction with a large model, the target object of the spatially directional action associated with the video frame during the video interaction process is determined; Based on the input information associated with the spatial directional action, determine the data processing instructions for the target object; Using the large model, the target object is processed according to the data processing instructions to obtain the data processing result.

2. The method according to claim 1, wherein, The step of determining the data processing instructions for the target object based on the input information associated with the spatial directional action includes: In response to the fact that the description portion of the target object in the recognition text of the input information is an implicit referential description, a semantic description text of the target object is generated; The data processing instructions are determined by combining the semantic description text and the recognition text.

3. The method according to claim 2, wherein, The step of combining the semantic description text and the identified text to determine the data processing instruction includes: By combining the semantic description text and the identified text, the fused text is determined; The visual data of the target object is combined with the description portion of the target object in the fused text to determine the data processing instructions.

4. The method according to claim 1, wherein, The step of determining the data processing instructions for the target object based on the input information associated with the spatial directional action includes: In response to the fact that the description of the target object in the recognition text of the input information is an explicit referential description, the visual data of the target object is combined with the description of the target object in the recognition text to determine the data processing instruction.

5. The method according to any one of claims 1-4, wherein, The input information is associated with multiple spatial directional actions, and The step of determining the data processing instructions for the target object based on the input information associated with the spatial directional action includes: For multiple spatially directional actions, in response to the implicit referential description of the target object targeted by the spatially directional action in the recognition text of the input information, a semantic description text of the target object is generated; The recognized text of the input information and the multiple spatial directional actions are temporally aligned to determine the temporal correspondence between the multiple spatial directional actions and the recognized text; Based on the temporal correspondence, and in combination with the semantic description text and the recognition text, the data processing instruction is determined.

6. The method according to claim 5, wherein, The step of determining the data processing instruction based on the temporal correspondence, combined with the semantic description text and the identified text, includes: Based on the temporal correspondence, and combining the semantic description text and the identified text, the fused text is determined; For multiple spatial directional actions, the visual data of the target object targeted by the spatial directional action is combined with the description part of the target object in the fused text to obtain the data processing instruction.

7. The method according to claim 1, wherein, The process of using the large model and processing the target object according to the data processing instructions to obtain data processing results includes: Using the large model, the target object is processed according to the data processing instructions, the context of the input information, and the video frame to obtain the data processing result.

8. The method according to claim 1, wherein, During video interaction with a large model, determining the target object of the spatially directional action associated with the video frame during the video interaction includes: During the video interaction, the position of the spatial directional action on the video screen is determined; Determine the target object at the location in the video frame.

9. The method according to claim 8, wherein, Determining the target object at the location in the video frame includes: Based on the description portion of the input information that occurs at the same time as the spatial directional action, the target object at the location in the video frame is determined.

10. The method according to claim 8, wherein, Determining the position of the spatial directional action on the video screen during the video interaction includes: During the video interaction, the target location range is determined based on the position points of the motion trajectory represented by the spatial directional action on multiple video frames; and Determining the target object at the location in the video frame includes: Object recognition is performed within the target location range in the video frame to determine the target object.

11. The method according to claim 10, wherein, During the video interaction, determining the target location range based on the position points of the motion trajectory represented by the spatial directional action on multiple video frames includes: During the video interaction, the hot zone targeted by the spatial directional action is determined based on the position points of the action trajectory represented by the spatial directional action on multiple video frames and the attribute information of the action trajectory. The target location range is determined based on the hot zone.

12. The method according to claim 8, wherein, Determining the position of the spatial directional action on the video screen during the video interaction includes: During the video interaction, the position point of the instantaneous action represented by the spatial directional action on the video frame is determined; and Determining the target object at the location in the video frame includes: Object recognition is performed at the location point in the video frame to determine the target object.

13. The method according to any one of claims 1-12, wherein, The video interaction process includes video call process, screen sharing process, and video dialogue process.

14. A video interactive device based on a large model, comprising: The object determination unit is configured to determine the target object of the spatially directional action associated with the video frame during the video interaction with the large model. The instruction determination unit is configured to determine a data processing instruction for the target object based on the input information associated with the spatial directional action. The data processing unit is configured to use a large model to process the target object according to the data processing instructions and obtain the data processing result.

15. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-13.

16. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-13.

17. A computer program product comprising: A computer program that, when executed by a processor, implements the method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Method and device for optimizing microphone connection transmission protocol, equipment and medium

    CN110881135A

  • Method executed by intelligent equipment and related equipment

    CN115993783A

  • Information interaction method and device, electronic equipment and storage medium

    CN116680376A

  • Multi-modal interaction method and system based on holographic equipment

    CN116880701A

  • Contextual assistant using mouse pointing or touch cues

    CN119137658A