Display device and film watching interaction method
Through multimodal fusion and graph neural network technology, video, audio and text features of the video are extracted and character relationship features are generated, which solves the problem of insufficient understanding of single modal information in intelligent movie viewing programs, and achieves more accurate interactive feedback and better movie viewing experience.
Patent Information
- Application Number
- CN202510264548.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-08
AI Technical Summary
The existing intelligent movie viewing program uses single modal information to understand video content, making it difficult to fully obtain diverse information in complex video content, resulting in low accuracy in content understanding and affecting user viewing experience.
By extracting the video features, audio features and text features of the played clip data, the multimodal fusion model is used to perform feature fusion, and a character network diagram is generated, combining the graph neural network and the adversarial generation network to generate interactive feedback results.
It improves the accuracy of the understanding of the video content by the intelligent movie viewing program, enhances the accuracy and credibility of interactive feedback, and improves the user's viewing experience.
Smart Images

Figure CN120281946A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of display devices, and in particular, to a display device and a viewing interaction method. Background Art
[0002] A display device is a terminal device for displaying pictures. Based on the application programs installed in the display device, the display device can output different display pictures. Based on the playback function of the display device, users can watch media data such as movies through the display device.
[0003] The display device can embed an intelligent viewing program, which is a movie analysis program based on an intelligent understanding algorithm. During the playback of a movie by the display device, the intelligent viewing program can analyze the content of the movie in response to the interaction questions raised by the user about the movie, and provide the analysis results to the user to facilitate the user's understanding of the movie content during the viewing process.
[0004] However, the intelligent viewing program uses single-modal information to understand video content, making it difficult to comprehensively obtain diverse information in complex video content. This results in a low accuracy of the intelligent viewing program's understanding of the movie content, causing the output analysis results to not match the user's interaction questions and affecting the user's viewing experience. Summary of the Invention
[0005] This application provides a display device and a viewing interaction method to solve the problem of low accuracy in understanding the content of a movie using single-modal information.
[0006] In a first aspect, some embodiments of this application provide a display device, including a display, a memory, and a controller. The display is configured to display a media interface corresponding to media data; the memory is configured to store a multi-modal fusion model; and the controller is configured to:
[0007] In response to a first interaction instruction input by the user based on target media data, obtain the played segment data of the target media data according to the playback progress of the target media data;
[0008] Extract the video features, audio features, and text features of the played segment data;
[0009] Perform multi-modal fusion on the video features, audio features, and text features through the multi-modal fusion model to obtain multi-modal fusion features;
[0010] Generate a role network diagram based on the role information identified from the video features, audio features, and text features. The role network diagram is used to represent the role relationships between roles;
[0011] Extract the role relationship features corresponding to the role network diagram through a graph neural network;
[0012] Generate an interaction feedback result corresponding to the first interaction instruction according to the multi-modal fusion feature and the role relationship feature.
[0013] The above technical solutions have the following beneficial effects or advantages: By extracting the video features, audio features, and text features of the played clip data, and performing feature fusion to obtain multi-modal features, an interaction feedback result is generated from the role relationship features generated according to the graph neural network and the multi-modal fusion features. Through the multi-modal fusion features of this application, multiple modalities such as video, audio, and text are fused to improve the accuracy of voice interaction feedback.
[0014] In some embodiments, the memory further stores a role relationship model, and the step of the controller executing to extract the role relationship features corresponding to the role network diagram through the graph neural network is specifically configured as:
[0015] Input the role network diagram into the role relationship model to calculate the weight values of the role relationships in the role network diagram through the role relationship model;
[0016] Extract the role relationship features through the graph neural network according to the role network diagram and the weight values of the role relationships.
[0017] The above technical solutions have the following beneficial effects or advantages: By inputting the role network diagram into the role relationship model to calculate the weight values, the role relationship features are generated through the graph neural network, the role network diagram, and the weight values of the role relationships, improving the credibility of the interaction feedback.
[0018] In some embodiments, before the controller executes the step of inputting the role network diagram into the role relationship model, it is further configured as:
[0019] Obtain a role network diagram sample for training the role relationship model;
[0020] Input the role network diagram sample into the role relationship model to be trained in the training stage to calculate the weight values in the training stage through the role relationship model to be trained;
[0021] Input the role network diagram sample and the weight values in the training stage into the graph neural network to output the role relationship features in the training stage;
[0022] Calculate the role relationship training loss of the role relationship features in the training stage through a self-supervised learning task;
[0023] When the loss of the role relationship training is less than or equal to the loss threshold, the role relationship model is output based on the current model parameters of the role relationship model to be trained.
[0024] The above technical solutions have the following beneficial effects or advantages: When training the role relationship model, by performing a pre-training task of self-supervised learning on the role relationship model, the dependence on labeled data is reduced, the manual labeling cost is reduced, and the efficiency of model training is improved.
[0025] In some embodiments, the step in which the controller generates an interaction feedback result corresponding to the first interaction instruction according to the multi-modal fusion feature and the role relationship feature is specifically configured as:
[0026] A target generator that has completed adversarial generative network training generates a media resource understanding result according to the multi-modal fusion feature and the role relationship feature;
[0027] Parse the first target problem of the first interaction instruction;
[0028] Generate the interaction feedback result according to the first target problem and the media resource understanding result.
[0029] The above technical solutions have the following beneficial effects or advantages: The target generator obtained through adversarial generative network training can, through the way of network confrontation, generate a targeted interaction feedback result according to the parsed first target problem, improving the accuracy of generating the interaction feedback result.
[0030] In some embodiments, before the step in which the controller generates a media resource understanding result according to the multi-modal fusion feature and the role relationship feature through a target generator that has completed adversarial generative network training, it is further configured as:
[0031] Obtain multi-modal fusion feature samples and role relationship feature samples;
[0032] Train the generator in the training stage through the multi-modal fusion feature samples and the role relationship feature samples to generate a media resource understanding result in the training stage;
[0033] The discriminator outputs a discrimination result of the media resource understanding result in the training stage according to the true label, and the true label is used to represent the true content understanding result of the multi-modal fusion feature samples and the role relationship feature samples;
[0034] When the discrimination result is the first discrimination result, iterative training is performed on the generator in the training stage to update the generator parameters;
[0035] When the discrimination result is the second discrimination result, the target generator is output according to the current generator parameters.
[0036] The above technical solution has the following beneficial effects or advantages: During the process of the generative adversarial network, the discriminator discriminates the media understanding result generated by the generator, and then iteratively trains the generator according to the discrimination result output by the discriminator, so as to improve the accuracy of the target generator in generating the media understanding result.
[0037] In some embodiments, the controller executes the step of outputting the discrimination result of the media understanding result in the training stage by the discriminator according to the true label, and is specifically configured to:
[0038] Calculate the self-supervised loss between the media understanding result in the training stage and the true label through a self-supervised learning task;
[0039] When the generator loss is greater than the loss threshold, obtain the first discrimination result output by the discriminator;
[0040] When the generator loss is less than or equal to the loss threshold, obtain the second discrimination result output by the discriminator.
[0041] The above technical solution has the following beneficial effects or advantages: By combining the generative adversarial network with self-supervised learning, the dependence on labeled data is reduced, the manual labeling cost is reduced, and the training efficiency of the generator is improved.
[0042] In some embodiments, when the playback progress changes, the controller executes the step of performing multimodal fusion on the video feature, the audio feature, and the text feature through the multimodal fusion model to obtain the multimodal fusion feature, and is further configured to:
[0043] Update the played segment data according to the changed playback progress;
[0044] Update the video feature, the audio feature, and the text feature according to the updated played segment data;
[0045] Perform multimodal fusion on the updated video feature, the updated audio feature, and the updated text feature through the multimodal fusion model to obtain the updated multimodal fusion feature.
[0046] The above technical solution has the following beneficial effects or advantages: When the playback progress changes, the corresponding video feature, audio feature, and text feature will also change accordingly. Therefore, the display device can update the multimodal fusion feature when the playback progress changes, so as to update the media understanding result in real time according to the current playback progress.
[0047] In some embodiments, after the controller executes the step of generating an interaction feedback result corresponding to the first interaction instruction according to the multi-modal fusion feature and the role relationship feature, it is further configured to:
[0048] In response to a second interaction instruction generated by the user based on the interaction feedback result, parse the second target problem of the second interaction instruction;
[0049] Update the interaction feedback result according to the second target problem and the media asset understanding result.
[0050] The above technical solution has the following beneficial effects or advantages: The user can input a second interaction instruction according to the generated interaction feedback result to further perform problem interaction on the basis of the interaction feedback result, so as to improve the user's understanding of the target media asset data.
[0051] In some embodiments, after the controller executes the step of extracting the video feature, audio feature, and text feature of the played segment data, it is further configured to:
[0052] Obtain the video frames of the played segment data;
[0053] Identify the role expression feature of the role included in the video frame;
[0054] Perform multi-modal fusion on the video feature, the audio feature, the text feature, and the role expression feature through the multi-modal fusion model to obtain a multi-modal fusion feature.
[0055] The above technical solution has the following beneficial effects or advantages: By identifying the role expression feature of the role in the video frame and using the role expression feature as a modal feature to perform multi-modal feature fusion with the video feature, audio feature, and text feature, the display device can improve the accuracy of generating the media asset understanding result according to the role expression feature.
[0056] In a second aspect, some embodiments of the present application provide a viewing interaction method, which is applied to the display device described in the first aspect. The method includes:
[0057] In response to a first interaction instruction input by the user based on the target media asset data, obtain the played segment data of the target media asset data according to the playback progress of the target media asset data;
[0058] Extract the video feature, audio feature, and text feature of the played segment data;
[0059] Perform multi-modal fusion on the video feature, the audio feature, and the text feature through the multi-modal fusion model to obtain a multi-modal fusion feature;
[0060] Generate a character network diagram based on the character information identified from the video features, the audio features, and the text features, where the character network diagram is used to represent the character relationships between characters;
[0061] Extract the character relationship features corresponding to the character network diagram through a graph neural network;
[0062] Generate an interaction feedback result corresponding to the first interaction instruction according to the multimodal fusion features and the character relationship features.
[0063] As can be seen from the above technical solutions, the present application provides a display device and a viewing interaction method. The method responds to a first interaction instruction input by a user based on target media data, obtains the played segment data according to the playback progress, and extracts video features, audio features, and text features. The video features, audio features, and text features are fused through multimodal fusion technology features to obtain multimodal fusion features. Character relationship features are extracted from the character network diagram generated based on the video features, audio features, and text features, thereby generating an interaction feedback result. Through multimodal fusion technology, the present application intelligently and accurately performs content understanding on target media data through a variety of feature fusion methods, thereby providing an interaction feedback result that conforms to the content of the target media data to the user and improving the accuracy of content understanding. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0065] Figure 1 Schematic diagram of the operation scenario between the display device and the control device provided by some embodiments of the present application;
[0066] Figure 2 Schematic diagram of the hardware configuration of the display device provided by some embodiments of the present application;
[0067] Figure 3 Schematic diagram of the software configuration of the display device provided by some embodiments of the present application;
[0068] Figure 4 Flowchart of the display device executing the viewing interaction method provided by some embodiments of the present application;
[0069] Figure 5 Sequence diagram of the display device executing the viewing interaction method provided by some embodiments of the present application;
[0070] Figure 6Schematic diagram of the first embodiment for a display device provided in some embodiments of the present application to determine played segment data;
[0071] Figure 7 Schematic diagram of the second embodiment for a display device provided in some embodiments of the present application to determine played segment data;
[0072] Figure 8 Flowchart for a display device provided in some embodiments of the present application to obtain multimodal fusion features;
[0073] Figure 9 Flowchart for a display device provided in some embodiments of the present application to generate a role network diagram;
[0074] Figure 10 Flowchart for a display device provided in some embodiments of the present application to train a role relationship model by combining self-supervised learning;
[0075] Figure 11 Flowchart for a display device provided in some embodiments of the present application to train a generator;
[0076] Figure 12 Flowchart for a display device provided in some embodiments of the present application to update multimodal fusion features according to the updated playback progress. Detailed implementation manners
[0077] The embodiments will be described in detail below, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application detailed in the claims.
[0078] It should be noted that the brief description of the terms in the present application is only for facilitating the understanding of the following described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.
[0079] The terms "first", "second", "third", etc. in the specification, claims and the above drawings of the present application are used to distinguish similar or homogeneous objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.
[0080] The terms "include" and "have" and any variations thereof are intended to cover but not exclusively include. For example, a product or device including a series of components does not necessarily have to be limited to all the clearly listed components, but may include other components not clearly listed or inherent to these products or devices.
[0081] The term "module" refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware and / or software code that can perform functions related to that element.
[0082] In the embodiments of the present application, the display device 200 generally refers to a device with the ability to display images and process data. For example, the display device 200 includes, but is not limited to, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.
[0083] Figure 1 It is a schematic diagram of the operation scenario between the display device and the control device provided for some embodiments of the present application. As Figure 1 shown, the user can operate the display device 200 through touch operations, the mobile terminal 300, and the control device 100. Among them, the control device 100 is used to receive the operation instructions input by the user and convert the operation instructions into control instructions that the display device 200 can recognize and respond to. For example, the control device 100 can be a remote control, a stylus, a handle, etc.
[0084] The mobile terminal 300 can be used as a control device to perform human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device to establish a communication connection with the display device 200 for data interaction. In some embodiments, software applications can be installed on the mobile terminal 300 and the display device 200, and connection communication can be achieved through network communication protocols to achieve the purpose of one-to-one control operations and data communication. It is also possible to transmit the audio and video content displayed on the mobile terminal 300 to the display device 200 to achieve the synchronous display function.
[0085] In some embodiments, the mobile terminal 300 or other electronic devices can also simulate the functions of the control device 100 by running an application program for controlling the display device 200.
[0086] As Figure 1 also shown, the display device 200 also communicates with the server 400 for data communication through various communication methods. The display device 200 is allowed to establish a communication connection through a local area network (LAN), a wireless local area network (WLAN), and other networks.
[0087] The display device 200 can provide a broadcast reception TV function, and can also additionally provide a smart network TV function with computer support functions, including but not limited to, network TV, smart TV, Internet Protocol TV (IPTV), etc.
[0088] Figure 2 Provided for some embodiments of the present application Figure 1Hardware configuration block diagram of the display device 200 is shown.
[0089] In some embodiments, the display device 200 may include at least one of a tuner demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0090] In some embodiments, the detector 230 is used to collect signals from the external environment or for external interaction. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures. Or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.
[0091] In some embodiments, the display 260 includes a display function component for presenting a picture and a driving component for driving image display. The display 260 is used to receive an image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, components of a menu manipulation interface, and a user manipulation UI interface, etc.
[0092] In some embodiments, the communication device 220 is a component for communicating with an external device or a server 400 according to various communication protocol types. The display device 200 may be provided with multiple communication devices 220 according to different supported communication methods. For example, when the display device 200 supports wireless network communication, the display device 200 may be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.
[0093] The communication device 220 can communicate and connect the display device 200 with an external device or a server 400 in a wireless or wired connection manner. Among them, the wired connection can connect the display device 200 with an external device through components such as a data cable and an interface. The wireless connection can connect the display device 200 with an external device through a wireless signal or a wireless network. The display device 200 can directly establish a connection relationship with an external device or indirectly establish a connection relationship through a gateway, a router, a connection device, etc.
[0094] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and first to nth interfaces for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.
[0095] In some embodiments, the controller 250 and the tuner demodulator 210 may be located in different separate devices, that is, the tuner demodulator 210 may also be in an external device of the main device where the controller 250 is located, such as an external set-top box, etc.
[0096] In some embodiments, if a user enters a user command through the graphical user interface (GUI) displayed on the display 260, the user input interface receives the user input command through the graphical user interface (GUI).
[0097] In some embodiments, the audio output device 270 may be a built-in speaker of the display device 200, or may be an external audio output device connected to the display device 200. Among them, for the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device may be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.
[0098] In some embodiments, the user input interface 280 can be used to receive instructions from the user input.
[0099] To perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling the hardware resources and software resources in the display device 200. The operating system can control the display device to provide a user interface. For example, the operating system can directly control the display device to provide a user interface, or can provide a user interface by running an application program. The operating system also allows the user to interact with the display device 200.
[0100] It should be noted that the operating system may be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.
[0101] The operating system can be divided into different modules or levels according to the functions implemented. For example, as Figure 3As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (Applications) layer (referred to as "application layer"), the application framework layer (Application Framework) layer (referred to as "framework layer"), the system library layer and the kernel layer.
[0102] In some embodiments, the application layer is used to provide services and interfaces for applications so that the display device 200 can run applications and interact with users based on the applications. At least one application can be run in the application layer, and these applications can be window programs, system settings programs, clock programs, etc. that come with the operating system; they can also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.
[0103] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center that determines the actions that applications in the application layer take. Through the API interface, applications can access system resources and obtain system services during execution.
[0104] like Figure 3 As shown, Figure 3 A software configuration diagram of a display device provided for some embodiments of the present application. In some embodiments, the system of the display device 200 can be divided into three layers, namely, an application layer, a middleware layer, and a hardware layer from top to bottom.
[0105] The application layer mainly includes applications on the TV and the application framework. Among them, applications are mainly applications developed based on the browser, such as HTML5 APPs and native applications.
[0106] The Application Framework is a complete program model that has all the basic functions required by standard application software, such as file access, data exchange, etc., as well as the user interfaces of these functions (toolbars, status bars, menus, dialog boxes).
[0107] Native apps can support online or offline, message push or local resource access.
[0108] The middleware layer includes various middleware such as TV protocols, multimedia protocols, and system components. The middleware can use the basic services (functions) provided by the system software to connect various parts of the application system on the network or different applications, and can achieve the purpose of resource sharing and function sharing.
[0109] The hardware layer mainly includes the HAL interface, hardware, and drivers. Among them, the HAL interface is the unified interface for all TV chips to dock, and the specific logic is implemented by each chip. The drivers mainly include: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor drivers (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.
[0110] It should be noted that the above examples are only simple divisions of the functions of the operating system, and do not limit the specific form of the operating system of the display device 200 in the embodiments of the present application. According to factors such as the functions of the display device and the type of the operating system, the number of levels and the specific level types included in the operating system can be in other forms.
[0111] In some embodiments, the display device 200 can implement the functions of playing and interacting with media data by installing application programs. For example, the display device 200 can install a media playback application such as a video player, and the user can input a start instruction to start the video playback program based on the control icon corresponding to the video player in the user interface of the display device 200. The display device 200 can respond to the start instruction, run the video player, and display the video playback interface corresponding to the video player in the user interface. A variety of playable media data can be provided in the video playback interface, and the user can select the media data they want to watch to control the display device 200 to play.
[0112] To improve the user's viewing experience, the display device 200 can embed an intelligent viewing program. The intelligent viewing program is a film analysis program based on intelligent understanding algorithms, so as to understand the content of the media data during the process of the display device 200 playing the media data and provide the understanding results to the user. The intelligent viewing program can be embedded in the media playback application. After the display device 200 recognizes the installation of the media playback application, the intelligent viewing program can be embedded into the media playback application. In this way, when the display device 200 runs the media playback application, it can synchronously execute the content understanding function corresponding to the intelligent viewing program. The intelligent viewing program can also be directly embedded into the system of the display device 200 to directly understand the content of the media data played by the media playback application at the system layer.
[0113] In some embodiments, during the process of the display device 200 playing media data, the intelligent viewing program can analyze the content of the media data in response to an interaction question raised by the user based on the media data. For example, the user can input "What is the relationship between character A and character B?" After receiving the interaction question, the display device 200 starts to analyze the content of the media data and displays interaction feedback content on the user interface according to the analysis result, such as "Character A and character B are in a hostile relationship", so as to facilitate the user to understand the plot content of the media data according to the real-time viewing interaction.
[0114] The intelligent viewing program uses single-modal information to understand video content. For example, during the process of the display device 200 playing media data, the intelligent viewing program analyzes the content of the media data according to the lines of the characters in the media data, so as to determine the relationship between character A and character B based on the analysis of the lines of character A and character B. This way of using single-modal information to understand video content is difficult to comprehensively obtain diverse information in complex video content, resulting in a relatively low accuracy of the intelligent viewing program in understanding the content of the film, causing the output analysis result to not match the user's interaction question and affecting the user's viewing experience.
[0115] Based on the above scenario, some embodiments of the present application provide a display device 200, including a display 260, a memory, and a controller 250. Among them, the display 260 is configured to display a media interface corresponding to the media data, the memory is configured to store a multi-modal fusion model, and the controller 250 is configured to execute a viewing interaction method. Figure 4 It is a flowchart of the display device provided by some embodiments of the present application for executing viewing interaction. Figure 5 It is a timing diagram of the display device provided by some embodiments of the present application for executing viewing interaction. The display device 200 executes the viewing interaction method according to the timing relationship as Figure 5 shown. Refer to Figure 4 and Figure 5 , the method includes the following content:
[0116] S100: In response to a first interaction instruction input by the user based on the target media data, obtain the played segment data of the target media data according to the playback progress of the target media data.
[0117] The display device 200 can be used to play media data. For example, the display device 200 can play TV programs, play music, listen to the radio, etc. Based on different media data, the display device 200 can output different display screens. The display device 200 can respond to a start instruction for starting a playback application and display an application interface, and the application interface can include multiple media playback controls for indicating the media data corresponding to the playback controls.
[0118] The user can input a playback instruction based on the media asset playback control to control the display device 200 to play the media asset data selected by the user, that is, the target media asset data. During the process of the display device 200 playing the target media asset data, if the user has difficulty understanding information such as the plot and character relationships of the target media asset data, the user can input a first interaction instruction to the display device based on the media asset content of the target media asset data. The first interaction instruction can include the content that the user wants to ask, for example, "the relationship between character A and character B", "the future development trend of the plot", or the historical background when the target media asset data occurred, etc.
[0119] To facilitate providing the feedback result of the first interaction instruction to the user, the display device 200 can, in response to the first interaction instruction, obtain the played segment data of the target media asset data through the intelligent movie viewing program. The played segment data is the played part of the target media asset data. Therefore, the display device 200 can determine the played segment data according to the playback progress. For example, if the target media asset data is a 60-minute movie, as Figure 6 shown, and the current playback progress is 30 minutes at this time, that is Figure 6 the black part of the progress bar shown, it means that the played segment data is the target media asset data from 0 to 30 minutes. To avoid spoiling, the display device 200 does not use the content of the target media asset data within 30 - 60 minutes as the part for content understanding, so as to prevent the user from prematurely knowing the plot content of the unplayed media asset data segment after the playback progress.
[0120] In some embodiments, the playback progress can be the current playback progress or the historical playback progress. When the playback progress is the current playback progress, the played data segment can be continuously updated based on the current playback progress set by the user. For example, when the user is watching a movie and, due to not understanding some of the plot, sets the current playback progress of the movie from 40 minutes to 30 minutes and intends to watch the movie segment from 30 minutes to 40 minutes again, at this time, the current playback progress is traced back from 40 minutes to 30 minutes, that is, the played segment data is updated from the original 0 - 40 minutes to 0 - 30 minutes. Correspondingly, when the display device 200 responds to the first interaction instruction at this time, it needs to perform content understanding on the played segment data from 0 to 30 minutes.
[0121] The historical playback progress is the longest playback progress of the target media asset data. When the playback progress is the historical playback progress, the played segment data is not affected by the tracing back of the current playback progress. For example, as Figure 7As shown, the maximum playback progress of the video is 40 minutes. At this time, the user sets the current playback progress from 40 minutes to 30 minutes. The 40-minute time point is the historical playback progress of the target media data, that is, the played segment data is always 0 - 40 minutes. Correspondingly, since the historical playback progress remains unchanged during the playback backtracking process, when the display device 200 responds to the first interaction instruction, it continues to perform content understanding on the played segment data of 0 - 40 minutes.
[0122] It should be noted that the above embodiments are examples for explaining the played segment data when the historical playback progress remains unchanged. When the display device 200 plays continuously based on the maximum playback progress, the historical playback progress will be continuously updated based on the playback of the display device 200.
[0123] S200: Extract the video features, audio features, and text features of the played segment data.
[0124] After obtaining the played segment data, in order to facilitate content understanding of the played segment data, the display device 200 can combine the multi-modal features of the played segment, including video features, audio features, and text features. To obtain the video features, audio features, and text features respectively, as Figure 8 shown, the display device 200 can extract the video frames of the played segment data, the audio data of the played segment data, and the dialogue text data of the played segment data.
[0125] The display device 200 can input the video frames of the played segment data into a video encoder to extract the video features of the video frames through the video encoder. Among them, the video features can include character features and environmental features. The character features can include feature information such as the gender, appearance, and body type of the person, and the environmental features can include the time environment or location environment, etc., so that the display device can perform character relationship understanding or plot content understanding on the played segment data according to the extracted video features.
[0126] The display device 200 can input the audio data of the played segment data into an audio encoder to extract the audio features of the audio data. For example, the pitch, timbre, or speaking emotion of character A speaking to character B, etc., so as to understand the character relationship between character A and character B, as well as content such as emotional changes. The display device 200 can obtain the subtitle track of the target media data and extract the dialogue text data according to the playback progress. After obtaining the dialogue text data, the display device 200 can extract the text features of the dialogue text data through a text encoder.
[0127] It should be noted that the above video encoder, audio encoder, and text encoder can all use conventional encoding methods to encode the corresponding data, and this application does not make specific limitations on the encoding method of the played segment data.
[0128] S300: Perform multi-modal fusion on the video feature, the audio feature, and the text feature through the multi-modal fusion model to obtain a multi-modal fusion feature.
[0129] To improve the accuracy of content understanding for the target media asset data, the display device 200 can call the multi-modal fusion model stored in the memory, and input the extracted video feature, audio feature, and text feature into the multi-modal fusion model, so as to perform multi-modal fusion on the video feature, audio feature, and text feature through the multi-modal fusion model to obtain a multi-modal fusion feature.
[0130] In some embodiments, to further improve the accuracy of content understanding, the display device 200 can also obtain video frames of the played segment data and identify the character expression features of the characters included in the video frames. Among them, according to the time sequence of the video frames, the character expression features of the same character may change. For example, character A and character B are in an ally relationship, and character A has a friendly expression towards character B. When the alliance relationship breaks down, the relationship between character A and character B is converted into a hostile relationship. At this time, character A has a hostile expression towards character B.
[0131] Based on the above scenario, during the process of performing multi-modal fusion, the display device 200 can also input the character expression features together with the video feature, audio feature, and text feature into the multi-modal fusion model to obtain a multi-modal fusion feature, so that the display device 200 can, based on the timing relationship of the video frames, improve the fitting degree between the generated interaction feedback result and the plot change of the target media asset data as the playback progresses according to the change of the character's expression features.
[0132] S400: Generate a character network graph according to the character information identified from the video feature, the audio feature, and the text feature.
[0133] The character network graph is used to represent the character relationships between characters. As Figure 9 shown, the display device 200 can determine the character information according to the video frames of the played segment data, the audio data of the played segment data, and the line text data of the played segment data. The character information includes each character that appears and the number of characters that appear, and based on the character information, determine the relationship between characters according to the video feature, audio feature, and text feature. For example, character A and character B are in a teacher-student relationship, character C and character D are in a mother-son relationship, or character E and character F are in a superior-subordinate relationship, etc. For the same character, there can be multiple different character relationships with multiple characters. For example, character A and character B are in a mother-son relationship, at the same time in a father-son relationship with character C, and also in a teacher-student relationship with character D, etc. Correspondingly, character B and character C are in a husband-wife relationship.
[0134] In a neural network graph, multiple nodes can be included, as well as edges for connecting nodes to nodes. Among them, nodes are used to represent the roles that appear in the played segment data, and edges are used to represent the role relationships between roles (between nodes).
[0135] The display device 200 can generate a role network graph according to the extracted character relationships. The role network graph can include the role relationships between each role that appears in the played segment data and other roles. In addition to the role relationships, the role network graph can also include the emotional changes between each role and other roles, and specifically to the time points corresponding to the playback progress. For example, when playing to 15 minutes, the relationship between role A and role B is a stranger relationship. As the playback progress and the plot of the target media data change, when playing to 30 minutes, the stranger relationship between role A and role B is converted into a friendly relationship.
[0136] In some embodiments, for target media data with a large number of roles, during the process of generating a role network graph by the display device 200, the appearance times and appearance times of roles can also be recorded according to video features, and the preset number of main roles can be determined according to the appearance times and appearance times. Thus, the construction of the network graph of secondary roles with relatively few appearance times can be screened out. For example, roles with appearance times less than the preset times, such as roles that only appear once, or roles with appearance times less than the preset time, are screened out and not included in the generation of the role network graph, so as to reduce the time for constructing the network graph part of secondary roles and retain the main roles to construct the network graph of main roles.
[0137] In some embodiments, the display device 200 can also verify the role relationships in the role network graph to determine the correctness of the role relationships. For example, when role A is the mother of role B, then role B cannot be the mother of role A. The display device 200 can improve the correctness of the role relationships in the role network graph through the verification of the role relationships.
[0138] S500: Extract the role relationship features corresponding to the role network graph through a graph neural network.
[0139] After generating the role network graph, the display device 200 can extract the role relationship features corresponding to the role network graph. To this end, a role relationship model is also stored in the memory of the display device 200. After inputting the role network graph into the role relationship model, the role relationship model can calculate the weight value of the role relationship in the role network graph. The weight value is used to represent the relationship weight between roles. For example, the weight value of the cooperation relationship between role A and role B is 0.8. As the playback progress, the weight value can change with the plot of the target media data. For example, if role A and role B have a disagreement when discussing the cooperation method in the plot of the target media data, the display device 200 will update the role network graph based on the newly added video features, newly added audio features, and newly added text features extracted according to the playback progress, so that the role relationship model recalculates the weight value of the cooperation relationship between role A and role B for the updated role network graph. Based on the feature of "having a disagreement", the role relationship model can reduce the weight value of the cooperation relationship between role A and role B, for example, from 0.8 to 0.6, so as to realize updating the weight value of the role relationship in real time according to the playback content of the target media data.
[0140] After calculating the weight value, the display device 200 can input the role network graph and the weight value of the role relationship into a Graph Neural Network (GNN) at the same time. The graph neural network is a deep learning method specifically used to process graph-structured data. It can update the representation of nodes by defining the connection relationship between nodes in the role network graph and using the adjacent information of the nodes, so as to transmit and learn the information in the role network graph, so that the graph neural network extracts the role relationship features according to the role network graph and the weight value of the role relationship.
[0141] S600: Generate an interaction feedback result corresponding to the first interaction instruction according to the multi-modal fusion feature and the role relationship feature.
[0142] After obtaining the multi-modal fusion feature and the role relationship feature, the display device 200 can parse the first target question corresponding to the first interaction instruction and generate a corresponding interaction feedback result according to the multi-modal fusion feature and the role relationship feature. For example, the first target question is "What is the relationship between role A and role B?" The display device 200 can obtain the multi-modal fusion features of role A and role B, and combine the role relationship features output by the graph neural network, such as "cooperation index 0.8, hostility index 0", to generate interaction feedback content, such as "Role A and role B are in a cooperative relationship, jointly against role C, but there are differences in tactics."
[0143] As can be seen from the above technical solutions, the display device 200 provided by the present application can generate multi-modal fusion features and a role network graph by extracting video features, audio features, and text features of the played segment data, and extract role relationship features in the role network graph through a graph neural network, so as to generate an interactive feedback result according to the multi-modal fusion features and the role relationship features, thereby improving the content understanding ability of the display device 200 for target media data and improving the accuracy of the interactive feedback result.
[0144] In order for the display device 200 to facilitate the execution of the above-mentioned viewing interaction method, it is also necessary to perform a specific training process on the multi-modal fusion model and the role relationship model. In order to facilitate the display device 200 to identify the information of the training data, these training data are manually annotated data sets. Manual annotation requires manual processing of unprocessed data such as speech, pictures, text, and videos, such as classification, bounding box drawing, annotation, and annotation operations, so that the display device 200 can identify the information in the training data based on the annotation.
[0145] In order to omit manual annotation, embodiments of the present application can combine self-supervised learning tasks during the training process. Taking the training of the multi-modal fusion model as an example, in order to facilitate the distinction between the application stage and the training stage of the multi-modal fusion model, the multi-modal fusion model in the training stage is defined as the multi-modal fusion model to be trained in this embodiment. Before the training process of the display device 200, training data can be obtained. The training data can be video data including video frames, audio data, and line text data. Since subsequent self-supervised learning tasks are adopted, the training data here can be unannotated video data, for example, video data recorded by the user himself.
[0146] The display device 200 extracts video features of video frames through a video encoder, extracts audio features of audio data through an audio encoder, and extracts text features of line text data through a text encoder. After inputting the video features, audio features, and text features into the multi-modal fusion model to be trained, the multi-modal fusion model to be trained will perform multi-modal fusion on the video features, audio features, and text features to obtain multi-modal fusion features in the training stage. Since the accuracy of the multi-modal fusion features in the training stage may not meet the output standard, the display device 200 can calculate the relationship features between the multi-modal fusion features in the training stage and the true multi-modal fusion feature labels through self-supervised learning tasks to calculate the feature fusion loss during multi-modal feature fusion.
[0147] When the feature fusion loss is less than or equal to the multi-modal fusion loss threshold, the display device 200 can output the multi-modal fusion model according to the current model parameters of the multi-modal fusion model to be trained. When the feature fusion loss is greater than the multi-modal fusion loss threshold, the display device 200 needs to perform iterative training on the multi-modal fusion model to be trained until the feature fusion loss is less than or equal to the multi-modal fusion loss threshold.
[0148] In some embodiments, the role relationship model can adopt a training process similar to that of the multi-modal fusion model. However, since the functions performed by the role relationship model and the multi-modal fusion model are different, there are some different steps in the training processes of the role relationship model and the multi-modal fusion model. Specifically, as Figure 10 shown, in the training process of the role relationship model, the display device 200 can obtain a role network graph sample for training the role relationship model, that is, the training data corresponding to the role relationship model. After inputting the role network graph sample into the role relationship model to be trained in the training stage, the role relationship model to be trained can identify the role relationship vectors between the roles according to the role network graph sample and calculate the role relationship weight value in the training stage according to the role relationship vectors.
[0149] The display device 200 can input the role network graph sample and the weight value in the training stage into the graph neural network to output the role relationship feature in the training stage through the graph neural network, so as to determine whether the role relationship feature meets the output feature accuracy through the graph neural network. In the process of the graph neural network generating the role relationship feature, the role relationship training loss between the role relationship feature in the training stage and the true feature label is calculated through a self-supervised learning task, that is, the self-supervised loss of the role relationship model to be trained. When the role relationship training loss is less than or equal to the role relationship loss threshold, the role relationship model is output based on the current model parameters of the role relationship model to be trained. When the role relationship training loss is greater than the role relationship loss threshold, iterative training is performed on the role relationship model to be trained until the role relationship training loss is less than or equal to the role relationship loss threshold.
[0150] It should be noted that in order to improve the accuracy of the role relationship feature in the training stage, the role network graph sample and the weight value in the training stage should be in a corresponding relationship to avoid incorrect role relationship features generated due to the non-correspondence between the weight value and the role network graph sample, which affects the calculation accuracy of the role relationship training loss.
[0151] In some embodiments, the display device 200 can also generate a media understanding result of the played clip data through a target generator that has completed the training of the adversarial generation network. Since the target generator is trained using the adversarial generation network, the accuracy of understanding the played clip data can be improved. The media understanding result includes the relationship between characters and the plot content among the characters in the played clip data. After generating the media understanding result, the display device 200 can parse the first target question of the first interaction instruction, determine the media understanding content part corresponding to the first interaction instruction according to the relationship between characters and the plot content among the characters in the media understanding result, and generate an interaction feedback result based on this part of the media understanding content.
[0152] In some embodiments, during the training process of the target generator, self-supervised learning tasks can also be combined to reduce the dependence of the target generator on labeled data. To this end, before the target generator enters the application stage, the display device 200 can train the target generator through a generative adversarial network that combines self-supervised learning tasks. During the training process of the target generator, a discriminator for competing with the generator can be set. The display device 200 can obtain multi-modal fusion feature samples and character relationship feature samples as the training data for training the generator, and input the multi-modal fusion feature samples and character relationship feature samples into the generator in the training stage for training. During this process, as Figure 11 shown, the generator in the training stage can generate a media understanding result in the training stage according to the multi-modal fusion feature samples and character relationship feature samples. To determine the generation accuracy of the media understanding result in the training stage, the display device 200 will, in the adversarial generation network training stage, use the discriminator to discriminate the media understanding result in the training stage, that is, to discriminate the media understanding result in the training stage from the true label, and output a discrimination result. Among them, the true label is used to represent the true content understanding result of the multi-modal fusion feature samples and character relationship feature samples. Through discrimination, the parameters of the generator in the training stage can be continuously updated, thereby improving the accuracy of the media understanding result in the training stage. When the discrimination result is the first discrimination result, for example, the discrimination result is false, it means that the discriminator can distinguish the media understanding result generated by training from the true label. At this time, the generator in the training stage can be iteratively trained to update the generator parameters, thereby improving the accuracy of the media understanding result generated in the next training process. The discriminator can be continuously iteratively trained based on the training process of the generator to achieve the purpose of adversarial training.
[0153] When the discrimination result is the second discrimination result, for example, the discrimination result is true, it means that the discriminator cannot distinguish the media understanding result in the training stage from the true label at this time. At this time, the generator can generate a media understanding result that is sufficiently close to or the same as the true label. The display device 200 can output the target generator according to the current generator parameters.
[0154] In some embodiments, during the process of training the generator, the display device 200 may calculate the generator loss between the media understanding result and the ground truth label in the training phase through a self-supervised learning task. Among them, the self-supervised learning task may be to perform perturbation processing on a part of the training data for training the generator. For example, perform perturbation processing such as masking, deleting, adding noise, etc. on part of the content of the training data, so that during the training process, the generator infers the perturbed information based on the unperturbed content, thereby realizing self-supervised learning.
[0155] It should be noted that during the training process of the multi-modal fusion model and the role relationship model, the self-supervised learning task can also adopt the above-mentioned perturbation processing method, which will not be elaborated in this embodiment.
[0156] After calculating the generator loss, when the generator loss is greater than the preset loss threshold of the generator, the discriminator outputs the first discrimination result. After obtaining the first discrimination result, the display device 200 may iteratively train the generator. When the generator loss is less than or equal to the loss threshold, the discriminator outputs the second discrimination result. After obtaining the second discrimination result, the display device 200 may output the target generator based on the current generator.
[0157] In some embodiments, as the target media data is played, the playback progress is also continuously increasing. Based on the playback progress, the display screen of the target media data may include newly appeared characters or new plot times, which may lead to changes in the role relationship and the plot. In order to make the generated interactive feedback conform to the plot and role relationship of the target media data, the display device 200 may update the video feature, audio feature, and text feature in real time. For this purpose, as Figure 12 shown, the display device 200 may update the played segment data according to the changed playback progress, where the changed playback progress may be the progress of natural playback over time or the playback progress set by the user himself. For example, in the foregoing example, the user repeatedly watches the media data content of a certain segment. The display device 200 may update the video feature, audio feature, and text feature according to the updated played segment data, and perform multi-modal feature fusion on the updated video feature, updated audio feature, and updated text feature through the multi-modal fusion feature to obtain the updated multi-modal fusion feature.
[0158] In some embodiments, the role network diagram may also be updated based on the updated video features, updated audio features, and updated text features, and the role relationship features may be regenerated. Among them, according to the appearance and disappearance of roles and the direction of plot development, the role relationships may be reduced or increased during the process of updating the role network diagram. For example, for a newly added role, the corresponding part of the role network diagram may be generated according to the relationship between this role and the roles that have appeared. For a role that has disappeared, the display device 200 may delete the part of the role in the role network diagram to simplify the role network diagram.
[0159] In some embodiments, the display device 200 may set update conditions for updating the multi-modal fusion features and role relationship features. For example, an update is performed at the playback progress every preset duration, or an update is performed every time a new role appears. In addition, for a series of dramas, it may include multiple consecutive movies or multiple serial dramas. Therefore, a role that first appears in the target media data may have appeared in other media data in the series of dramas. So, the display device 200 may identify a role that first appears in the target media data. If this role has appeared in other media data in the series of dramas, additional identification information may be generated for this role. For example, "The role appears in Movie A in the series of dramas and is a friend of the main character."
[0160] In some embodiments, after the display device 200 generates an interactive feedback result, the user may input a second interactive instruction according to the interactive feedback result displayed by the display device 200 on the user interface to continuously ask questions to the intelligent viewing program. For example, the user's first interactive instruction is "Is character A the murderer?" The display device 200 may generate an interactive feedback result such as "Character A has the suspicion of being the murderer. However, according to character A's alibi, the probability of character A being the murderer is 30%." If the user wants to continue asking questions, the user may input a second interactive instruction to the display device 200 based on the interactive feedback result, such as "Please help me predict the probability that a character other than character A is the murderer." The display device 200 may parse the second target question of the second interactive instruction and update the interactive feedback result according to the media understanding result. Among them, based on the first target question, the second interactive instruction may be related to the plot content of the played segment data, that is, the second target question is a sub-interactive question of the first target question. The user may also input a second interactive instruction without relying on the interactive feedback result to ask questions to the display device 200 again according to other plot content or character relationships.
[0161] Some embodiments of the present application further provide a viewing interaction method, which is applied to a display device 200. The display device 200 includes a display 260, a memory, and a controller 250. The display 260 is configured to display media data corresponding to media resources; the memory is configured to store a multimodal fusion model; the method includes:
[0162] S100: In response to a first interaction instruction input by the user based on the target media data, obtain the played segment data of the target media data according to the playback progress of the target media data.
[0163] S200: Extract the video features, audio features, and text features of the played segment data.
[0164] S300: Perform multimodal fusion on the video features, the audio features, and the text features through the multimodal fusion model to obtain multimodal fusion features.
[0165] S400: Generate a role network diagram according to the role information identified from the video features, the audio features, and the text features.
[0166] Wherein, the role network diagram is used to represent the role relationships between roles.
[0167] S500: Extract the role relationship features corresponding to the role network diagram through a graph neural network.
[0168] S600: Generate an interaction feedback result corresponding to the first interaction instruction according to the multimodal fusion features and the role relationship features.
[0169] As can be seen from the above technical solutions, the present application provides a display device and a viewing interaction method. The method responds to a first interaction instruction input by the user based on the target media data, obtains the played segment data according to the playback progress, and extracts the video features, audio features, and text features. The video features, audio features, and text features are fused through multimodal fusion technology to obtain multimodal fusion features. The role relationship features are extracted from the role network diagram generated based on the video features, audio features, and text features, so as to generate an interaction feedback result. The present application uses multimodal fusion technology to intelligently and accurately perform content understanding on the target media data through a variety of feature fusion methods, so as to provide the user with an interaction feedback result that conforms to the content of the target media data and improve the accuracy of content understanding.
[0170] For the same and similar parts among the various embodiments in this specification, reference can be made to each other and will not be repeated here.
[0171] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of the various embodiments or some parts of the embodiments of the present invention.
[0172] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.
[0173] For the sake of convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above exemplary discussions are not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
Claims
1. A display device, characterized in that, Comprising: A display configured to display a media interface corresponding to media data; A memory configured to store a multimodal fusion model; A controller configured to: In response to a first interaction instruction input by a user based on target media data, obtain played segment data of the target media data according to the playback progress of the target media data; Extract video features, audio features, and text features of the played segment data; Perform multimodal fusion on the video features, the audio features, and the text features through the multimodal fusion model to obtain multimodal fusion features; Generate a role network graph according to role information identified from the video features, the audio features, and the text features, where the role network graph is used to represent role relationships between roles; Extract role relationship features corresponding to the role network graph through a graph neural network; Generate an interaction feedback result corresponding to the first interaction instruction according to the multimodal fusion features and the role relationship features.
2. The display device according to claim 1, wherein The memory further stores a role relationship model, and the step of the controller extracting role relationship features corresponding to the role network graph through a graph neural network is specifically configured to: Input the role network graph into the role relationship model to calculate weight values of role relationships in the role network graph through the role relationship model; Extract the role relationship features through the graph neural network according to the role network graph and the weight values of the role relationships.
3. The display device according to claim 2, wherein Before the controller executes the step of inputting the role network graph into the role relationship model, it is further configured to: Obtain a role network graph sample for training the role relationship model; Input the role network graph sample into the role relationship model to be trained in the training stage to calculate weight values in the training stage through the role relationship model to be trained; Input the role network graph sample and the weight values in the training stage into the graph neural network to output role relationship features in the training stage; Calculate a role relationship training loss of the role relationship features in the training stage through a self-supervised learning task; When the role relationship training loss is less than or equal to a loss threshold, output the role relationship model based on the current model parameters of the role relationship model to be trained.
4. The display device according to claim 1, wherein The step of the controller generating an interaction feedback result corresponding to the first interaction instruction according to the multimodal fusion features and the role relationship features is specifically configured to: Generate a media understanding result according to the multimodal fusion features and the role relationship features through a target generator that has completed adversarial generation network training; Parse a first target question of the first interaction instruction; Generate the interaction feedback result according to the first target question and the media understanding result.
5. The display device according to claim 4, wherein Before the controller generates a media understanding result according to the multimodal fusion features and the role relationship features through a target generator that has completed adversarial generation network training, it is further configured to: Obtain a multimodal fusion feature sample and a role relationship feature sample; Train a generator in the training stage through the multimodal fusion feature sample and the role relationship feature sample to generate a media understanding result in the training stage; The discriminator outputs a discrimination result of the media resource understanding result in the training stage according to the true label, and the true label is used to represent the true content understanding result of the multi-modal fusion feature sample and the role relationship feature sample; When the discrimination result is the first discrimination result, iterative training is performed on the generator in the training stage to update the generator parameters; When the discrimination result is the second discrimination result, the target generator is output according to the current generator parameters.
6. The display device according to claim 5, wherein The controller executes the step of outputting a discrimination result of the media resource understanding result in the training stage by the discriminator according to the true label, and is specifically configured to: Calculate the self-supervised loss between the media resource understanding result in the training stage and the true label through a self-supervised learning task; When the generator loss is greater than the loss threshold, obtain the first discrimination result output by the discriminator; When the generator loss is less than or equal to the loss threshold, obtain the second discrimination result output by the discriminator.
7. The display device according to claim 1, characterized in that When the playback progress changes, the controller executes the step of performing multi-modal fusion on the video feature, the audio feature, and the text feature through the multi-modal fusion model to obtain a multi-modal fusion feature, and is also configured to: Update the played segment data according to the changed playback progress; Update the video feature, the audio feature, and the text feature according to the updated played segment data; Perform multi-modal fusion on the updated video feature, the updated audio feature, and the updated text feature through the multi-modal fusion model to obtain an updated multi-modal fusion feature.
8. The display device according to claim 4, wherein After the controller executes the step of generating an interaction feedback result corresponding to the first interaction instruction according to the multi-modal fusion feature and the role relationship feature, it is also configured to: In response to a second interaction instruction generated by the user based on the interaction feedback result, parse the second target question of the second interaction instruction; Update the interaction feedback result according to the second target question and the media resource understanding result.
9. The display device according to claim 1, wherein After the controller executes the step of extracting the video feature, the audio feature, and the text feature of the played segment data, it is also configured to: Obtain the video frames of the played segment data; Identify the role expression features of the roles included in the video frames; Perform multi-modal fusion on the video feature, the audio feature, the text feature, and the role expression features through the multi-modal fusion model to obtain a multi-modal fusion feature.
10. A movie-watching interaction method, characterized in that, Applied to a display device, the display device includes a display, a memory, and a controller, and the display is configured to display media data corresponding to the media resource data; The memory is configured to store a multi-modal fusion model; the method includes: In response to a first interaction instruction input by the user based on the target media resource data, obtain the played segment data of the target media resource data according to the playback progress of the target media resource data; Extract the video feature, the audio feature, and the text feature of the played segment data; Perform multi-modal fusion on the video feature, the audio feature, and the text feature through the multi-modal fusion model to obtain a multi-modal fusion feature; Generate a character network graph based on the character information identified from the video features, the audio features, and the text features, where the character network graph is used to represent the character relationships between characters; Extract the character relationship features corresponding to the character network graph through a graph neural network; Generate an interaction feedback result corresponding to the first interaction instruction according to the multimodal fusion features and the character relationship features.