Display device and bullet screen display method
By integrating audio and video acquisition devices into display devices, real-time data collection generates scene description information and extracts personalized portraits, which solves the problem of low accuracy of barrage in existing technologies and improves the accuracy and pertinence of personalized barrage.
Patent Information
- Application Number
- CN202510728163.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-16
AI Technical Summary
Existing display devices display personalized bullet screens based on user login accounts, resulting in low bullet screen accuracy and failure to meet the personalized needs of different users in different scenarios.
By integrating audio and video acquisition devices into the display device, audio data and video data are collected in real time, scene description information is generated, the real-time personalized portrait of the viewer is extracted, and personalized barrage is determined based on the portrait for display.
It improves the accuracy and pertinence of barrage, realizes personalized display across apps and signal sources, and enhances the user experience.
Smart Images

Figure CN120658905A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of display devices, and in particular to a display device and a bullet screen display method. Background Art
[0002] With the development of display devices (such as smart TVs), in order to enhance the fun of content presentation and improve the viewing experience, bullet comments are often superimposed on the content displayed on the display. Typically, the bullet comments displayed on the display are often provided by the content provider and input by the user who views the displayed content.
[0003] However, current display devices customize the barrage based on the user's logged-in account. This approach results in low accuracy in the displayed barrage. For example, if a user who enjoys martial arts films logs into a display device using the account of a user who enjoys romance films, the display device will display personalized barrage based on the user's login behavior. Summary of the Invention
[0004] The present application provides a display device and a barrage display method. During the process of content display on the display of the display device, the barrage is personalized for the viewers in the current scene with the display device as the unit, and the granularity of the personalized barrage is further refined to improve the accuracy and pertinence of the displayed barrage.
[0005] In a first aspect, some embodiments provide a display device, including:
[0006] monitor;
[0007] a communication device configured to connect to the audio and video acquisition device;
[0008] and a controller connected to the communication device and the display, configured to:
[0009] When a bullet screen triggering event is detected during content display on the display, audio data and video data of the current scene of the display device collected by the audio and video collection device are received, and image data of the content currently displayed on the display are obtained; wherein the scene description information is used to describe the environmental characteristics and character characteristics of the current scene; the character characteristics include at least one of the content viewed by the character, the character's dialogue, and the character's basic attributes;
[0010] generating scene description information of the current scene according to the audio data, the video data, and the image data;
[0011] Extracting a real-time personalized portrait of the viewing party in the current scene from the scene description information;
[0012] Determining personalized bullet comments based on the real-time personalized portrait of the viewer in the current scene;
[0013] Control the display to overlay and display the personalized bullet screen.
[0014] In the above embodiment, during the process of displaying content on the display of the display device, the controller of the display device can receive the audio data and video data of the current scene of the display device collected by the audio and video acquisition device, and obtain the image data of the content currently displayed on the display, when detecting the barrage triggering event. Then, based on the above audio data, video data and image data, the scene description information of the current scene is generated. Since the scene description information is used to describe the environmental characteristics and character characteristics of the current scene, and the character characteristics include at least one of the character viewing content, character dialogue and character basic attributes, the real-time personalized portrait of the viewing party in the current scene can be extracted from the scene description information. Furthermore, based on the real-time personalized portrait of the viewing party in the current scene, a personalized barrage can be determined, and the display can be controlled to overlay the display of the determined personalized barrage. In this way, on the one hand, audio and video data can be collected with the help of an audio and video acquisition device. Since audio and video data can reflect the environmental characteristics of the current scene and the personal characteristics of the people in the current scene, the real-time personalized portrait constructed based on audio and video data is essentially to refine the granularity of the real-time personalized portrait based on the characteristics of the display device's usage environment and the user, making up for the fine-grained personalized portrait that the login account cannot perceive; on the other hand, personalized barrage is determined based on the real-time personalized portrait. Since the real-time personalized portrait is determined based on the audio and video data of the current scene and the content currently displayed on the display, the real-time and pertinence of the determined personalized barrage can be improved. Furthermore, the above scheme expands the scope of traditional personalized determination based on the login account, realizing cross-APP (Application) and cross-signal source behavior perception without the user's awareness. In the process of content display on the display of the display device, the barrage is personalized for the viewer in the current scene based on the display device, further refining the granularity of the personalized barrage and improving the accuracy and pertinence of the displayed barrage.
[0015] In a second aspect, some embodiments further provide a bullet screen display method, applied to a display device, comprising:
[0016] When a bullet screen triggering event is detected during content display on the display device, receiving audio data and video data of the current scene of the display device collected by an audio and video collection device, and obtaining image data of the content currently displayed on the display;
[0017] Generate scene description information of the current scene based on the audio data, the video data, and the image data; wherein the scene description information is used to describe environmental characteristics and character characteristics of the current scene; the character characteristics include at least one of the content viewed by the character, the conversations between the character, and basic attributes of the character;
[0018] Extracting a real-time personalized portrait of the viewing party in the current scene from the scene description information;
[0019] Determining personalized bullet comments based on the real-time personalized portrait of the viewer in the current scene;
[0020] Control the display to overlay and display the personalized bullet screen.
[0021] In the above embodiment, during the process of displaying content on the display of the display device, the controller of the display device can receive the audio data and video data of the current scene of the display device collected by the audio and video acquisition device, and obtain the image data of the content currently displayed on the display, when detecting the barrage triggering event. Then, based on the above audio data, video data and image data, the scene description information of the current scene is generated. Since the scene description information is used to describe the environmental characteristics and character characteristics of the current scene, and the character characteristics include at least one of the character viewing content, character dialogue and character basic attributes, the real-time personalized portrait of the viewing party in the current scene can be extracted from the scene description information. Furthermore, based on the real-time personalized portrait of the viewing party in the current scene, a personalized barrage can be determined, and the display can be controlled to overlay the display of the determined personalized barrage. In this way, on the one hand, audio and video data can be collected with the help of an audio and video acquisition device. Since audio and video data can reflect the environmental characteristics of the current scene and the personal characteristics of the people in the current scene, the real-time personalized portrait constructed based on audio and video data is essentially to refine the granularity of the real-time personalized portrait based on the characteristics of the display device's usage environment and the user, making up for the fine-grained personalized portrait that the login account cannot perceive; on the other hand, personalized barrage is determined based on the real-time personalized portrait. Since the real-time personalized portrait is determined based on the audio and video data of the current scene and the content currently displayed on the display, the real-time and pertinence of the determined personalized barrage can be improved. Furthermore, the above scheme expands the scope of traditional personalized determination based on the login account, realizing cross-APP (Application) and cross-signal source behavior perception without the user's awareness. In the process of content display on the display of the display device, the barrage is personalized for the viewer in the current scene based on the display device, further refining the granularity of the personalized barrage and improving the accuracy and pertinence of the displayed barrage.
[0022] In a third aspect, some embodiments further provide a bullet screen display device, applied to a display device, comprising:
[0023] a data acquisition module for, when a bullet screen triggering event is detected during content display on the display of the display device, receiving audio data and video data of the current scene of the display device collected by the audio and video acquisition device, and acquiring image data of the content currently displayed on the display;
[0024] an information generation module, configured to generate scene description information of the current scene based on the audio data, the video data, and the image data; wherein the scene description information is used to describe environmental characteristics and character characteristics of the current scene; the character characteristics include at least one of the content viewed by the character, the character's dialogue, and the character's basic attributes;
[0025] A portrait extraction module, configured to extract a real-time personalized portrait of the viewing party in the current scene from the scene description information;
[0026] A barrage determination module, configured to determine personalized barrages based on a real-time personalized portrait of the viewer in the current scene;
[0027] The bullet screen display module is used to control the display to overlay and display the personalized bullet screen.
[0028] In the above embodiment, during the process of displaying content on the display of the display device, the controller of the display device can receive the audio data and video data of the current scene of the display device collected by the audio and video acquisition device, and obtain the image data of the content currently displayed on the display, when detecting the barrage triggering event. Then, based on the above audio data, video data and image data, the scene description information of the current scene is generated. Since the scene description information is used to describe the environmental characteristics and character characteristics of the current scene, and the character characteristics include at least one of the character viewing content, character dialogue and character basic attributes, the real-time personalized portrait of the viewing party in the current scene can be extracted from the scene description information. Furthermore, based on the real-time personalized portrait of the viewing party in the current scene, a personalized barrage can be determined, and the display can be controlled to overlay the display of the determined personalized barrage. In this way, on the one hand, audio and video data can be collected with the help of an audio and video acquisition device. Since audio and video data can reflect the environmental characteristics of the current scene and the personal characteristics of the people in the current scene, the real-time personalized portrait constructed based on audio and video data is essentially to refine the granularity of the real-time personalized portrait based on the characteristics of the display device's usage environment and the user, making up for the fine-grained personalized portrait that the login account cannot perceive; on the other hand, personalized barrage is determined based on the real-time personalized portrait. Since the real-time personalized portrait is determined based on the audio and video data of the current scene and the content currently displayed on the display, the real-time and pertinence of the determined personalized barrage can be improved. Furthermore, the above scheme expands the scope of traditional personalized determination based on the login account, realizing cross-APP (Application) and cross-signal source behavior perception without the user's awareness. In the process of content display on the display of the display device, the barrage is personalized for the viewer in the current scene based on the display device, further refining the granularity of the personalized barrage and improving the accuracy and pertinence of the displayed barrage.
[0029] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the method provided in some embodiments of the second aspect are implemented.
[0030] In a fifth aspect, a computer program product is provided, which includes: a computer program, which, when executed by a processor, implements the steps of the method provided in some embodiments of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0032] Figure 1 A schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application;
[0033] Figure 2 A schematic diagram of the hardware configuration of a display device provided in some embodiments of the present application;
[0034] Figure 3 A schematic diagram of the hardware configuration of a control device provided in some embodiments of the present application;
[0035] Figure 4 A schematic diagram of software configuration of a display device provided in some embodiments of the present application;
[0036] Figure 5 A flowchart of a bullet screen display method provided in some embodiments of the present application;
[0037] Figure 6 A schematic diagram of a process for determining personalized bullet comments provided in some embodiments of the present application;
[0038] Figure 7 A flowchart of a bullet screen display method provided in some other embodiments of the present application;
[0039] Figure 8 A flowchart of a bullet screen display method provided in some embodiments of the present application;
[0040] Figure 9 A flowchart of a bullet screen display method provided in some further embodiments of the present application;
[0041] Figure 10A signaling diagram of the bullet screen display method provided in some embodiments of the present application;
[0042] Figure 11 A schematic structural diagram of a bullet screen display device provided in some embodiments of the present application. DETAILED DESCRIPTION
[0043] The following embodiments are described in detail, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following embodiments are not intended to represent all possible implementations consistent with the present application. They are merely examples of systems and methods consistent with certain aspects of the present application, as detailed in the claims.
[0044] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.
[0045] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.
[0046] The terms "comprise," "comprises," and "having," and any variations thereof, are intended to cover but not exclude inclusion; for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0047] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functionality associated with that element.
[0048] In the embodiments of the present application, the display device 200 generally refers to a device capable of displaying images and processing data. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.
[0049] Figure 1 This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG, a user can operate the display device 200 through touch operation, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote controller, a stylus pen, a handle, etc.
[0050] The mobile terminal 300 can function as a control device for performing human-computer interaction between a user and the display device 200. The mobile terminal 300 can also function as a communication device for establishing a communication connection with the display device 200 for data exchange. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, enabling communication via a network communication protocol for one-to-one control and data communication. Audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 for synchronized display.
[0051] like Figure 1 As shown in FIG, the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.
[0052] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart TV, Internet Protocol television (IPTV), etc.
[0053] Figure 2 Some embodiments of this application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.
[0054] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0055] In some embodiments, detector 230 is used to collect signals from the external environment or external interactions. For example, detector 230 may include a light receiver, such as a sensor for collecting ambient light intensity; or an image collector, such as a camera, for collecting external environmental scenes, user attributes, or user interaction gestures; or a sound collector, such as a microphone, for receiving external sounds.
[0056] In some embodiments, the display 260 includes a display component for presenting images and a driver component for driving image display. The display 260 is configured to receive image signals output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces.
[0057] In some embodiments, the communication device 220 is a component used to communicate with an external device or server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 depending on the supported communication methods. For example, if the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 including WiFi functionality. If the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including Bluetooth functionality.
[0058] The communication device 220 can establish a communication connection between the display device 200 and an external device or server 400 via a wireless or wired connection. A wired connection can connect the display device 200 to an external device via a data cable, an interface, or other components. A wireless connection can connect the display device 200 to an external device via a wireless signal or wireless network. The display device 200 can establish a connection with an external device directly or indirectly through a gateway, router, or connection device.
[0059] In some embodiments, the controller 250 may include at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processor, and a power processor, and first to nth interfaces for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in a memory. The controller 250 controls the overall operation of the display device 200.
[0060] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0061] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).
[0062] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may further be provided with an external audio output terminal, through which the audio output device may be connected to the display device 200 to output the sound of the display device 200.
[0063] In some embodiments, the user input interface 280 may be configured to receive instructions from a user.
[0064] Figure 3 Some embodiments of this application provide Figure 1 The hardware configuration diagram of the control device in the figure is as follows. Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.
[0065] The control device 100 is configured to control the display device 200 , and can receive user input operation instructions, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, thereby acting as an interaction intermediary between the user and the display device 200 .
[0066] In some embodiments, the control device 100 may be a smart device. For example, the control device 100 may be installed with various applications for controlling the display device 200 according to user needs.
[0067] In some embodiments, as Figure 1 As shown, the mobile terminal 300 or other intelligent electronic devices can play a similar function as the control device 100 after installing the application for controlling the display device 200 .
[0068] Controller 110 includes a processor 112, random access memory (RAM) 113, read-only memory (ROM) 114, a communication interface 130, and a communication bus. Controller 110 is used to control the operation and functionality of control device 100, as well as communication and coordination between internal components and external and internal data processing.
[0069] Under the control of the controller 110, the communication interface 130 communicates control signals and data signals with the display device 200. The communication interface 130 may include at least one of a WiFi chip 131, a Bluetooth module 132, an NFC (Near Field Communication) module 133, or other near field communication modules.
[0070] The user input / output interface 140 includes at least one of a microphone 141 , a touch panel 142 , a sensor 143 , a button 144 and other input interfaces.
[0071] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with the communication interface 130, such as a WiFi, Bluetooth, or NFC module, to encode user input commands via the WiFi protocol, Bluetooth protocol, or NFC protocol and transmit them to the display device 200.
[0072] The memory 190 is used to store various operating programs, data and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.
[0073] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.
[0074] To facilitate user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling the hardware and software resources in the display device 200. The operating system may provide a user interface (control the display device), allow the user to interact with the display device 200, and support the running of various application programs.
[0075] It should be noted that the operating system may be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.
[0076] The operating system can be divided into different modules or layers according to the functions implemented.
[0077] For example, Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely the application layer (abbreviated as "application layer"), the application framework layer (abbreviated as "framework layer"), the system library layer and the kernel layer.
[0078] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer can host at least one application, which can include built-in window programs, system settings programs, clock programs, and the like, or applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the examples above.
[0079] The framework layer provides applications with an application programming interface (API) and programming framework. The application framework layer includes predefined functions. The application framework layer acts as a processing center, determining the actions taken by applications in the application layer. Through the API, applications can access system resources and services during execution.
[0080] like Figure 4As shown, in the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc., wherein the view system can design and implement the interface and interaction of the application, and the view system includes lists, grids, text boxes, buttons, etc. The manager includes at least one of the following modules: an activity manager for interacting with all activities running in the system; a location manager for providing system services or applications with access to the system location service; a package manager for retrieving various information related to the application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0081] In some embodiments, the activity manager is used to manage the lifecycle of each application and common navigation back functions, such as controlling application exit, opening, and back. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, taking screenshots, and controlling changes in display windows, such as shrinking, shaking, or distorting the display window.
[0082] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions to be implemented by the framework layer.
[0083] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. Figure 4 As shown, the kernel layer can be configured with hardware drivers, and the drivers included in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB (Universal Serial Bus) driver, HDMI (High Definition Multimedia Interface) driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.
[0084] It should be noted that the above example is only a simple division of the operating system functions and does not constitute a limitation on the specific operating system form of the display device 200 in the embodiment of the present application. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.
[0085] With the development of display devices (such as smart TVs), in the process of displaying content on the display of the current display device, in order to enhance the fun of the content display and improve the viewing experience of the viewer, barrages are often superimposed on the content displayed on the display. Usually, the barrages displayed on the display are often derived from the real barrages input by the viewing users of the displayed content provided by the content provider. However, the current display device displays barrages in a personalized manner based on the user's logged-in account. This approach makes the accuracy of the barrages presented low. Based on this, in some embodiments, a barrage display method is provided. Among them, the barrage display method can be implemented by a display device.
[0086] In an exemplary embodiment, the bullet screen display method is applied to a controller in a display device as an example, wherein the display device includes a display, a communication device, and a controller.
[0087] The communication connection between the display and the controller can be a wired communication connection or a wireless communication connection.
[0088] The communication connection between the communication device and the audio and video acquisition device can be a wired communication connection or a wireless communication connection.
[0089] like Figure 5 As shown, the bullet screen display method applied to the controller in the display device may include the following steps:
[0090] S501, when a bullet screen triggering event is detected during content display on the display, audio data and video data of the current scene of the display device collected by the audio and video collection device are received, and image data of the content currently displayed on the display is obtained.
[0091] Typically, barrages are displayed superimposed on the content displayed on the display to increase the interest of the displayed content and improve the viewing experience of the viewer. Therefore, barrages are usually displayed while the display is displaying content.
[0092] In addition, the display of barrage is usually triggered by some triggering events, for example, the viewer sends a barrage display instruction to the display device by clicking the relevant button on the remote control, the display plays a specific type of display content (such as TV series, movies, variety shows), "barrage" and other related information appear in the communication information of the viewer in the current scene, and specific content (such as specific actors, specific buildings, specific plots) appears in the content displayed on the display.
[0093] In an optional embodiment, detecting a bullet screen triggering event may include any of the following:
[0094] 1. When it is detected that the display starts to display content (such as the display device is turned on), it is determined that a bullet screen triggering event is detected.
[0095] 2. When it is detected that preset content (such as preset buildings, preset actors, preset plots, etc.) appears in the content displayed on the display, it is determined that a barrage triggering event is detected.
[0096] 3. When it is detected that the audio data of the current scene contains preset content (such as barrage, opening barrage, displaying barrage, etc.), it is determined that a barrage trigger event is detected.
[0097] 4. When it is detected that the display starts to display a specified type of content (such as a live broadcast, TV series, movie, variety show, etc.), it is determined that a barrage triggering event is detected.
[0098] 5. When it is detected that a bullet screen display instruction (such as a voice interaction instruction, an instruction sent through a remote control, etc.) is received, it is determined that a bullet screen triggering event is detected.
[0099] Based on this, when a bullet screen triggering event is detected during content display on the display, it can be determined that the bullet screen display is triggered, and it is necessary to further determine the bullet screen to be displayed on the display. In order to be able to determine the personalized bullet screen for the viewer in the current scene where the display device is located, so as to realize the personalized bullet screen display for the viewer in the above-mentioned on-site scene, it is possible to first receive the audio data and video data of the current scene where the display device is located collected by the audio and video collection device, and obtain the image data of the content currently displayed on the display.
[0100] In some embodiments, the controller can send a data acquisition request to the audio and video acquisition device, and the audio and video acquisition device responds to the data acquisition request and feeds back to the controller the audio data and video data of the current scene in which the display device is located. Then, the controller can receive the audio data and video data of the current scene in which the display device is located collected by the audio and video acquisition device.
[0101] Optionally, the audio and video acquisition device is a device that integrates audio acquisition and video acquisition functions. The audio and video acquisition device includes an independent audio acquisition device (such as a microphone, a microphone array) for audio acquisition and a video acquisition device (such as a camera, etc.) for video acquisition. Accordingly, in the above S501, obtaining the audio data and video data of the current scene of the display device collected by the audio and video acquisition device may include: obtaining the audio data of the current scene of the display device collected by the audio acquisition device, and obtaining the video data of the current scene of the display device collected by the video acquisition device.
[0102] Optionally, the above-mentioned audio and video acquisition device can be provided on the display device, or can be independent of the display device. In an optional embodiment, the above-mentioned audio and video acquisition device includes an audio acquisition device and a video acquisition device, and the display device further includes: an audio acquisition device configured to acquire audio data, and / or a video acquisition device configured to acquire video data. That is to say, in the case where the above-mentioned audio acquisition device includes a separately independent audio acquisition device and video acquisition device, the display device can be provided with only an audio acquisition device, while the video acquisition device is provided outside the display device; the display device can also be provided with only a video acquisition device, while the audio acquisition device is provided outside the display device; the display device can also be provided with both an audio acquisition device and a video acquisition device.
[0103] Optionally, the audio data includes, but is not limited to, communication information between people in the current scene, such as comments on the content displayed on the display, conversation information between people, etc. For example, the audio data may also include ambient sounds of the current scene (such as wind noise, audio information output by other electronic devices), voice interaction instructions from people in the current scene to the display device, etc. The video data includes both ambient data and personnel data of the current scene.
[0104] Optionally, the display device may have a screenshot function, so that the display may be screenshoted while the display is displaying content, and thus each captured screen image is image data of the content currently displayed on the display.
[0105] Optionally, an image acquisition device for capturing images of the display screen may be provided, so that, while the display is displaying content, the image acquisition device captures images of the display screen to obtain image data of the content currently displayed on the display. The image acquisition device may be provided on the display device or may be independent of the display device. In an optional embodiment, the display device further includes an image acquisition device configured to capture image data. Accordingly, obtaining image data of the content currently displayed on the display in S501 may include obtaining image data of the content currently displayed on the display captured by the image acquisition device.
[0106] In order to ensure that the personalized bullet screen displayed is targeted and real-time to the viewers in the current scene, the above-mentioned audio data and video data are collected in real time by the audio and video acquisition device, and the image data of the content currently displayed on the above-mentioned display is also generated in real time. Based on this, when the display is displaying content, when a bullet screen triggering event is detected, the controller of the display device can receive in real time the audio data and video data of the current scene of the display device collected in real time by the audio and video acquisition device, and obtain in real time the image data of the content currently displayed on the display generated in real time.
[0107] S502: Generate scene description information of the current scene based on the audio data, video data, and image data.
[0108] The scene description information of the current scene is used to describe the environmental characteristics and character characteristics of the current scene; the character characteristics include at least one of the content viewed by the character, the character dialogue, and the character's basic attributes.
[0109] After obtaining the above-mentioned audio data, video data and image data, the above-mentioned audio data, video data and image data can be processed in a composite or single-modal manner using a multimodal large model, an expert model in a specific field, etc., to obtain scene description information of the current scene.
[0110] Scenario description information is a detailed description of a specific situation or environment, clarifying elements such as "when, where, who, what, what needs / goals, and the surrounding environment." It transforms abstract scenarios into understandable and analyzable descriptions through structured or semi-structured information.
[0111] Based on this, the scene description information of the current scene is used to describe the environmental characteristics and character characteristics of the current scene. Among them, the information describing the environmental characteristics includes information describing the characteristics of the environment in which the display device is located, for example, the room type (such as living room, bedroom, etc.), furniture style, light brightness, room area, etc. The information describing the character characteristics includes at least one of information describing the content viewed by the character, information describing the character conversation, and information describing the basic attributes of the character. Among them, the information describing the content viewed by the character includes information describing the content displayed by the display, for example, the content type (such as games, movies, sports, shopping), changes in the content type, the characters and scenes in the content displayed by the display, and the search keywords entered by the user and the search result content displayed in response to the search keywords. The information describing the character conversation includes information describing the conversation voice of the characters in the current scene, the commentary voice of the characters in the current scene on the content displayed on the display, and the human-computer interaction voice between the characters in the current scene and the display device. In some embodiments, the characters in the current scene may include not only the viewing party in the current scene but also characters passing by the current scene. While passing by the current scene, the characters may engage in conversation with the viewing party in the current scene or comment on the content displayed on the display. Therefore, the information describing the characters' conversations may also include information describing the voices of the characters passing by the current scene. The information describing the basic attributes of the characters includes information related to the characteristics of the characters in the current scene, such as number, age, gender, clothing style, and emotion.
[0112] Optionally, for the above audio data, an ASR (Automatic Speech Recognition) model can be used for speech recognition to obtain speech information such as the evaluation speech of the characters in the current scene, the characters' conversational speech, and the interactive speech between the characters and the display device. For the above video data, a CV (Computer Vision) model can be used to perform recognition processing such as face recognition, age recognition, emotion recognition, behavior recognition, and environment recognition to obtain basic attribute information of the characters in the current scene (such as age, gender, emotion, number, behavior, etc.) and environmental attribute information (such as home furnishings, decoration style, brightness information, scene area, etc.). For the above image data, multiple models such as CNN (Convolutional Neural Network) models and TSN (Temporal Segment Network) models can be used to identify the content and characters in the image data to obtain viewing content information, such as content type (such as games, movies, sports, shopping), fine-grained classification within specific types (such as martial arts, romance, action, and era within the movie genre), changes in content types, and dwell time within each type. Thus, by combining the various types of information identified above, scene description information for the current scene can be comprehensively generated.
[0113] S503: Extracting a real-time personalized portrait of the viewer in the current scene from the scene description information.
[0114] The viewers in the current scene may be all viewers in the current scene, or may be some viewers in the current scene.
[0115] Optionally, when the viewing party in the current scene is some of the viewers in the current scene, the biometrics of the viewers in the current scene can be determined based on the above-mentioned audio data and / or video data, and thus, by matching the biometrics of the above-mentioned viewers with the biometrics of the registered persons, the registered persons in the current scene can be determined as the viewing party in the current scene.
[0116] The so-called personalized portrait refers to the collection and analysis of multi-dimensional data of an individual to extract a virtual image summary that can accurately describe his or her characteristics, behaviors, preferences, needs and other attributes.
[0117] As previously mentioned, the audio data, video data, and image data described above can all be collected in real time within the current scene of the display device. Therefore, the so-called real-time personalized portrait of the viewer in the current scene can be understood as a real-time, accurate description of the viewer's current state, behavior, preferences, needs, and other attributes extracted based on the audio data, video data, and image data collected in real time within the current scene. In other words, the real-time, accurate description of the viewer's personal characteristics in the current scene can be provided in real time.
[0118] In this way, after the scene description information of the current scene is generated based on the above-mentioned audio data, video data and image data, a real-time personalized portrait of the viewer in the current scene can be extracted from the scene description information.
[0119] Because the scene description information of the current scene may include multi-dimensional information for describing the environmental characteristics and character characteristics of the current scene. Optionally, the multi-dimensional information includes viewing personnel information (scene description information describing basic attributes of characters), viewing environment information (scene description information describing environmental characteristics), viewing content information (scene description information describing the content viewed by characters), and viewing dialogue information (scene description information describing dialogues between characters). Of course, the multi-dimensional information may also include information of other dimensions, which is not specifically limited. For example, the device information of the display device, such as model, brand, memory size, etc., is extracted as the device information included in the description information of the current scene.
[0120] Furthermore, personalized information in four dimensions, namely, viewing person information, viewing environment information, viewing content information, and viewing conversation information, can be extracted from the scene description information to obtain a real-time personalized portrait of the viewing party in the current scene.
[0121] For example, in the dimension of viewer information, the number of viewers, age of viewers, gender of viewers, clothing characteristics of viewers (such as fashionable, elegant, cute, etc.), and emotions of viewers are extracted as personalized information; in the dimension of viewing environment, the room category (such as living room, bedroom, etc.), lighting brightness, home style, and scene area are extracted as personalized information; in the dimension of viewing content information, the content watched by viewers, the content searched by viewers, the applications used by viewers, the people displayed on the display that viewers are interested in, and the objects displayed on the display that viewers are interested in are extracted as personalized information; in the dimension of viewing conversation information, the content discussed by viewers, the content of conversations between viewers, and the human-computer interaction content between viewers and display devices are extracted as personalized information. At this point, the personalized information extracted in each dimension is used as part of the real-time personalized portrait of the viewer in the current scene, thereby obtaining a real-time personalized portrait of the viewer in the current scene.
[0122] S504: Determine personalized bullet comments based on the real-time personalized portrait of the viewer in the current scene.
[0123] The so-called personalized barrage refers to the dynamic generation or recommendation of barrage content that meets the needs of different users based on data such as their personal characteristics, preferences, behavioral habits or real-time scenarios.
[0124] As mentioned above, the real-time personalized portrait of the viewer in the current scene can describe the real-time personalized characteristics of the viewer in the current scene. Based on the real-time personalized portrait of the viewer in the current scene, personalized barrage can be determined to meet the barrage needs of the viewer in the current scene and realize accurate barrage display for the viewer in the current scene.
[0125] Optionally, when the content displayed on the display has barrage resources, content tags can be added to the barrage resources based on the content of each barrage resource. Thus, the barrage resources whose content tags match the real-time personalized portrait of the viewer in the current scene can be determined as personalized barrages.
[0126] Optionally, based on the information of each dimension represented by the real-time personalized portrait of the viewing party in the current scene, a barrage with content that conforms to the real-time personalized portrait of the viewing party in the above-mentioned current scene can be generated as a personalized barrage. For example, personalized barrage is generated by AIGC (Artificial Intelligence Generated Content). Exemplarily, when the content displayed on the display does not have barrage resources, or the barrage resources are relatively small, based on the information of each dimension represented by the real-time personalized portrait of the viewing party in the current scene, a barrage with content that conforms to the real-time personalized portrait of the viewing party in the above-mentioned current scene can be generated as a personalized barrage.
[0127] S505, controlling the display to overlay and display the personalized bullet screen.
[0128] After determining the above-mentioned personalized barrage, the display can be controlled to display the above-mentioned personalized barrage on the upper layer of the displayed content during the process of content display. Then, during the process of content display, the viewer in the current scene can not only view the content originally displayed by the display, but also view the personalized barrage for the viewer in the current scene, thereby improving the viewing experience of the viewer in the current scene.
[0129] Optionally, in an optional embodiment, during the process of displaying content, when a barrage triggering event is detected, the above S501-S505 can be executed periodically. Thus, during the process of displaying content, personalized barrage can be continuously superimposed and displayed on the display according to a set period to meet the viewing needs of viewers in the current scene and improve the viewing experience of viewers in the current scene.
[0130] In the above embodiment, during the process of displaying content on the display of the display device, the controller of the display device can receive the audio data and video data of the current scene of the display device collected by the audio and video acquisition device, and obtain the image data of the content currently displayed on the display, when detecting the barrage triggering event. Then, based on the above audio data, video data and image data, the scene description information of the current scene is generated. Since the scene description information is used to describe the environmental characteristics and character characteristics of the current scene, and the character characteristics include at least one of the character viewing content, character dialogue and character basic attributes, the real-time personalized portrait of the viewing party in the current scene can be extracted from the scene description information. Furthermore, based on the real-time personalized portrait of the viewing party in the current scene, a personalized barrage can be determined, and the display can be controlled to overlay the display of the determined personalized barrage. In this way, on the one hand, audio and video data can be collected with the help of an audio and video acquisition device. Since audio and video data can reflect the environmental characteristics of the current scene and the personal characteristics of the people in the current scene, the real-time personalized portrait constructed based on audio and video data is essentially to refine the granularity of the real-time personalized portrait based on the characteristics of the display device's usage environment and the user, making up for the fine-grained personalized portrait that the login account cannot perceive; on the other hand, personalized barrage is determined based on the real-time personalized portrait. Since the real-time personalized portrait is determined based on the audio and video data of the current scene and the content currently displayed on the display, the real-time and pertinence of the determined personalized barrage can be improved. Furthermore, the above scheme expands the scope of traditional personalized determination based on the login account, realizing cross-APP (Application) and cross-signal source behavior perception without the user's awareness. In the process of content display on the display of the display device, the barrage is personalized for the viewer in the current scene based on the display device, further refining the granularity of the personalized barrage and improving the accuracy and pertinence of the displayed barrage.
[0131] Based on the above embodiment, in an exemplary embodiment, the generation of the scene description information in S502 is further refined. The scene description information can be generated by inputting audio data, video data, and image data into a multimodal large model to obtain scene description information of the current scene.
[0132] In this embodiment, a multimodal large model (such as a CLIP (Contrastive Language-Image Pre-training) model) is an artificial intelligence model that can simultaneously process and understand information from multiple different modalities, where the multiple different modalities may include text, images, audio, video, etc. Therefore, the above-mentioned audio data, video data, and image data can be input into the multimodal large model to obtain descriptive information about the scene obtained by the multimodal large model through processing and recognition of various types of data, which serves as the scene description information of the current scene.
[0133] Among them, the multimodal large model can convert data from different modalities into data features in the same semantic space through modal alignment. The so-called semantic space is used to describe the semantic relationships and structured representations of information such as symbols, language, images, and sounds. The semantic space converts abstract semantics into a computable spatial structure through mathematical modeling, enabling computers to understand and process the semantic associations of human language or multimodal data. Therefore, the multimodal large model can accurately map, associate, and interact with data from different modalities through modal alignment, eliminate differences in the representation of data from different modalities, and establish cross-modal semantic consistency to complete cross-modal tasks.
[0134] Based on this, in an optional embodiment, the multimodal large model may include a modal alignment module and a semantic recognition module. Then, the above-mentioned inputting of audio data, video data and image data into the multimodal large model to obtain scene description information of the current scene may include inputting the audio data, video data and image data into the modal alignment module for modal alignment to obtain data features of the audio data, video data and image data in the same semantic space; inputting the data features into the semantic recognition module for semantic recognition to obtain scene description information of the current scene.
[0135] In this embodiment, the modal alignment module in the multimodal large model is used to perform modal alignment on input data of different modalities to convert the input data of different modalities into data features in the same semantic space. Therefore, the above-mentioned audio data, video data, and image data can be first input into the modal alignment module of the multimodal large model for modal alignment. The modal alignment module can then output feature data of the converted audio data, video data, and image data in the same speech space.
[0136] Furthermore, the semantic recognition module of the multimodal large model can identify the deep semantic information implied in the feature data obtained by the above conversion, so as to extract the meaning, emotion, scene, logical relationship and other information expressed by the above audio data, video data and image data, and obtain the scene description information of the current scene reflected by the above audio data, video data and image data.
[0137] Optionally, the modal alignment module in the above-mentioned multimodal large model can be trained by taking the audio data, video data and image data of the historical scene in which the display device is located as input, and taking the description information of the audio data, video data and image data of the historical scene in which the display device is located in the same semantic space as the label; and the semantic recognition module of the above-mentioned multimodal large model can be trained by taking the audio data, video data and image data of the historical scene in which the display device is located as input, and taking the scene description information of the historical scene in which the display device is located as the label. At this point, when the audio data, video data and image data of the above-mentioned current scene are input into the trained multimodal large model, the scene description information of the current scene can be output by the multimodal large model.
[0138] In this embodiment, the use of a multimodal large model to generate scene description information for the current scene can integrate scene information represented by data from different modalities, such as audio data, video data, and image data, breaking through the limitations of a single modality. This makes the generated scene information for the current scene richer, more comprehensive, and more accurate. Furthermore, data from different modalities can complement each other, which can also improve the robustness of the scene description information obtained for the current scene. This can further improve the accuracy and pertinence of the extracted real-time personalized portrait of the viewer in the current scene, further improving the accuracy and pertinence of the displayed barrage.
[0139] Based on the above embodiments, in an exemplary embodiment, the extraction of the real-time personalized portrait in S503 is further refined. The real-time personalized portrait can be extracted by extracting historical operation data of the viewer on the display device in the current scene from the local storage data of the display device; and extracting the real-time personalized portrait of the viewer in the current scene from the scene description information and the historical operation data.
[0140] In this embodiment, it is understood that while the display is displaying content, people in the current scene can perform various operations on the display device through various methods, such as voice commands, operating a pointing remote control, and touchscreen operations. For example, they can use voice commands to control the display to switch displayed content, operate a pointing remote control to open an application, use a touchscreen operation on the display to circle an actor's face, and operate a pointing remote control to move the pointing remote cursor around the actor's facial contour in the display. Therefore, the display device's local storage data can record historical operation data of the viewer on the display device in the current scene, such as the display device's operation log.
[0141] It is understandable that the historical operation data of the viewer on the display device in the current scene can also reflect the viewer's personalized characteristics to a certain extent. For example, historical operation data of frequently opening a certain application can reflect the viewer's high attention to the application's functions in the current scene, while historical operation data of controlling the remote control cursor to move around the facial contours of an actor on the display can reflect the viewer's high attention to the actor in the current scene. Therefore, the viewer's historical operation data on the display device in the current scene can participate in the process of extracting the viewer's personalized portrait in the current scene.
[0142] Based on this, after generating the scene description information of the current scene, we can first extract the historical operation data of the viewer on the display device in the current scene from the local storage data of the display device, and then extract the real-time personalized portrait of the viewer in the current scene from the above scene description information and historical operation data.
[0143] Optionally, viewing content feature information describing the content being viewed by a person can be extracted from the scene description information and historical operation data, respectively. The viewing content feature information extracted from the historical operation data can then be used to adjust the viewing content feature information extracted from the scene description information. For example, the weights of attention paid to different content types can be adjusted, or the number of people with high attention can be increased or decreased. This adjusted viewing content feature information is then used as a portrait in the viewing content information dimension of the personalized portrait. Furthermore, portraits in the viewing person information, viewing environment information, and viewing conversation information dimensions can be extracted from the scene description information. This results in a real-time personalized portrait of the viewer in the current scene.
[0144] In this embodiment, the historical operation data of the display device by the viewer in the current scene participates in the process of extracting the personalized portrait of the viewer in the current scene, which can enrich the extracted personalized portrait of the viewer in the current scene. In addition, the scene description information of the current scene can be corrected through the historical operation data of the display device by the viewer in the current scene, which can also improve the accuracy and pertinence of the extracted personalized portrait of the viewer in the current scene, and further improve the accuracy and pertinence of the displayed barrage.
[0145] It's understandable that human preferences, needs, and other characteristics can change over time, and thus the personalized profile of the viewer in the current scenario can also change over time. For example, in a family TV viewing scenario, the viewer's preference for different actors, their preference for different types of content, as well as their personality, clothing style, and furniture style can all change over time. Therefore, when determining the viewer's personalized profile, it is necessary to fully consider both the short-term changes and long-term stability of personalized characteristics.
[0146] Short-term changes in a viewer's personalized characteristics can be reflected through their real-time personalized profile, enabling real-time tracking of preferences that change over time. Long-term stability of a viewer's personalized characteristics can be reflected through their historical personalized profile. Furthermore, personalized comments can be determined based on both the viewer's real-time and historical personalized profiles.
[0147] Based on this, on the basis of the above embodiment, in an exemplary embodiment, the determination of the personalized bullet screen in the above S504 is further refined. Figure 6 As shown, the following steps may be included:
[0148] S601, performing image fusion on the historical personalized portrait and the real-time personalized portrait associated with the viewer to obtain a comprehensive personalized portrait.
[0149] The historical personalized portrait associated with the viewer includes information describing the features of the content viewed by the viewer within a preset historical period.
[0150] Understandably, the so-called historical personalized profile is a collection of user feature tags constructed through data analysis and algorithmic models based on the user's behavioral data over a period of time. It focuses on the user's historical behavior patterns and is used to describe the user's long-term and stable interests, habits, and needs.
[0151] Based on this, the historical personalized portrait associated with the viewer includes information describing the characteristics of the content viewed by the viewer within a preset historical period. In other words, the historical personalized portrait associated with the viewer can reflect the long-term stability of the viewer's preferences for viewing content within the preset historical period. Through the historical personalized portrait associated with the viewer, it is possible to extract the content type characteristics (such as games, movies, sports, shopping) of the viewing content that the viewer is interested in viewing within the preset historical period, the fine-grained classification characteristics of specific types (such as martial arts, romance, action, era, etc. under the film and television type), the character characteristics of the viewing content that the viewer is interested in viewing (such as celebrities, actors, athletes, etc.), and the event characteristics of the viewing content that the viewer is interested in viewing (such as entertainment events, etc.).
[0152] Optionally, the historical personalized portrait associated with the viewer may also include information describing the viewer's basic attributes during a preset historical period. In other words, the historical personalized portrait associated with the viewer may reflect the long-term stability of the viewer's clothing style, etc., during the preset historical period. Optionally, the historical personalized portrait associated with the viewer may also include information describing the environmental characteristics of the display device's environment during the preset historical period. In other words, the historical personalized portrait associated with the viewer may also reflect the long-term stability of the display device's environment, such as furniture style and lighting brightness, during the preset historical period.
[0153] Among them, the specific length of the above-mentioned preset historical period and its time relationship with the current moment can be set based on experience values, test values of multiple tests, and the needs of actual application scenarios, and no specific limitation is made on this.
[0154] After obtaining the real-time personalized portrait of the viewer in the current scene, we can further obtain the historical personalized portrait associated with the viewer in the current scene, and then fuse the historical personalized portrait and the real-time personalized portrait to obtain a comprehensive personalized portrait.
[0155] Among them, the comprehensive personalized portrait can have the long-term characteristics of the viewer in the current scene based on the display device, as well as the real-time characteristics under the real-time viewing state. Therefore, the error jitter of the real-time personalized portrait caused by randomness can be reduced, and the specificity of the real-time personalized portrait can be enriched.
[0156] Optionally, by integrating historical personalized portraits and real-time personalized portraits, using weighted calculations and combining time decay factors to handle conflicting or complementary features, a comprehensive personalized portrait that takes into account both stability and timeliness can be formed to achieve accurate characterization and dynamic adaptation of the user's diverse needs. For example, the historical personalized portraits and real-time personalized portraits associated with the viewing party in the current scene may include information of the same dimension, and different weights can be assigned to the information of the same dimension in the historical personalized portraits and the real-time personalized portraits. Thus, by weighting and summing the information of each same dimension in the historical personalized portraits and the real-time personalized portraits, new information of the same dimension can be obtained. The new information of each same dimension is the information of the dimension in the comprehensive personalized portrait, and then the comprehensive personalized portrait is obtained. Among them, the above-mentioned weights can be set according to experience values, test values of multiple tests, and the needs of actual application scenarios, and there is no specific limitation on this.
[0157] In an optional embodiment, the historical personalized portraits and real-time personalized portraits associated with the viewing party in the current scene can be used as model inputs, so that the model can summarize and generate a comprehensive personalized portrait through specific model fine-tuning or prompt word technology.
[0158] Among them, the pre-trained model is self-supervised learning based on a large amount of public data, and it has certain intelligent emergence capabilities, such as language understanding and logical reasoning. However, its effect on vertical domain tasks may not be good, so it is necessary to train the model specifically for the vertical domain based on the pre-trained model and combined with the vertical domain data. This training is model fine-tuning, for example, the LORA (Low-Rank Adaptation) method. In this embodiment, model fine-tuning refers to fine-tuning vertical domain models such as barrage generation, personalized portrait extraction, personalized portrait update, and personalized portrait synthesis.
[0159] Prompt word technology refers to a method that uses the design and optimization of natural language instructions (i.e., "prompt words") to guide an AI model (such as a large language model) to generate a desired response. In this embodiment, natural language instructions are used to guide the AI model to generate a desired comprehensive personalized profile based on the historical and real-time personalized profiles associated with the viewer in the current scene.
[0160] S602, when the content displayed on the display is the first category of content, screen the barrage that matches the comprehensive personalized portrait from the existing barrages of the content displayed on the display as the personalized barrage.
[0161] Typically, when users watch videos on video websites, they can actively input barrages. Therefore, during the display of content, the content displayed by the display may have a large number of barrage resources input by other users when watching. Then, when determining personalized barrages for the viewer in the current scene, personalized barrages can be directly filtered from the existing barrage resources based on personalized portraits. Correspondingly, for display content that does not have barrage resources or has fewer barrage resources, since personalized barrages cannot be directly filtered from existing barrage resources, it is necessary to generate personalized barrages based on personalized portraits. Based on this, the content displayed by the display can be classified according to the number of barrage resources. Among them, the first category of content is content with barrage resources that meet the quantity requirements, such as media content with barrage resource access provided by content providers. The second category of content is content that does not have barrage resources that meet the quantity requirements, such as live content, screen casting content, content provided by third-party APPs, and content provided by HDMI's own film sources, as well as new media content with barrage resource access but fewer barrage resources provided by content providers.
[0162] Among them, the first category of content is content with barrage resources that meet the quantity requirements.
[0163] As mentioned above, when the content displayed on the display is the first category of content, the barrage that matches the comprehensive personalized portrait can be screened from the existing barrages of the content displayed on the display as personalized barrages.
[0164] For example, a content label can be added to each existing barrage based on its content, so that a barrage with a content label that matches the comprehensive personalized portrait can be determined as a personalized barrage.
[0165] S603: When the content displayed on the display is the second type of content, a bullet screen that matches the comprehensive personalized portrait is generated as a personalized bullet screen.
[0166] Among them, the second category of content is content that does not have barrage resources that meet the quantity requirements.
[0167] Correspondingly, when the content displayed on the display is the second type of content, a barrage that matches the comprehensive personalized portrait is generated as a personalized barrage.
[0168] For example, through the AIGC method, based on the information of each dimension represented by the comprehensive personalized portrait, a barrage with content that conforms to the above comprehensive personalized portrait is generated as a personalized barrage.
[0169] In this embodiment, the content displayed on the display is classified according to the number of existing barrage resources, and personalized barrage resources are determined in different ways when the display displays different types of content. Thus, for the content displayed on the display that has no barrage resource access or has barrage resource access but fewer barrage resources, personalized barrages that match the comprehensive personalized portrait can be generated to achieve personalized barrage display for the viewing party in the current scene when there is no barrage source or fewer barrages. Furthermore, the use of comprehensive personalized portraits to determine personalized barrages can also reduce the error jitter caused by randomness in the real-time personalized portrait of the viewing party in the current scene, and can enrich the specificity of the real-time personalized portrait, thereby further improving the accuracy and pertinence of the displayed barrages.
[0170] It can be understood that the historical personalized portrait associated with the viewer in the current scene changes over time, and can be gradually formed based on the real-time personalized portrait that changes over time. That is to say, the real-time personalized portrait of the viewer in the current scene can contribute to the historical personalized portrait associated with the viewer in subsequent scenes. Therefore, the real-time personalized portrait of the viewer in the current scene can be used to construct a new historical personalized portrait associated with the viewer in the current scene, which can be used to generate a comprehensive personalized portrait in subsequent scenes.
[0171] Based on this, on the basis of the above embodiments, in an exemplary embodiment, the historical personalized portrait associated with the viewing party in the above current scene is the historical personalized portrait within the current portrait update cycle, such as Figure 7 As shown, the bullet screen display method may include the following steps:
[0172] S701, after the previous portrait update cycle of the current portrait update cycle ends, the real-time personalized portraits extracted in the preset historical period before the current portrait update cycle are subjected to portrait fusion to obtain the historical personalized portraits within the current portrait update cycle.
[0173] In order to obtain a comprehensive personalized portrait by integrating the historical personalized portrait and the real-time personalized portrait of the viewer in the current scene within the current portrait cycle, it is necessary to first determine the historical personalized portrait within the current cycle.
[0174] Based on this, after the previous portrait update cycle of the current portrait update cycle ends, the real-time personalized portrait extracted in the preset historical period before the current portrait update cycle can be obtained, and the obtained real-time personalized portraits can be fused to use the obtained fused personalized portrait as the historical personalized portrait within the current portrait update cycle.
[0175] Correspondingly, after the current portrait update cycle ends, the real-time personalized portraits extracted in the preset historical period before the next portrait update cycle of the current portrait update cycle can be fused to obtain the historical personalized portraits within the next portrait update cycle of the current portrait update cycle, so as to realize personalized barrage display for the viewer in the next portrait update cycle of the current portrait update cycle.
[0176] In an optional embodiment, the preset historical period may include multiple profile update cycles, serving as multiple historical profile update cycles. Thus, the real-time personalized profiles extracted in each of the historical profile update cycles prior to the current profile update cycle may be fused to obtain the historical personalized profiles for the current profile update cycle. The number of real-time personalized profiles extracted in each profile update cycle may be at least one.
[0177] Optionally, after the previous portrait update cycle of the current portrait update cycle ends, all real-time personalized portraits extracted in the historical portrait update cycles before the current portrait update cycle can be fused to obtain the historical personalized portraits in the current portrait update cycle.
[0178] Optionally, when there are multiple real-time personalized portraits extracted in each portrait update cycle, some of the real-time personalized portraits extracted in the historical portrait update cycles before the current portrait update cycle can be fused to obtain the historical personalized portraits in the current portrait update cycle. For example, some of the real-time personalized portraits extracted in the historical portrait update cycles are randomly selected for fusion to obtain the historical personalized portraits in the current portrait update cycle. For another example, based on the frequency of use of the display device in different time periods in each historical portrait update cycle, the real-time personalized portraits extracted in the time period with higher frequency of use are selected for fusion to obtain the historical personalized portraits in the current portrait update cycle, etc.
[0179] In an optional embodiment, during each portrait update cycle, the extracted real-time personalized portrait can be stored in a preset historical portrait database. Thus, after the portrait update cycle ends, the real-time personalized portrait stored in the preset historical period before the next portrait update cycle of the portrait update cycle is read from the historical portrait database, and the read real-time personalized portrait is subjected to portrait fusion to obtain a new historical personalized portrait, which serves as the historical personalized portrait for the next portrait update cycle of the portrait update cycle. For example, if each portrait update cycle is 1 day and the preset historical period is 30 days, the historical personalized portrait used each day is obtained by performing portrait fusion on the real-time personalized portraits extracted within 30 days before that day.
[0180] Optionally, for each portrait update cycle, real-time personalized portraits extracted within a preset historical period prior to the portrait update cycle are subjected to portrait fusion. When determining the historical personalized portrait for the portrait update cycle following the portrait update cycle, the real-time personalized portraits extracted within the preset historical period may be nonlinearly weighted based on a principle similar to an RLS (Recursive Least Squares Filter). The weight of the real-time personalized portrait extracted within the historical portrait update cycle closer to the portrait update cycle is higher. For example, if the preset historical period includes multiple historical portrait update cycles, first, for each historical portrait update cycle, information of the same dimension in each real-time personalized portrait extracted within the historical portrait update cycle is averagely weighted. Then, a weighted sum of information of the same dimension in each real-time personalized portrait extracted within the historical portrait update cycle is performed to obtain the corresponding fused personalized portrait for the historical portrait update cycle. Subsequently, a weight is assigned to each historical portrait update cycle based on the time from each historical portrait update cycle to the current portrait update cycle, with the weight of the historical portrait update cycle closer to the portrait update cycle being higher. In this way, the weighted sum of the information of the same dimension in the corresponding integrated personalized portraits within each historical portrait update cycle can be performed to obtain new information of the same dimension, which is used as the information of the same dimension in the current portrait update cycle, thereby obtaining the historical personalized portraits within the current portrait update cycle. The above weights can be set based on empirical values, test values from multiple experiments, and the needs of actual application scenarios, and are not specifically limited to this.
[0181] S702, when a bullet screen triggering event is detected during content display on the display, audio data and video data of the current scene of the display device collected by the audio and video collection device are received, and image data of the content currently displayed on the display is obtained.
[0182] S703: Generate scene description information of the current scene based on the audio data, video data and image data.
[0183] S704: Extracting a real-time personalized portrait of the viewer in the current scene from the scene description information.
[0184] The specific implementation of the above S702-S704 is the same as the specific implementation of the above S501-S503, and will not be repeated here.
[0185] S705: Fusing the historical personalized portrait and the real-time personalized portrait associated with the viewer to obtain a comprehensive personalized portrait.
[0186] S706: When the content displayed on the display is the first category of content, select the barrage that matches the comprehensive personalized portrait from the existing barrages of the content displayed on the display as the personalized barrage.
[0187] S707: When the content displayed on the display is the second type of content, a bullet screen that matches the comprehensive personalized portrait is generated as a personalized bullet screen.
[0188] The specific implementation of the above S705-S707 is the same as the specific implementation of the above S601-S603, and will not be repeated here.
[0189] S708, controlling the display to overlay and display the personalized bullet screen.
[0190] The specific implementation of the above S708 is the same as that of the above S505 and will not be repeated here.
[0191] In this embodiment, by performing portrait fusion on the real-time personalized portraits extracted in the preset historical period before the current portrait update cycle, the historical personalized portraits within the current portrait update cycle are obtained. While paying attention to long-term personalized tracking, it is possible to ensure that the personalization changes over time, thereby ensuring the matching of the historical personalized portraits with the characteristics of the viewers that change over time, avoiding the comprehensive personalized portrait errors caused by the historical personalized portraits being too fixed, improving the accuracy of the comprehensive personalized portraits, and further improving the accuracy and pertinence of the displayed barrage.
[0192] Typically, when a viewer in the current scene is viewing content displayed on a display, they often want to see bullet comments related to the content displayed on the display, such as evaluation information, introduction information, question answering information, etc. Therefore, when generating personalized bullet comments, in addition to considering the comprehensive personalized profile of the viewer in the current scene, the content displayed on the display can also be further considered.
[0193] Based on this, on the basis of the above embodiments, in an exemplary embodiment, as Figure 8 As shown, the bullet screen display method may include the following steps:
[0194] S801, when a bullet screen triggering event is detected during content display on the display, audio data and video data of the current scene of the display device collected by the audio and video collection device are received, and image data of the content currently displayed on the display is obtained.
[0195] S802: Generate scene description information of the current scene based on the audio data, video data, and image data.
[0196] S803: Extracting a real-time personalized portrait of the viewer in the current scene from the scene description information.
[0197] S804: Fusing the historical personalized portrait and the real-time personalized portrait associated with the viewer to obtain a comprehensive personalized portrait.
[0198] Among them, the specific implementation method of the above S801-S804 is the same as the specific implementation method of the above S702-S705, and will not be repeated here.
[0199] S805: Obtain content analysis information of the content displayed on the display.
[0200] The content analysis information includes information describing at least one of characters, scenes, and plots in the content displayed on the display.
[0201] During the process of displaying content on the display, the content displayed on the display may be analyzed to obtain content analysis information of the content displayed on the display.
[0202] The so-called parsing of the content displayed on the display may include performing a variety of parsing operations on the displayed content, such as behavior recognition, object and scene classification, and event monitoring, thereby obtaining character behavior, scene type, object type, and event information (such as event occurrence sequence, causal relationship, etc.) in the target content. Therefore, the content parsing information of the content displayed on the display includes information describing at least one of the characters, scenes, and plots in the content displayed on the display. The information describing the characters included in the above-mentioned content parsing information may include information such as the number of characters, gender of the characters, relationship between the characters, language of the characters, and behavior of the characters; the information describing the scenes included in the above-mentioned content parsing information may include information such as scene category, scene location, and scene name; and the information describing the plot included in the above-mentioned content parsing information may include information such as plot type, plot occurrence sequence, plot cause, and plot cause relationship.
[0203] Optionally, during the display's content display process, the displayed content can be cached according to a preset cache period, and after each cache period, the cached content can be parsed to obtain content parsing information. For example, when parsing the cached content, the content cached within the current preset cache period can be parsed, or all currently cached target content can be parsed, or the content cached within a preset duration with the current time as the end time can be parsed. All of these are reasonable.
[0204] In addition, as mentioned above, the content displayed by the display may be a variety of content such as live broadcast content, screen projection content, content provided by a third-party APP, content provided by HDMI's own film source, and content provided by a content provider. Depending on the source of the content displayed by the display, whether the complete content of the content displayed this time can be obtained during the display process may be different. For example, when the content displayed by the display is live broadcast content, screen projection content, content provided by a third-party APP, and content provided by HDMI's own film source, it may not be possible to obtain the complete content of the content displayed this time during the display process. However, when the content displayed by the display is content provided by a content provider, the complete content of the content displayed this time can be obtained during the display process. Therefore, depending on the source of the content displayed by the display, different methods can be used to obtain content parsing information of the content displayed by the display.
[0205] Based on this, in an optional embodiment, the acquisition of the content analysis information in the above S805 is further limited, and the above S805 may include the following steps:
[0206] 1. When the content displayed on the display is live content, the content displayed on the display is cached according to a preset cache period during the display display process, and after each cache is completed, the cached content is parsed to obtain content parsing information.
[0207] As mentioned above, for live broadcast content, it is impossible to obtain the complete content of this live broadcast during the display process, but only the content that has been displayed in this live broadcast can be obtained. Therefore, during the process of displaying content, the content displayed on the display can be cached according to a preset cache period (such as 10 seconds, etc.), and after each caching is completed, the cached content is parsed to obtain content parsing information.
[0208] Optionally, for display content such as screen projection content, content provided by third-party APPs, and content provided by HDMI's own video source, which cannot obtain the complete content of the displayed content during the display process, the content displayed on the display can be cached according to a preset cache period during the display display process, and after each cache is completed, the cached content can be parsed to obtain content parsing information.
[0209] Optionally, when parsing cached content, the content cached within the current preset cache period can be parsed, or all currently cached content can be parsed, or the content cached within a preset duration with the current time as the end time can be parsed. All of these are reasonable.
[0210] In addition, the above-mentioned preset cache period can be set according to the specific content displayed on the display, the estimated total display time, the viewing habits of the viewer, etc., and no specific limitation is made to this.
[0211] 2. When the content displayed on the display is on-demand content, content analysis information of the content displayed on the display is obtained from the server.
[0212] On-demand content refers to content provided by content providers for users to select and display. Typically, the complete content of the on-demand content is available before or during the display of the on-demand content. Examples include TV series, movies, and variety shows provided by various video websites.
[0213] Correspondingly, when the content displayed on the display is on-demand content, since the complete content of the on-demand content can be obtained during the display process, although the display has not yet displayed the complete content of the on-demand content, the complete content of the on-demand content can be directly obtained, and thus, the content parsing information of the content displayed on the display can be directly obtained from the server.
[0214] Optionally, the complete content of the on-demand content may be obtained from the server, and then the obtained complete content may be parsed to obtain content parsing information.
[0215] Optionally, the server may pre-store content parsing information of the complete content to which the on-demand content belongs, and the content parsing information of the complete content to which the on-demand content belongs may be directly obtained from the server without performing a parsing operation.
[0216] S806, when the content displayed on the display is the first category of content, select the barrage that matches both the content analysis information and the comprehensive personalized portrait from the existing barrages of the content displayed on the display as the personalized barrage.
[0217] After obtaining the above-mentioned content analysis information, as mentioned above, the first category of content is content with barrage resources that meet the quantity requirements. In the case that the content displayed on the display is the first category of content, you can directly filter out the barrages that match the content analysis information and the comprehensive personalized portrait from the existing barrages of the content displayed on the display as personalized barrages.
[0218] For example, a content label can be added to each existing barrage based on its content. Thus, a barrage whose content label matches the content analysis information and the comprehensive personalized portrait can be determined as a personalized barrage.
[0219] S807, when the content displayed on the display is the second type of content, generate a barrage that matches both the content analysis information and the comprehensive personalized portrait as a personalized barrage.
[0220] Correspondingly, the second category of content is content that does not have barrage resources that meet the quantity requirements. When the content displayed on the display is the second category of content, a barrage that matches both the content analysis information and the comprehensive personalized portrait can be generated as a personalized barrage.
[0221] For example, the content analysis information and the comprehensive personalized portrait can be fused to obtain fused information, and the fused information is used to describe the characteristics that need to be matched by the barrage to be displayed. Then, through the AIGC method, according to the barrage characteristics represented by the fused information, a barrage with content that conforms to the above-mentioned content analysis information and the comprehensive personalized portrait is generated as a personalized barrage.
[0222] S808, controlling the display to overlay and display the personalized bullet screen.
[0223] The specific implementation of the above S808 is the same as the specific implementation of the above S505, and will not be repeated here.
[0224] In this embodiment, by obtaining the content analysis information of the target content, the content analysis information can be involved in the determination of personalized barrages. Thus, while ensuring that the determined personalized barrages meet the comprehensive personalized portrait of the viewer in the current scene, the matching degree between the determined personalized barrages and the content displayed on the display can be improved. Thus, the comprehensive matching degree of the determined personalized barrages to the viewers and the displayed content in the current scene can be improved. While further improving the accuracy and pertinence of the displayed barrages, the fun of the displayed barrages and their appeal to viewers are improved, further improving the viewing experience of viewers.
[0225] It is understandable that for a scene with relatively fixed members, the characters in the current scene are usually fixed during the display of content. Therefore, the personalized portraits such as the real-time personalized portrait, historical personalized portrait, and comprehensive personalized portrait are determined for the fixed members of the scene. Therefore, the personalized portraits are determined and personalized bullet comments are displayed for the fixed members of the scene as a whole. For example, for a family, the family members are usually considered as a whole to determine the personalized portraits and then display personalized bullet comments.
[0226] However, in some cases, during the display of content, characters that appear occasionally may appear in the current scene where the display device is located, or not all fixed members may exist in the current scene where the display device is located. In order to avoid the personalized portrait of the occasionally appearing character from affecting the real-time personalized portrait of the fixed members in the scene, and further to avoid affecting the historical personalized portrait of the fixed members in the scene, which may lead to deviations in the subsequent personalized barrage display for the fixed members in the scene, thereby reducing accuracy and pertinence; and in order to improve the accuracy and pertinence of the personalized barrage for the fixed members in the current scene, characters that are fixed members of the scene can be identified among the various characters in the current scene, so as to determine the personalized portrait for the fixed member and display the personalized barrage.
[0227] Based on this, on the basis of the above embodiments, in an exemplary embodiment, as Figure 9 As shown, the bullet screen display method may include the following steps:
[0228] S901, when a bullet screen triggering event is detected during content display on the display, audio data and video data of the current scene of the display device collected by the audio and video collection device are received, and image data of the content currently displayed on the display is obtained.
[0229] S902: Generate scene description information of the current scene based on the audio data, video data, and image data.
[0230] Among them, the specific implementation method of the above S901-S902 is the same as the specific implementation method of the above S501-S502, and will not be repeated here.
[0231] S903: Determine biological characteristics of people in the current scene based on the audio data and / or video data.
[0232] Biometrics are physiological attributes used to uniquely identify a person. Examples include voiceprint information and facial image information. Voiceprint recognition on audio data can be used to obtain the voiceprint information of each person in the current scene, while facial image recognition on video data can be used to obtain the facial image information of each person in the current scene.
[0233] Optionally, the voiceprint information of each person in the current scene may be determined based on the audio data as the biometric characteristics of each person in the current scene.
[0234] Optionally, facial image information of each person in the current scene may be determined based on the video data as the biometric features of each person in the current scene.
[0235] Optionally, the voiceprint information of each person in the current scene can be determined based on the audio data, and the facial image information of each person in the current scene can be determined based on the video data, and then the above voiceprint information and facial image information can be used as the biometric characteristics of each person in the current scene.
[0236] S904: Match the biological characteristics with the biological characteristics of a preset person to obtain a target person in the current scene who is a preset person.
[0237] Based on the members of the scene in which the display device is located, preset characters that can serve as viewers can be pre-set. For example, for a family, family members can be set as preset characters, and the biometric characteristics of the preset characters can be pre-collected and stored. Thus, the biometric characteristics of each person in the current scene can be matched with the biometric characteristics of the preset characters to obtain the target person in the current scene that is a preset character.
[0238] S905: Determine the target person or a person among the target persons who matches the content displayed on the display as the viewing party.
[0239] Optionally, after the target person is determined, the target person can be directly determined as the viewer.
[0240] Optionally, after determining the target person, based on the specific information of the displayed content, the target person can be screened for people who are currently paying attention to the displayed content or people who pay more attention to the content displayed on the display. The screened person will pay more attention to whether the barrage displayed on the display matches his or her own preferences and needs. Thus, personalized barrage display can be performed for the screened person, and the target person who matches the content displayed on the display can be determined as the viewer.
[0241] For example, the determined target persons include person A and person B, and person A usually pays more attention to variety shows, while person B usually pays more attention to news content. If the content displayed on the monitor is a variety show, person A can be determined as the viewer.
[0242] S906: Extracting a real-time personalized portrait of the viewer in the current scene from the scene description information.
[0243] S907, determining personalized bullet comments based on the real-time personalized portrait of the viewer in the current scene.
[0244] S908, controlling the display to overlay and display the personalized bullet screen.
[0245] Among them, the specific implementation method of the above S906-S908 is the same as the specific implementation method of the above S503-S505, and will not be repeated here.
[0246] Optionally, in this embodiment, there are historical personalized portraits for all preset personnel in the current portrait update cycle, and there are historical personalized portraits for each preset character. Then, after the previous portrait update cycle of the current portrait update cycle ends, the real-time personalized portraits extracted in the preset historical period before the current portrait update cycle can be fused to obtain the historical personalized portraits of all preset personnel in the current portrait update cycle, and the historical personalized portraits of each preset character in the current portrait update cycle.
[0247] In this embodiment, personalized bullet screen display can be performed for specific characters in the current scene, thereby further refining the granularity of personalized bullet screen display and improving the accuracy and pertinence of the displayed bullet screen display.
[0248] Based on the above embodiments, in an exemplary embodiment, Figure 10 As shown, the bullet screen display method may include the following steps:
[0249] S1001, after the previous portrait update cycle of the current portrait update cycle ends, the controller performs portrait fusion on the real-time personalized portraits extracted in the preset historical period before the current portrait update cycle to obtain the historical personalized portraits in the current portrait update cycle.
[0250] S1002: The audio and video acquisition device acquires audio data and video data of the current scene where the display device is located.
[0251] S1003: When the display is displaying content, the image acquisition device acquires image data of the content currently displayed on the display.
[0252] S1004: When a bullet screen triggering event is detected, the controller sends a data acquisition request to the audio acquisition device and the image acquisition device respectively;
[0253] S1005: The audio and video acquisition device sends audio data and video data to the controller.
[0254] S1006: The image acquisition device acquires image data and sends it to the controller.
[0255] S1007: The controller generates scene description information of the current scene based on the audio data, video data, and image data.
[0256] S1008: The controller extracts a real-time personalized portrait of the viewer in the current scene from the scene description information.
[0257] S1009: The controller fuses the historical personalized portrait and the real-time personalized portrait associated with the viewer to obtain a comprehensive personalized portrait.
[0258] S1010: When the content displayed on the display is live content, the controller caches the content displayed on the display according to a preset cache period, and parses the cached content after each caching to obtain content parsing information.
[0259] S1011 : When the content displayed on the display is on-demand content, the controller obtains content analysis information of the content displayed on the display from the server.
[0260] S1012, when the content displayed on the display is the first category of content, the controller selects the barrage that matches both the content analysis information and the comprehensive personalized portrait from the existing barrages of the content displayed on the display as the personalized barrage.
[0261] S1013, when the content displayed on the display is the second type of content, the controller generates a barrage that matches both the content analysis information and the comprehensive personalized portrait as a personalized barrage.
[0262] S1014, the controller sends personalized bullet comments to the display.
[0263] S1015, the display overlays and displays the personalized bullet screen.
[0264] Among them, the specific implementation method of the above S1001-S1015 is the same as the specific implementation method in the above embodiments, and will not be repeated here.
[0265] Based on the same inventive concept, some embodiments further provide a bullet screen display device for implementing the aforementioned bullet screen display method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the bullet screen display device provided below can be found in the limitations of the bullet screen display method described above and will not be repeated here.
[0266] In an exemplary embodiment, Figure 11 As shown, a barrage display device is provided, which is applied to a display device, including a data acquisition module 1110, an information generation module 1120, a portrait extraction module 1130, a barrage determination module 1140 and a barrage display module 1150.
[0267] The data acquisition module 1110 is used to receive audio data and video data of the current scene of the display device collected by the audio and video acquisition device when a bullet screen triggering event is detected during the process of displaying content on the display of the display device, and to obtain image data of the content currently displayed on the display;
[0268] Information generation module 1120, configured to generate scene description information of the current scene based on the audio data, video data, and image data; wherein the scene description information is used to describe the environmental characteristics and character characteristics of the current scene; the character characteristics include at least one of the content viewed by the character, the character's dialogue, and the character's basic attributes;
[0269] The portrait extraction module 1130 is used to extract a real-time personalized portrait of the viewer in the current scene from the scene description information;
[0270] A bullet comment determination module 1140 is configured to determine a personalized bullet comment based on a real-time personalized portrait of the viewer in the current scene;
[0271] The bullet screen display module 1150 is used to control the display to overlay and display personalized bullet screens.
[0272] In an exemplary embodiment, the barrage determination module 1140 includes: a portrait fusion unit, which is used to perform portrait fusion on the historical personalized portrait and the real-time personalized portrait associated with the viewer to obtain a comprehensive personalized portrait; wherein the historical personalized portrait includes information describing the characteristics of the content viewed by the viewer within a preset historical period; a first determination unit, which is used to, when the content displayed on the display is the first category of content, filter the barrage that matches the comprehensive personalized portrait from the existing barrage of the content displayed on the display as a personalized barrage; wherein the first category of content is content with barrage resources that meet the quantity requirements; a second determination unit, which is used to, when the content displayed on the display is the second category of content, generate the barrage that matches the comprehensive personalized portrait as a personalized barrage; wherein the second category of content is content that does not have barrage resources that meet the quantity requirements.
[0273] In an exemplary embodiment, the historical personalized portrait associated with the viewing party is a historical personalized portrait within the current portrait update cycle. The barrage display device also includes: a portrait update module, which is used to perform portrait fusion on the real-time personalized portrait extracted in the preset historical period before the current portrait update cycle after the previous portrait update cycle of the current portrait update cycle ends, to obtain the historical personalized portrait within the current portrait update cycle.
[0274] In an exemplary embodiment, the barrage display device also includes: an information determination module, which is used to determine the biometric characteristics of the person in the current scene based on audio data and / or video data before extracting the real-time personalized portrait of the viewer in the current scene from the scene description information; wherein the biometric characteristics are physiological attribute characteristics used to uniquely identify a person; an information matching module, which is used to match the biometric characteristics with the biometric characteristics of a preset person to obtain a target person belonging to the preset person in the current scene; and a person determination module is used to determine the target person or a person among the target persons who matches the content displayed on the display as the viewer.
[0275] In an exemplary embodiment, the barrage display device also includes: an information acquisition module, used to obtain content analysis information of the content displayed on the display; wherein the content analysis information includes information describing at least one of the characters, scenes and plots in the content displayed on the display; a first determination unit, specifically used to filter out barrages that match both the content analysis information and the comprehensive personalized portrait from the existing barrages of the content displayed on the display, as personalized barrages; a second determination unit, specifically used to generate barrages that match both the content analysis information and the comprehensive personalized portrait, as personalized barrages.
[0276] In an exemplary embodiment, the information determination module is specifically used to: when the content displayed on the display is live content, during the process of the display displaying the content, cache the content displayed on the display according to a preset cache period, and after each caching is completed, parse the cached target content to obtain content resolution information; when the content displayed on the display is on-demand content, obtain content resolution information of the content displayed on the display from the server.
[0277] In an exemplary embodiment, the information generation module 1120 includes: an information generation module, specifically used to input audio data, video data and image data into the multimodal large model to obtain scene description information of the current scene.
[0278] In an exemplary embodiment, the multimodal large model includes a modal alignment module and a semantic recognition module; the information generation module is specifically used to: output the audio data, video data and image data to the modal alignment module for modal alignment, and obtain the data features of the audio data, video data and image data in the same semantic space; input the data features into the semantic recognition module for semantic recognition, and obtain the scene description information of the current scene.
[0279] In an exemplary embodiment, the portrait extraction module 1130 is specifically used to: extract the historical operation data of the viewer on the display device in the current scene from the local storage data of the display device; and extract the real-time personalized portrait of the viewer in the current scene from the scene description information and the historical operation data.
[0280] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0281] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0282] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0283] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile memory and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a programmable logic unit (PLC), a data processing logic unit based on quantum computing, an artificial intelligence (AI) processor, and the like.
[0284] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0285] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A display device, characterized in that: include: monitor; a communication device configured to connect to the audio and video acquisition device; and a controller connected to the communication device and the display, configured to: When a bullet screen triggering event is detected during content display on the display, receiving audio data and video data of the current scene of the display device collected by the audio and video collection device, and obtaining image data of the content currently displayed on the display; Generate scene description information of the current scene based on the audio data, the video data, and the image data; wherein the scene description information is used to describe environmental characteristics and character characteristics of the current scene; the character characteristics include at least one of the content viewed by the character, the conversations between the character, and basic attributes of the character; Extracting a real-time personalized portrait of the viewing party in the current scene from the scene description information; Determining personalized bullet comments based on the real-time personalized portrait of the viewer in the current scene; Control the display to overlay and display the personalized bullet screen.
2. The display device according to claim 1, wherein When the controller determines the personalized bullet comment based on the real-time personalized portrait of the viewer in the current scene, the controller is configured to: Performing a portrait fusion on the historical personalized portrait associated with the viewer and the real-time personalized portrait to obtain a comprehensive personalized portrait; wherein the historical personalized portrait includes information describing characteristics of the content viewed by the viewer within a preset historical period; When the content displayed on the display is the first category of content, screening the barrage that matches the comprehensive personalized portrait from the existing barrages of the content displayed on the display as the personalized barrage; wherein the first category of content is content with barrage resources that meet the quantity requirement; In the case where the content displayed on the display is the second type of content, a barrage matching the comprehensive personalized portrait is generated as a personalized barrage; wherein the second type of content is content that does not have barrage resources that meet the quantity requirements.
3. The display device according to claim 2, wherein The historical personalized portrait associated with the viewing party is a historical personalized portrait within the current portrait update period; the controller is further configured to: After the previous portrait update cycle of the current portrait update cycle ends, portrait fusion is performed on the real-time personalized portraits extracted in the preset historical period before the current portrait update cycle to obtain the historical personalized portraits within the current portrait update cycle.
4. The display device according to any one of claims 1 to 3, characterized in that: Before the controller extracts the real-time personalized portrait of the viewer in the current scene from the scene description information, the controller is further configured to: Determining a biometric characteristic of a person in the current scene based on the audio data and / or the video data; wherein the biometric characteristic is a physiological attribute characteristic used to uniquely identify a person; Matching the biological characteristics with the biological characteristics of a preset person to obtain a target person in the current scene who is the preset person; The target person or a person among the target persons who matches the content displayed on the display is determined as the viewing party.
5. The display device according to claim 2, wherein: The controller is further configured to: Obtaining content analysis information of the content displayed on the display; wherein the content analysis information includes information describing at least one of a character, a scene, and a plot in the content displayed on the display; The controller is configured to: Filtering, from existing bullet comments of the content displayed on the display, bullet comments that match both the content analysis information and the comprehensive personalized portrait as personalized bullet comments; The controller is configured to generate a bullet comment matching the comprehensive personalized portrait as a personalized bullet comment: Generate a barrage that matches both the content analysis information and the comprehensive personalized portrait as a personalized barrage.
6. The display device according to claim 5, wherein: When the controller executes the process of acquiring content analysis information of the content displayed on the display, the controller is configured to: In the case where the content displayed on the display is live content, during the process of the display displaying the content, the content displayed on the display is cached according to a preset cache period, and after each caching is completed, the cached content is parsed to obtain content parsing information; In a case where the content displayed on the display is on-demand content, content parsing information of the content displayed on the display is obtained from a server.
7. The display device according to any one of claims 1 to 6, characterized in that: When the controller generates scene description information of the current scene according to the audio data, the video data, and the image data, the controller is configured to: The audio data, the video data and the image data are input into a multimodal large model to obtain scene description information of the current scene.
8. The display device according to claim 7, wherein: The multimodal large model includes a modality alignment module and a semantic recognition module; when the controller inputs the audio data, the video data, and the image data into the multimodal large model to obtain scene description information of the current scene, the controller is configured to: Outputting the audio data, the video data, and the image data to the modality alignment module for modality alignment to obtain data features of the audio data, the video data, and the image data in the same semantic space; The data features are input into the semantic recognition module for semantic recognition to obtain scene description information of the current scene.
9. The display device according to any one of claims 1 to 8, characterized in that: When the controller extracts the real-time personalized portrait of the viewing party in the current scene from the scene description information, the controller is configured to: Extracting historical operation data of the viewing party on the display device in the current scene from local storage data of the display device; A real-time personalized portrait of the viewer in the current scene is extracted from the scene description information and the historical operation data.
10. A bullet screen display method, characterized in that: Applied to a display device, the method includes: When a bullet screen triggering event is detected during content display on the display device, receiving audio data and video data of the current scene of the display device collected by an audio and video collection device, and obtaining image data of the content currently displayed on the display; Generate scene description information of the current scene based on the audio data, the video data, and the image data; wherein the scene description information is used to describe environmental characteristics and character characteristics of the current scene; the character characteristics include at least one of the content viewed by the character, the conversations between the character, and basic attributes of the character; Extracting a real-time personalized portrait of the viewing party in the current scene from the scene description information; Determining personalized bullet comments based on the real-time personalized portrait of the viewer in the current scene; Control the display to overlay and display the personalized bullet screen.
Citation Information
Patent Citations
Bullet screen information processing method, device and equipment
CN110166811A
Interactive video generation method and device based on artificial intelligence, equipment and medium
CN110708595A
Bullet screen editing method, intelligent terminal and storage medium
CN110740387A
Bullet screen display method and device, electronic equipment and medium
CN114245222A
Bullet screen generation method and device, electronic equipment and readable storage medium
CN117499733A