Display device and method for playing media content with fused virtual character
By detecting the frame data of media assets and generating feature scaling and positioning parameters, virtual character models are rendered for integrated display, solving the problem of attention distraction caused by the large distance between the floating window and the video screen in the display device, and improving the display effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HISENSE VISUAL TECH CO LTD
- Filing Date
- 2023-05-08
- Publication Date
- 2026-05-12
AI Technical Summary
When a display device simultaneously shows footage captured by a camera while playing video, the floating window is too far from the video screen, causing the user's attention to be distracted and making it impossible to focus on both screens at the same time, thus affecting the display effect.
By detecting the frames of the media asset data, extracting and adjusting feature data to generate feature scaling and positioning parameters, rendering the virtual character model, and making its display layer higher than the media asset data, the virtual character is integrated with the playback screen.
It improves the display effect of virtual characters, allowing users to simultaneously view video footage and camera feeds, thus enhancing user interaction and the overall display effect of the display device.
Smart Images

Figure CN116506678B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of display device technology, and in particular to a display device and a method for playing media assets that integrate virtual characters. Background Technology
[0002] Display devices refer to terminal devices capable of outputting specific display images, such as smart TVs, communication terminals, smart advertising screens, and projectors. Taking smart TVs as an example, smart TVs are television products based on Internet application technologies, possessing open operating systems and chips, and having open application platforms. They enable two-way human-computer interaction and integrate multiple functions such as audio-visual, entertainment, and data to meet diverse and personalized user needs.
[0003] The display device can also simultaneously display footage captured by a connected camera while playing video. For example, when displaying an action video, the device can also show the user's real-time tracking footage captured by the camera within the user interface, allowing the user to compare their own movements with the video's. Therefore, the display device can display one of these frames as a floating window within the user interface.
[0004] However, to avoid interfering with the user's video viewing experience, the floating window needs to be displayed in a position that doesn't obstruct the video feed, such as the top left or top right corner. This results in a significant distance between the video feed and the camera feed, easily distracting the user and preventing them from simultaneously focusing on both feeds. In other words, the display device cannot simultaneously display both the video feed and the camera feed, leading to poor display quality. Summary of the Invention
[0005] This application provides a display device and a method for playing media assets that integrate virtual characters, in order to solve the problem of poor display effect of the display device.
[0006] In a first aspect, some embodiments of this application provide a display device, including a display, an image interface, and a controller. The display is configured to display a user interface; the image interface is configured to acquire a user image, the user image including reference feature data; and the controller is configured to execute the following program steps:
[0007] In response to a playback command for media asset data, detect the frame of the media asset data;
[0008] Extract the adjustment feature data of the aforementioned frame;
[0009] Based on the adjusted feature data, feature scaling parameters and feature localization parameters are generated;
[0010] The baseline feature model is scaled according to the feature scaling parameters to generate the target feature model, wherein the baseline feature model is a virtual character model generated based on the baseline feature data;
[0011] The target feature model is rendered according to the feature positioning parameters, and the display level of the target feature model is higher than the display level of the playback screen corresponding to the media asset data.
[0012] Secondly, some embodiments of this application also provide a method for playing media assets incorporating virtual characters, including:
[0013] In response to a playback command for media asset data, detect the frame of the media asset data;
[0014] Extract the adjustment feature data of the aforementioned frame;
[0015] Based on the adjusted feature data, feature scaling parameters and feature localization parameters are generated;
[0016] The baseline feature model is scaled according to the feature scaling parameters to generate the target feature model, wherein the baseline feature model is a virtual character model generated based on the baseline feature data;
[0017] The target feature model is rendered according to the feature positioning parameters, and the display level of the target feature model is higher than the display level of the playback screen corresponding to the media asset data.
[0018] As can be seen from the above technical solutions, the display device and media asset playback method integrating virtual characters provided in some embodiments of this application can respond to the playback command of media asset data and detect the frame of the media asset data. Then, adjustment feature data of the frame is extracted, and feature scaling parameters and feature positioning parameters are generated based on the adjustment feature data. A reference feature model is scaled according to the feature scaling parameters to generate a target feature model. The reference feature model is a virtual character model generated based on the reference feature data. The target feature model is then rendered according to the feature positioning parameters, wherein the display level of the target feature model is higher than the display level of the corresponding playback screen of the media asset data. This method can display the virtual character and the displayed character in the playback screen at the same height and bottom, thus integrating the virtual character with the playback screen of the media asset data and improving the display effect of the virtual character. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application;
[0021] Figure 2 This is a schematic diagram of the hardware configuration of a display device provided in some embodiments of this application;
[0022] Figure 3 This is a schematic diagram of the hardware configuration of the control device provided in some embodiments of this application;
[0023] Figure 4 This is a schematic diagram of the software configuration of a display device provided in some embodiments of this application;
[0024] Figure 5 This application provides some embodiments of the interaction diagrams illustrating the display device obtaining media asset data from the server.
[0025] Figure 6 This is a schematic diagram illustrating the process of decoding media asset data provided in some embodiments of this application;
[0026] Figure 7 A schematic diagram illustrating the process of rendering a target feature person according to some embodiments of this application;
[0027] Figure 8 A flowchart illustrating a method for playing media assets incorporating virtual characters, provided in some embodiments of this application;
[0028] Figure 9 A schematic diagram illustrating the effect of a baseline feature model provided in some embodiments of this application;
[0029] Figure 10 A schematic diagram illustrating the process of generating feature scaling parameters based on adjusting feature data, provided for some embodiments of this application;
[0030] Figure 11 A flowchart illustrating the process of generating feature localization parameters based on adjusting feature data, provided for some embodiments of this application;
[0031] Figure 12 A schematic diagram illustrating the rendering effect of the target feature model provided in some embodiments of this application;
[0032] Figure 13 A schematic diagram illustrating the hierarchical effect of the playback window and the target feature model provided in some embodiments of this application;
[0033] Figure 14 This is a schematic diagram illustrating the effect of fusing target feature models and media asset data in some embodiments of this application. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of the exemplary embodiments of this application clearer, the technical solutions in the exemplary embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described exemplary embodiments are only some embodiments of this application, and not all embodiments.
[0035] Based on the exemplary embodiments shown in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application. Furthermore, although the disclosures in this application are presented by way of one or more exemplary examples, it should be understood that each aspect of these disclosures can constitute a complete technical solution on its own.
[0036] It should be understood that the terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate, for example, to allow implementation in orders other than those given in the embodiments illustrated or described in this application.
[0037] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclusively include, for example, a product or device that includes a series of components is not necessarily limited to those that are explicitly listed, but may include other components that are not explicitly listed or that are inherent to such product or device.
[0038] The display device provided in this application can have various implementation forms, such as a television, a smart television, a laser projection device, a monitor, an electronic bulletin board, an electronic table, etc. Figure 1 and Figure 2 This is one specific embodiment of the display device of this application.
[0039] Figure 1 This is a schematic diagram illustrating the operational scenario between the display device and the control unit according to the embodiment. Figure 1 As shown, the user can operate the display device 200 through the smart device 300 or the control device 100.
[0040] In some embodiments, the control device 100 may be a remote control. Communication between the remote control and the display device includes infrared protocol communication, Bluetooth protocol communication, and other short-range communication methods, controlling the display device 200 wirelessly or via wired means. Users can control the display device 200 by inputting user commands through buttons on the remote control, voice input, control panel input, etc.
[0041] In some embodiments, a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) may also be used to control the display device 200. For example, an application running on the smart device may be used to control the display device 200.
[0042] In some embodiments, the display device may receive instructions not through the aforementioned smart devices or control devices, but through touch or gestures.
[0043] In some embodiments, the display device 200 can also be controlled in ways other than the control device 100 and the smart device 300. For example, it can be controlled by directly receiving the user's voice commands through a module configured inside the display device 200 for acquiring voice commands, or it can be controlled by receiving the user's voice commands through a voice control device set outside the display device 200.
[0044] In some embodiments, the display device 200 also communicates with the server 400. The display device 200 may communicate via a local area network (LAN), wireless local area network (WLAN), and other networks. The server 400 may provide various content and interactive features to the display device 200. The server 400 may be a cluster or multiple clusters, and may include one or more types of servers.
[0045] Figure 2 An exemplary block diagram of the configuration of the control device 100 according to an exemplary embodiment is shown. Figure 2 As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 can receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.
[0046] like Figure 3 The display device 200 includes at least one of the following: a tuner 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.
[0047] In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, RAM, ROM, and a first interface to an nth interface for input / output.
[0048] The display 260 includes a display screen assembly for presenting images, a driving assembly for driving image display, a component for receiving image signals from the controller output, and a user control UI interface for displaying video content, image content, menu control interface, and user control UI interface.
[0049] The display 260 can be an LCD display, an OLED display, or a projection display, and can also be a projection device and a projection screen.
[0050] The communicator 220 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of the following: a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The display device 200 can establish the transmission and reception of control signals and data signals with the control device 100 or the server 400 through the communicator 220.
[0051] The user interface can be used to receive control signals from the control device 100 (such as an infrared remote control).
[0052] Detector 230 is used to collect signals from the external environment or to interact with the external environment. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.
[0053] The external device interface 240 may include, but is not limited to, one or more of the following: High Definition Multimedia Interface (HDMI), analog or high-definition component input interface (component), composite video input interface (CVBS), USB input interface (USB), RGB port, etc. It may also be a composite input / output interface formed by multiple interfaces mentioned above.
[0054] The tuner / demodulator 210 receives broadcast television signals via wired or wireless means, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals.
[0055] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0056] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200. For example, in response to receiving a user command to select a UI object to display on the monitor 260, the controller 250 can execute operations related to the object selected by the user command.
[0057] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), a RAM (Random Access Memory), a ROM (Read-Only Memory), a first to an nth interface for input / output, a communication bus, etc.
[0058] Users can input commands through a graphical user interface (GUI) displayed on the monitor 260, and the user input interface receives the user input commands through the GUI. Alternatively, users can input commands by entering specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.
[0059] A "user interface" is the medium through which an application or operating system interacts and exchanges information with the user. It converts information from its internal form to a form that the user can accept. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.
[0060] like Figure 4 In some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the Android runtime and system library layer (referred to as the "System Runtime Layer"), and the kernel layer.
[0061] In some embodiments, at least one application runs in the application layer. These applications may be Windows programs, system settings programs, or clock programs that come with the operating system; they may also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the examples above.
[0062] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.
[0063] like Figure 4 As shown, the application framework layer in this embodiment includes managers, content providers, etc., wherein the managers include at least one of the following modules: Activity Manager, which interacts with all activities running in the system; Location Manager, which provides access to system location services for system services or applications; Package Manager, which retrieves various information related to application packages currently installed on the device; Notification Manager, which controls the display and clearing of notification messages; and Window Manager, which manages icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0064] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling display window changes (e.g., shrinking the display window, shaking the display, distorting the display, etc.).
[0065] In some embodiments, the system runtime library layer provides support for the upper layer, namely the framework layer. When the framework layer is used, the Android operating system runs the C / C++ libraries contained in the system runtime library layer to implement the functions that the framework layer needs to perform.
[0066] In some embodiments, the kernel layer is a layer between hardware and software. For example... Figure 4As shown, the kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver.
[0067] Based on the aforementioned display device 200, the display device 200 can receive control commands from the user and perform data processing according to the control commands to present user interfaces with different content through the display 260. In some embodiments, the control commands can be input by the user through a control device 100, such as a remote control associated with the display device 200; or, the control commands can also be input by the user through voice or touch operation.
[0068] In some embodiments, the display device 200 can communicate with the server based on user-input control commands to achieve media asset data interaction. The user can send control commands to the display device 200 to launch the media asset playback application. After the media asset playback application is launched, the display device 200 displays the application interface. The application interface may include multiple media asset items, each corresponding to a network link, where the network link is the Uniform Resource Locator (URL) for the media asset data corresponding to that item.
[0069] like Figure 5 As shown, in some embodiments, during communication between the display device 200 and the server 400, the display device 200 may send a request to the server 400, including a network link, to obtain media asset data corresponding to the media asset item. Therefore, after receiving the request from the display device 200, the server 400 extracts the corresponding media asset data according to the request and returns the extracted media asset data to the display device 200.
[0070] In order to enable data interaction between display device 200 and server 400, display device 200 needs to establish a communication connection with server 400. For example, both display device 200 and server 400 can be connected to the Internet and transmit interactive data between them through Internet Protocol (IP).
[0071] It should be noted that the display device 200 and the server 400 can also establish a communication connection using other methods. These include wired broadband, wireless LAN, cellular network, Bluetooth, infrared, and radio frequency communication.
[0072] The connection between display devices 200 and server 400 can be "many-to-one," meaning multiple display devices 200 can establish communication connections with the same server 400, allowing server 400 to provide services to multiple display devices 200. Alternatively, the connection can be "many-to-many," meaning multiple display devices 200 can establish communication connections with multiple servers 400, allowing multiple servers 400 to provide different services to each display device 200. Clearly, in certain application scenarios, the connection between display devices 200 and server 400 can also be "one-to-one," where one server 400 provides services exclusively to one display device 200.
[0073] Display device 200 can acquire media asset data from server 400 in real time and continuously generate playback images of the media asset data through decoding and other processing. In some embodiments, when the playback of media asset data is interrupted, display device 200 can automatically record the playback progress of the media asset data to generate a history record. The history record allows the user to directly resume playback of the media asset data according to the recorded playback progress when playing it again, thus meeting the user's need to continue watching.
[0074] Furthermore, in some embodiments, the display device 200 also monitors the status of media asset items. When a media asset item is selected, a playback command for the corresponding media asset data is generated. In response to the playback command, the display device 200 sends a media asset data retrieval request to the server to load the media asset data of the media asset item.
[0075] like Figure 6 As shown, after obtaining the media asset data, the display device 200 processes the media asset data to convert it into a playback screen that can be viewed by the user. Specifically, in some embodiments, the display device 200 also creates a playback window for the media asset data in response to a playback command. Then, it decodes the media asset data to obtain the playback screen. The playback screen consists of frames arranged in chronological order. When displaying the playback screen, the display device 200 plays the frames one by one in chronological order. After decoding the media asset data, the display device 200 controls the monitor 260 to display the playback screen through the playback window.
[0076] For example, a user launches a fitness application on display device 200 and selects media asset A in the application interface. Display device 200 then generates a playback command for media asset A and sends a media asset data retrieval request to server 400 to obtain the media asset data of media asset A. Simultaneously, display device 200 also responds to the playback command for media asset A by creating a media player and playback window for media asset A. The media player can decode the data stream of the media asset data and then display the playback screen of media asset A through the playback window.
[0077] In some embodiments, the application layer of the display device 200 may include various types of applications, such as fitness applications, media playback applications, and camera applications. Users can obtain different types of media data by launching different types of applications on the display device 200. Furthermore, launching different types of applications can enable more functions. For example, while playing media data, the display device 200 can also simultaneously launch the image interface 241 to capture user images in real time.
[0078] For example, the application layer of display device 200 includes a fitness application. After a user launches the fitness application on display device 200, the display device 200 displays the homepage interface of the fitness application. The homepage interface of the fitness application includes multiple media asset items such as aerobics videos and yoga videos. After the user selects a media asset item, display device 200 sends a request to server 400 to load the corresponding media asset data. When the user exercises through the fitness application on display device 200, the fitness application can also activate the image acquisition device connected to image interface 241, and use the image acquisition device to acquire the user's image in real time, and then control the display 260 to display the user's image. In this way, the user can compare their own exercise status with the exercise video in real time, enhancing the user's interactive experience.
[0079] To simultaneously display media asset data and user images, in some embodiments, the display device 200 creates a floating window, and displays the playback screen of one of the media asset data or user images on the other screen, so that two playback screens can be displayed simultaneously on the display 260. Furthermore, to avoid affecting the user's viewing of the content under the floating window, the floating window can be positioned in a location that does not obstruct the underlying screen, such as the upper left or upper right corner.
[0080] However, because there is a significant distance between the floating window positioned at the edge and the underlying window, the user's attention is easily distracted when viewing both windows simultaneously, making it impossible to focus on both content and reducing the user's interactive experience. Conversely, placing the two playback windows side-by-side at the same horizontal level to reduce the distance between them would compress the window size, degrading the display quality.
[0081] Based on the above application scenarios, in order to improve the display effect when the display device 200 simultaneously displays media asset data and user images, some embodiments of this application provide a display device 200, such as... Figure 7 As shown, the device includes a display 260, an image interface 241, and a controller 250. The display 260 is configured to display a user interface; the image interface 241 is configured to capture user images so that the display device 200 can generate display screens or virtual models based on the user images.
[0082] In some embodiments, the image interface 241 can be connected to an external image acquisition device or a built-in image acquisition module of the display device 200 to acquire user images, including user limbs, in real time via the image acquisition device or the image acquisition module. For example, both the image acquisition device and the image acquisition module can be a camera.
[0083] In some embodiments, when the image acquisition device or image acquisition module is a camera, it can be uniformly controlled by the controller 250 in the display device 200, and the acquired user images can be directly sent to the controller 250. Obviously, to facilitate the acquisition of user images, the camera should be positioned at a specific location on the display device 200. For example, for a display device 200 such as a smart TV, the camera can be positioned on the top of the smart TV, and the camera's shooting direction should be the same as the direction of light emitted from the smart TV screen, thereby enabling real-time capture of user images in front of the smart TV screen.
[0084] like Figure 8 As shown, the controller 250 is configured to perform the following program steps:
[0085] S100: In response to playback commands for media asset data, detects the frame of the media asset data.
[0086] After receiving the playback command for the media asset data, the display device 200 also detects the frame images in the media asset data to identify the displayed content within each frame. To improve the display effect when simultaneously displaying user images and media asset data, the display device 200 can merge the user image with the playback frame of the media asset data. Therefore, to match the displayed user image with the playback frame of the media asset data, the display device 200 can detect the frame images of the media asset data and identify the displayed figures within each frame to determine their height and position.
[0087] Since the user image also includes redundant background content, in some embodiments, the display device 200 can also generate a virtual character model based on the user image acquired by the image interface 241, and control the display 260 to render the virtual character model in the user interface. Then, based on the body movements in the user image, the virtual character is driven to move synchronously, so as to present the user's motion state in the user interface through the virtual character model. The virtual character model allows for real-time presentation of the user's body movements on the display device while reducing interference from redundant background in the user image.
[0088] Furthermore, to associate the virtual character model with the user's limbs, the display device 200 can also draw the virtual character model based on the user's limb joints, serving as a reference feature model for the display device 200. Therefore, in some embodiments, the display device 200 monitors the startup status of the image interface 241 when it is activated. If the startup status of the image interface is normal, it enters the startup preview mode and acquires a user image as a preview frame image. The startup status of the image interface 241 includes whether the image interface 241 is started normally and whether the image acquisition device or image acquisition module connected to the image interface 241 is started normally. By acquiring the preview frame image, the user's reference limb joints in the preview frame can be detected, and a set of feature points can be generated based on these reference limb joints as reference feature data.
[0089] Conversely, if the image interface 241 is in an abnormal startup state, the display device 200 controls the monitor 260 to display an image acquisition error message window to prompt the user to troubleshoot. For example, if the abnormal state indicates a problem with poor contact between the image interface 241 and the external image acquisition unit, the user can reconnect the image interface 241 to the external image acquisition unit according to the prompt message in the window to resolve the abnormal state of the image interface 241.
[0090] In some embodiments, the display device 200 acquires a preview frame image, which is a user image captured by the image interface 241 when the preview is started. Then, it extracts baseline feature data from the preview frame image based on a limb detection algorithm. The baseline feature data is a set of feature points generated based on baseline limb joints, and these baseline limb joints include the coordinates of key points at the limb joint positions in the preview frame. That is, the display device 200 can identify the feature points of each limb joint in the preview frame and represent the corresponding feature points using preset key point coordinates to form baseline feature data.
[0091] For example, the preset key point coordinate format is: 0-nose, 1-right shoulder, 2-left shoulder, 3-right elbow, 4-left elbow, 5-right wrist, 6-left wrist, 7-right hip, 8-left hip, 9-right knee, 10-left knee, 11-right ankle, 12-left ankle. After the display device 200 recognizes the above limb key points in the preview, it forms the baseline feature data: 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12.
[0092] After extracting the baseline feature data from the preview frame, the display device 200 draws connection lines for the baseline limb joints according to configuration parameters to generate a baseline feature model. The configuration parameters include a preset connection order and model height. In other words, the display device 200 obtains the user's individual limb joints based on the baseline limb joints detected in the preview frame. Then, it generates the baseline feature model by connecting these baseline joints.
[0093] For example, the preset keypoint coordinate format is: 0-nose, 1-right shoulder, 2-left shoulder, 3-right elbow, 4-left elbow, 5-right wrist, 6-left wrist, 7-right hip, 8-left hip, 9-right knee, 10-left knee, 11-right ankle, 12-left ankle. The reference feature data extracted by the display device 200 in the preview frame is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12. Then, the reference joints of each limb are connected according to the following rules: head: 0; left shoulder connection: 1→3→5; right shoulder connection: 2→4→6; shoulder connection: joint point: 1→2; waist connection: joint point: 7→8. For example... Figure 9 As shown, Figure 9 This is a schematic diagram illustrating the effect of a benchmark feature model.
[0094] S200: Extract adjustment feature data of the frame.
[0095] Since the height of the displayed character in the frame may differ from that of the virtual character in the reference feature model, merging the virtual character with the media asset playback may result in visual inconsistencies and reduce the display quality of the display device 200. Therefore, after detecting a frame of the media asset data, the display device 200 extracts adjustment feature data from the frame, specifically the joint points of the displayed character's limbs, as adjustment parameters for the reference feature model. By adjusting this feature data, the height and position of the reference feature model can be adjusted to ensure greater harmony between the reference feature model and the merged media asset data.
[0096] S300: Generates feature scaling parameters and feature localization parameters based on adjusted feature data.
[0097] After acquiring the aforementioned adjustment feature data, the display device 200 can determine the scaling ratio and rendering position of the reference feature model based on the adjustment feature data, which serve as feature scaling parameters and feature positioning parameters. The feature scaling parameters are used to enlarge or reduce the reference feature model, while the feature positioning parameters are used to determine the rendering position of the reference feature model.
[0098] like Figure 10 As shown, in some embodiments, when the display device 200 generates feature scaling parameters based on adjustment feature data, it acquires the first frame of the media asset data. The first frame is a frame randomly selected by the display device 200 from the loaded media asset data. After acquiring the first frame from the media asset data, the display device 200 detects the adjustment feature data of the first frame, whereby the adjustment feature parameters include the number of limb joints in the first frame. Since in the playback of the media asset data, the displayed character may not show all limb joints due to being turned to the side, or a certain frame may not include the displayed character. If the frame contains few limb joints or no displayed character, the display device 200 cannot determine the position and height of the displayed character. Therefore, after acquiring the first frame, the display device 200 detects the number of limb joints included in the first frame. If the number of limb joints is greater than a joint threshold, it indicates that the first frame contains enough limb joints, and the distance between the limb joints is calculated as a feature scaling parameter.
[0099] For example, the joint threshold is 8. After acquiring the first frame of the media asset data, the display device 200 identifies the limb joints in the first frame based on the limb detection algorithm to obtain the number of limb joints included in the first frame. When the number of limb joints is greater than 8, the distance between adjacent limb joints is calculated and used as a feature scaling parameter for comparison with the benchmark feature model.
[0100] Obviously, when the data of the first frame is detected to be less than or equal to the joint point threshold, the display device 200 needs to re-acquire frame data from the media asset data until the acquired frame data includes enough limb joint points. Since the media asset data includes frame data from multiple time points, a frame data from a different time point than the first frame data can be acquired based on the corresponding time point. Therefore, if the number of limb joint points is less than or equal to the joint point threshold, a second frame data from the media asset data is acquired to calculate the feature scaling parameter based on the distance between limb joint points in the second frame data. Here, the second frame data refers to a frame data from a different time point than the first frame data.
[0101] like Figure 11As shown, in some embodiments, when the display device 200 generates feature positioning parameters based on adjustment feature data, it acquires the first frame of the media asset data. The first frame is a frame randomly selected by the display device 200 from the loaded media asset data. After acquiring the first frame from the media asset data, the display device 200 detects the adjustment feature data of the first frame. The adjustment feature data includes limb joints in the first frame, such as the ankle joint, the left shoulder joint, and the right wrist joint. Since the feature positioning parameters are used to determine the rendering position of the baseline feature model, a specific joint is needed to determine the position of the displayed person in the media asset data. Therefore, the display device 200 presets positioning feature points to determine the position of the displayed person. For example, the positioning feature point can be the nose or the ankle, or a combination of both.
[0102] If a limb joint has a positioning feature point, it means that the first frame includes a positioning feature point preset by the display device 200. The horizontal position of the displayed person can be determined through the first frame, and the horizontal position of the positioning feature point is obtained as a feature positioning parameter. Here, the positioning feature point is a key point at the preset limb joint position. If a limb joint does not have a positioning feature point, it means that the first frame does not contain a preset positioning feature point, and the horizontal position of the displayed person cannot be determined through the first frame. In this case, the second frame of the media data is obtained, and the feature positioning parameter is calculated based on the positioning feature point of the second frame. Here, the second frame is a frame from a different time point than the first frame.
[0103] In other words, the display device 200 can determine the horizontal position of the person displayed in the media asset data based on preset positioning feature points. Therefore, the display device 200 needs to acquire frame images from the media asset data and detect the position of the positioning feature points of the person displayed in the frame images. However, since the person displayed in the frame images of the media asset data may not appear in the current frame due to special actions or because the person is not displayed in the current frame, if there are no positioning feature points in the frame images, the display device 200 will reacquire frame images from the media asset data until there are positioning feature points in the frame images.
[0104] For example, the location feature point can be a key point on the nose or ankle. The display device 200 acquires the first frame of the media asset data. If a key point on the ankle exists in the first frame, the display device 200 acquires the horizontal position of the key point on the ankle as a feature location parameter. When rendering the virtual character model, the display device 200 renders according to the position corresponding to the feature location parameter.
[0105] Furthermore, since the posture of the person displayed in the media asset data is constantly changing, special poses may occur, such as squatting or lying on their side. To improve the positioning accuracy of the feature positioning parameters, in some embodiments, the display device 200 acquires a preset number of sampled frame images, such as 8 or 10 images. These sampled frame images are those in the media asset data that include positioning feature points. The horizontal positions of the positioning feature points in the sampled frame images are then acquired to form a sample set. The mode of the sample set is calculated as the feature positioning parameter. Therefore, the display device 200 can calculate the feature positioning parameters by acquiring multiple frame images, thereby mitigating the impact of special postures of the person displayed in the media asset data on the accuracy of the feature positioning parameters and improving the positioning accuracy of the feature positioning parameters.
[0106] S400: Scale the baseline feature model according to the feature scaling parameters to generate the target feature model.
[0107] After calculating the feature scaling parameters based on the adjusted feature data, the display device 200 scales up or down the baseline feature model proportionally according to the feature scaling parameters. The baseline feature model is a virtual character model generated based on the baseline feature data. After scaling the baseline feature model according to the feature scaling parameters, the display device 200 generates a target feature model. The height of the target feature model is consistent with the height of the character displayed in the media asset data, resulting in a more harmonious blended image.
[0108] In some embodiments, when the display device 200 scales the reference feature model according to feature scaling parameters, it detects the scaling height based on the feature scaling parameters. The scaling height is equal to the height of the person displayed in the media asset data. To ensure that the height of the target feature model is consistent with the height of the person displayed in the media asset data, the display device 200 can detect the height of the person displayed using the feature scaling parameters and compare it with the height of the reference feature model. After detecting the scaling height, the display device 200 calculates the ratio of the scaling height to the model height as the scaling ratio of the reference feature model. Then, it adjusts the length of the connecting lines according to the ratio to enlarge or reduce the reference feature model.
[0109] For example: The height of the baseline model is h. The display device 200 detects the scaling height as H based on the feature scaling parameters, which is the height of the person displayed in the media asset data. The scaling ratio of the baseline feature model is calculated according to the values of h and H. Then scale the baseline virtual human to create a scaled-up version. Figure 12 The target feature model is shown.
[0110] S500: Renders the target feature model based on the feature localization parameters.
[0111] After scaling the baseline feature model according to the feature scaling parameters, the display device 200 renders the target feature model in the user interface based on the feature positioning parameters to merge the playback screen of the media asset data with the target feature model for display. The display layer of the target feature model is higher than the display layer of the corresponding playback screen of the media asset data. In this way, the target feature model is displayed at the same height and background as the characters displayed in the media asset data, making the playback screen more harmonious. Users can effectively view the screen content of the media asset data while observing their own actions.
[0112] In some embodiments, when the display device 200 renders the target feature model according to the feature positioning parameters, it detects the positioning feature points of the target feature model and sets the positioning feature points of the target feature model at the horizontal height of the feature positioning parameters. Then, it renders the target feature model at the horizontal height so that the target feature model is aligned with the displayed person in the media asset data. That is, the display device 200 can detect the position of the positioning feature points in the target feature model, set the positioning feature points of the target feature model at the horizontal height of the feature positioning parameters, and render the target feature model at that horizontal height. Figure 12 As shown, the target feature model can be positioned on the same level as the displayed figures in the media asset data.
[0113] In some embodiments, the display device 200 displays the playback screen of media asset data through a playback window and draws the target feature model using OpenGL (Open Library). For example... Figure 13 As shown, the target feature model is located above the playback window, allowing each level to perform its own function without interfering with the others.
[0114] The display device 200 can, based on the above embodiments, integrate the target feature model with the playback screen of media asset data for display. To enhance the display effect of the target feature model in the playback screen, in some embodiments, the display device 200 also collects the pixel colors of the screen frames. Then, it extracts the pixel color with the largest proportion in the screen frame as the main color tone of the screen frame. For example, it extracts prominent colors from the screen frame using a Palette tool and determines the main color tone by obtaining the RGB proportions. After determining the main color tone, the display device 200 sets the rendering color of the target feature model according to the main color tone. The rendering color is a pixel color whose color difference value with the main color tone is greater than or equal to a color difference threshold. The color difference threshold is a preset color difference value. By setting the rendering color of the target feature model to a color with a certain color difference value from the background color, the display effect of the target feature model is improved.
[0115] To ensure a continuous color contrast between the rendered color of the target feature model and the background of the media asset data, in some embodiments, the display device 200 periodically collects the pixel colors of the screen frames, extracts the dominant color tone of the screen frames, and then adjusts the rendered color of the target feature model based on the dominant color tone. That is, the display device 200 can adjust the rendered color of the target feature model by collecting screen frame colors in real time. When the background color of the media asset data screen frames changes, the display device 200 can promptly adjust the rendered color of the target feature model to improve its display effect.
[0116] Furthermore, to enable the target feature model to follow the user's movements in real time, in some embodiments, the display device 200 monitors the target position of limb joints in the reference feature data. Then, it generates pose data of the limb joints based on the displacement data of the target position, and drives the target feature model through the pose data, so that the target feature model synchronously performs actions following the limb movements in the user's image.
[0117] For example, after rendering the target feature model, the display device 200 identifies the user's movements by observing changes in the positions of limb joints in the user's image, and generates pose data based on these position changes. The generated pose data then drives the target feature model's movements, such as... Figure 14 As shown, users can drive the target feature model to move synchronously by imitating the body movements in the media asset video, thereby enhancing the user's sense of interaction.
[0118] Based on the aforementioned display device 200, some embodiments of this application also provide a method for playing media assets incorporating virtual characters, such as... Figure 8 As shown, the method includes the following steps:
[0119] S100: In response to a playback command for media asset data, detect the frame of the media asset data;
[0120] S200: Extract the adjustment feature data of the image frame;
[0121] S300: Generate feature scaling parameters and feature positioning parameters based on the adjusted feature data;
[0122] S400: Scale the baseline feature model according to the feature scaling parameters to generate the target feature model, wherein the baseline feature model is a virtual character model generated based on the baseline feature data;
[0123] S500: Render the target feature model according to the feature positioning parameters, wherein the display level of the target feature model is higher than the display level of the playback screen corresponding to the media asset data.
[0124] As can be seen from the above technical solutions, the display device and media asset playback method integrating virtual characters provided in some embodiments of this application can respond to the playback command of media asset data and detect the frame of the media asset data. Then, adjustment feature data of the frame is extracted, and feature scaling parameters and feature positioning parameters are generated based on the adjustment feature data. A reference feature model is scaled according to the feature scaling parameters to generate a target feature model. The reference feature model is a virtual character model generated based on the reference feature data. The target feature model is then rendered according to the feature positioning parameters, wherein the display level of the target feature model is higher than the display level of the corresponding playback screen of the media asset data. This method can display the virtual character and the displayed character in the playback screen at the same height and bottom, thus integrating the virtual character with the playback screen of the media asset data and improving the display effect of the virtual character.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0126] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.
Claims
1. A display device, characterized in that, include: The monitor is configured to display the user interface; An image interface is configured to acquire user images, the user images including baseline feature data; The controller is configured as follows: Acquire a preview frame image, wherein the preview frame image is a user image captured by the image interface when the preview is started; Based on the limb detection algorithm, reference feature data is extracted from the preview frame image. The reference feature data is a set of feature points generated based on reference limb joints. The reference limb joints include the coordinates of key points of the limb joint positions in the preview frame. The connection lines of the reference limb joints are drawn according to the configuration parameters to generate a reference feature model; the configuration parameters include a preset connection order and model height, and the reference feature model is a virtual character model generated based on the reference feature data; In response to a playback command for media asset data, the first frame of the media asset data is acquired; The adjustment feature data of the first frame is detected, and the adjustment feature data includes the number of limb joints in the first frame. If the number of limb joints is greater than the joint threshold, the distance between the limb joints is calculated and used as a feature scaling parameter. If the number of limb joints is less than or equal to the joint threshold, a second frame of the media data is obtained to calculate the feature scaling parameter based on the distance between limb joints in the second frame. The second frame is a frame taken at a different time point than the first frame; In addition, a preset number of sampled frame images are acquired from the media asset data; Determine the horizontal position of the positioning feature points in a preset number of the sampled image frames; Calculate the mode of the horizontal position to obtain the feature positioning parameters; The scaling height is detected based on the feature scaling parameters, and the scaling height is equal to the height of the person displayed in the media asset data; Calculate the ratio of the scaled height to the model height; Adjust the length of the connecting line according to the ratio, and scale the reference feature model to generate the target feature model; The target feature model is rendered according to the feature positioning parameters, and the display level of the target feature model is higher than the display level of the playback screen corresponding to the media asset data; Monitor the target location of limb joints in the baseline feature data; The pose data of the limb joints is generated based on the displacement data of the target position; The pose data drives the target feature model so that the target feature model follows the movements of the person in the user image.
2. The display device according to claim 1, characterized in that, The controller generates feature localization parameters based on the adjusted feature data and is configured as follows: Obtain the first frame of the media asset data; Detect adjustment feature data of the first frame, the adjustment feature data including limb joint points in the first frame; If the limb joint has a positioning feature point, obtain the horizontal position of the positioning feature point as the feature positioning parameter; The positioning feature points are key points at preset limb joint positions; If the limb joint does not have the positioning feature point, the second frame of the media data is acquired to calculate the feature positioning parameter based on the positioning feature point of the second frame; the second frame is a frame at a different time point than the first frame.
3. The display device according to claim 2, characterized in that, The controller is configured to render the target feature model based on the feature localization parameters. Detect the localized feature points of the target feature model; Set the location feature points of the target feature model at the horizontal height of the feature location parameters; The target feature model is rendered at the specified horizontal height so that it is aligned with the displayed person in the media asset data.
4. The display device according to claim 1, characterized in that, The controller is also configured to: Collect the pixel colors of the aforementioned image frames; Extract the color of the pixel with the largest proportion in the frame and use it as the main color of the frame; The rendering color of the target feature model is set according to the main color tone, wherein the rendering color is the pixel color whose color difference value with the main color tone is greater than or equal to the color difference threshold.
5. The display device according to claim 1, characterized in that, The controller is configured to: In response to a playback command for the media asset data, a playback window for the media asset data is created; Decoding the media asset data is performed to obtain a playback screen, which consists of frame images arranged in chronological order. The display is controlled to show the playback screen of the media data through the playback window.
6. A method for playing media assets incorporating virtual characters, characterized in that, include: Acquire a preview frame image, which is a user image captured by the image interface when the preview is started; Based on the limb detection algorithm, reference feature data is extracted from the preview frame image. The reference feature data is a set of feature points generated based on reference limb joints. The reference limb joints include the coordinates of key points of the limb joint positions in the preview frame. The connection lines of the reference limb joints are drawn according to the configuration parameters to generate a reference feature model; the configuration parameters include a preset connection order and model height, and the reference feature model is a virtual character model generated based on the reference feature data; In response to a playback command for media asset data, the first frame of the media asset data is acquired; The adjustment feature data of the first frame is detected, and the adjustment feature data includes the number of limb joints in the first frame. If the number of limb joints is greater than the joint threshold, the distance between the limb joints is calculated and used as a feature scaling parameter. If the number of limb joints is less than or equal to the joint threshold, a second frame of the media data is obtained to calculate the feature scaling parameter based on the distance between limb joints in the second frame. The second frame is a frame taken at a different time point than the first frame; In addition, a preset number of sampled frame images are acquired from the media asset data; Determine the horizontal position of the positioning feature points in a preset number of the sampled image frames; Calculate the mode of the horizontal position to obtain the feature positioning parameters; The scaling height is detected based on the feature scaling parameters, and the scaling height is equal to the height of the person displayed in the media asset data; Calculate the ratio of the scaled height to the model height; Adjust the length of the connecting line according to the ratio, and scale the reference feature model to generate the target feature model; The target feature model is rendered according to the feature positioning parameters, and the display level of the target feature model is higher than the display level of the playback screen corresponding to the media asset data; Monitor the target location of limb joints in the baseline feature data; The pose data of the limb joints is generated based on the displacement data of the target position; The pose data drives the target feature model so that the target feature model follows the movements of the person in the user image.