Display device and avatar generation method
By extracting non-facial features from media asset video frames and user facial features from display devices, and combining them with a neural network model to generate personalized virtual avatars, the problem of display devices being unable to meet users' personalized needs is solved, thus improving user experience and immersion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JUHAOKAN TECH CO LTD
- Filing Date
- 2026-03-04
- Publication Date
- 2026-07-07
AI Technical Summary
The virtual avatars generated by existing display devices cannot meet users' personalized needs and cannot be linked to media content, resulting in a reduced user experience.
By extracting non-facial features of people in video frames at specific playback times from media asset videos and combining them with the user's facial features, a personalized virtual avatar is generated. The feature extraction and matching are performed using a neural network model, and the avatar supports user editing and export. The computation is performed using an edge-cloud collaborative architecture.
The generated virtual avatars are highly integrated with media content, enhancing user immersion and personalized experience, meeting diverse user needs, and reducing the computational load on display devices.
Smart Images

Figure CN122349034A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of display device technology, and in particular to a display device and a method for generating virtual images. Background Technology
[0002] Display devices refer to terminal devices capable of outputting specific display images, such as smart TVs, communication terminals, smart advertising screens, and projectors. Taking smart TVs as an example, smart TVs are television products based on Internet application technologies, possessing open operating systems and chips, and having open application platforms. They enable two-way human-computer interaction and integrate multiple functions such as audio-visual, entertainment, and data to meet diverse and personalized user needs.
[0003] Display devices can generate virtual avatars based on built-in templates. These templates are fixed, limiting users to choosing the appearance and style of the generated avatar. Clearly, this limitation makes it difficult to create virtual avatars that meet users' personalized needs. Furthermore, the static, pre-set templates prevent the generated avatars from being associated with the specific scenarios in which they are used, thus degrading the user experience. Summary of the Invention
[0004] This application provides a display device and a method for generating virtual avatars to solve the problem that the generated virtual avatars cannot meet the personalized needs of users.
[0005] In a first aspect, this application provides a display device, comprising: The display is configured to display a user interface, the user interface including media asset video; The controller is configured as follows: In response to a first instruction to generate a virtual avatar, the first playback time point of the media asset video is determined upon receiving the first instruction; Extract the first video frame corresponding to the first playback time point from the media asset video; Extract the non-facial features of the target person in the first video frame; the target person is the person appearing in the first video frame. Obtain the user's facial features; Based on the facial features, a basic head model is generated; The non-facial features are fused onto the base head model to obtain a virtual image; The display is controlled to show the virtual image on top of the media asset video.
[0006] The above technical solution has the following beneficial effects or advantages: When a display device plays media asset videos, it extracts non-facial features of characters from video frames corresponding to specific playback times and combines these features with the user's facial features to generate a virtual avatar. The generated virtual avatar is no longer a generic avatar, but a personalized avatar highly integrated with the style of the media asset content the user is watching. Furthermore, the virtual avatar can be displayed synchronously while the media asset video is playing, enhancing the user's immersion and personalized experience.
[0007] In some embodiments of this application, the controller extracts non-facial features of the target person in the first video frame, and is configured to: According to the first instruction, obtain descriptive information about the target person's image; Based on the description information, the image of the target person is identified in the first video frame; Extract the target image region containing the image of the target person; The target image region is input into the non-facial feature extraction model to extract non-facial features and obtain the non-facial features; the non-facial feature extraction model is a neural network model pre-trained based on sample images labeled with non-facial features.
[0008] The above technical solution has the following beneficial effects or advantages: The display device achieves accurate extraction of non-facial features of a target person by combining instruction semantics with a neural network model. This solves the target recognition problem in multi-person scenes. When multiple people appear in a video frame, recognition is performed based on the descriptive information of the target person, accurately locating the specific person the user intends to point to within the video frame and avoiding the extraction of incorrect non-facial features. Furthermore, by utilizing a pre-trained non-facial feature extraction model, it can deeply understand and extract complex visual features, improving the accuracy of feature extraction.
[0009] In some embodiments of this application, the controller acquires the user's facial features and is configured to: Obtain the user's facial information; the facial information includes facial images and / or facial videos; The facial information is input into the facial feature extraction model to extract facial features and obtain the facial features; the facial feature extraction model is a neural network model pre-trained based on sample images labeled with facial feature tags.
[0010] The above technical solution has the following beneficial effects or advantages: The display device uses a pre-trained facial feature extraction model to extract facial features, which can extract structured, high-quality feature vectors from the user's facial images and / or facial videos, thereby improving the accuracy of feature extraction.
[0011] In some embodiments of this application, the controller generates a basic head model based on the facial features and is configured to: Call the header model library; the header model library includes multiple preset header models; Calculate the matching degree between the facial features and each of the head models; Based on the matching degree between the facial features and each head model, a target head model is determined; the target head model is a head model in the head model library whose matching degree with the facial features is greater than a preset matching degree threshold. The facial features are fused onto the target head model to obtain the base head model.
[0012] The above technical solution has the following beneficial effects or advantages: The display device comes pre-installed with a head model library to assist in generating a basic head model. By selecting a head model with a high degree of matching from the library as the target head model, the computational overhead of modeling can be avoided, thus improving the generation speed. Furthermore, the user's facial features are integrated into the target head model to generate a basic head model that meets the user's personalized needs.
[0013] In some embodiments of this application, a first communication device and a second communication device are also included; the first communication device is configured to establish a communication connection with a mobile terminal; the second communication device is configured to establish a communication connection with a server. After the controller extracts the first video frame corresponding to the first playback time point from the media asset video, it is configured to: Send a first notification to the mobile terminal, the first notification being used to instruct the mobile terminal to upload the user's facial information; Receive the user's facial information sent by the mobile terminal; The first video frame and the facial information are sent to the server so that the server generates the virtual image based on the first video frame and the facial information; Receive the virtual image sent by the server.
[0014] The above technical solution has the following beneficial effects or advantages: Virtual avatars are generated through an edge-cloud collaborative architecture. Mobile terminals are used to collect users' facial information, improving user convenience. The computationally intensive virtual avatar generation task is offloaded to the server, leveraging the computing power of the cloud server to ensure the quality of the generated virtual avatar while reducing the computational load on the display device.
[0015] In some embodiments of this application, after the controller fuses the non-facial features onto the base head model to obtain the virtual image, it is configured to: In response to a user's editing command for the virtual avatar, the display is controlled to show an editing interface; the editing interface includes at least one editing control and a preview area, and each editing control corresponds to at least one attribute of the virtual avatar; The attributes of the virtual avatar are adjusted based on the attribute parameters input by the user using the editing control. Control the display to show the adjusted virtual image in the preview area.
[0016] The above technical solution has the following beneficial effects or advantages: The display device provides a virtual avatar editing function, allowing users to further adjust the virtual avatar through the editing interface and preview the adjusted virtual avatar in real time, thus meeting users' personalized needs.
[0017] In some embodiments of this application, after the controller fuses the non-facial features onto the base head model to obtain the virtual image, it is configured to: In response to the export command of the virtual avatar, the display is controlled to show the export interface; the export interface includes at least one export control, and each export control corresponds to the export format of a virtual avatar; Determine the target export format corresponding to the target export control selected by the user; the target export control is any one of the export controls in the export interface. Based on the virtual avatar, media asset content corresponding to the target export format is generated; the media asset content includes the virtual avatar. Output the media asset content.
[0018] The above technical solution has the following beneficial effects or advantages: Display devices can transform virtual avatars into media content in various formats, enabling their application in various platforms and applications such as social sharing, short video production, and games, to meet diverse user needs.
[0019] In some embodiments of this application, after the controller determines the first playback time point of the media asset video upon receiving the first instruction, it is further configured to: Based on the first playback time point, the target duration is traced back to obtain the second playback time point of the media asset video; the target duration is a preset duration or the duration required to receive the first instruction. Extract the second video frame corresponding to the second playback time point from the media asset video; Extract the non-facial features of the target person in the second video frame.
[0020] The above technical solution has the following beneficial effects or advantages: By tracing the playback time backwards to the target duration, display devices can more accurately pinpoint the video frame that the user actually saw and that triggered their intent, thereby extracting the non-facial features of the person the user truly wants and reducing the false extraction rate.
[0021] In some embodiments of this application, after the controller determines the first playback time point of the media asset video upon receiving the first instruction, it is further configured to: In the media video, a first video segment within a first duration before the first playback time point and a second video segment within a second duration after the first playback time point are extracted. Extract a third video frame from the first video segment and the second video segment. The third video frame is a video frame that includes the image of the target person. Extract the non-facial features of the target person in the third video frame.
[0022] The above technical solution has the following beneficial effects or advantages: The display device provides multiple video frames from the time period before and after the first playback point as candidates, and selects video frames containing the target character's image from these frames to extract non-facial features. This avoids the failure of virtual character generation due to poor quality of a single video frame, thus improving the user experience.
[0023] Secondly, this application provides a method for generating a virtual avatar, including: In response to the first instruction to generate a virtual avatar, determine the first playback time point of the media asset video upon receiving the first instruction; Extract the first video frame corresponding to the first playback time point from the media asset video; Extract the non-facial features of the target person in the first video frame; the target person is the person appearing in the first video frame. Obtain the user's facial features; Based on the facial features, a basic head model is generated; The non-facial features are fused onto the base head model to obtain a virtual image; The virtual avatar is displayed on top of the media asset video.
[0024] The above technical solution has the following beneficial effects or advantages: When playing media asset videos, non-facial features of characters in video frames corresponding to specific playback times are extracted and combined with the user's facial features to generate a virtual avatar. The generated virtual avatar is no longer a generic avatar, but a personalized avatar highly integrated with the style of the media asset content the user is watching. Furthermore, the virtual avatar can be displayed synchronously while the media asset video is playing, enhancing the user's immersion and personalized experience. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application; Figure 2 This is a schematic diagram of the hardware configuration of a display device provided in some embodiments of this application; Figure 3 This is a schematic diagram of the software configuration of a display device provided in some embodiments of this application; Figure 4 A flowchart illustrating a virtual avatar generation method provided in some embodiments of this application; Figure 5 A schematic diagram illustrating the process of extracting non-facial features provided for some embodiments of this application; Figure 6 A schematic diagram of a character selection interface provided in some embodiments of this application; Figure 7 A schematic diagram illustrating the process of extracting a second video frame provided for some embodiments of this application; Figure 8 A schematic diagram illustrating the process of extracting a third video frame provided for some embodiments of this application; Figure 9 A schematic diagram of a video frame selection interface provided in some embodiments of this application; Figure 10 This is a schematic diagram of the interface for video frame selection provided in some embodiments of this application; Figure 11 A schematic diagram illustrating the process of extracting facial features provided in some embodiments of this application; Figure 12 A schematic diagram illustrating the process of generating a basic header model provided for some embodiments of this application; Figure 13These are schematic diagrams illustrating the display effects of virtual avatars provided in some embodiments of this application. Figure 14 A schematic diagram of an editing interface provided for some embodiments of this application; Figure 15 A schematic diagram illustrating the export interface provided in some embodiments of this application; Figure 16 Timing diagrams for virtual image generation provided in some embodiments of this application; Figure 17 This is a schematic diagram of the module of the end-to-cloud collaborative architecture provided in some embodiments of this application. Detailed Implementation
[0027] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.
[0028] In this embodiment, the display device 200 generally refers to a device with screen display and data processing capabilities. For example, the display device 200 includes, but is not limited to, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.
[0029] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application. For example... Figure 1 As shown, a user can operate the display device 200 via touch operation, a mobile terminal 300, and a control device 100. The control device 100 receives user input commands and converts them into control commands that the display device 200 can recognize and respond to. For example, the control device 100 can be a remote control, a stylus, a gamepad, etc.
[0030] The mobile terminal 300 can function as a control device for human-computer interaction between the user and the display device 200. It can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can have software applications installed on it and communicate with the display device 200 via network communication protocols to achieve one-to-one control and data communication. Furthermore, it can transmit audio and video content displayed on the mobile terminal 300 to the display device 200 for synchronized display.
[0031] In some embodiments, the mobile terminal 300 or other electronic devices may also simulate the functions of the control device 100 by running an application that controls the display device 200.
[0032] like Figure 1 The diagram also shows that the display device 200 communicates with the server 400 via various communication methods. This allows the display device 200 to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.
[0033] Display device 200 can provide broadcast television reception function, and can also be equipped with intelligent network television function that provides computer support function, including but not limited to network television, smart television, Internet Protocol television (IPTV), etc.
[0034] Figure 2 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of display device 200.
[0035] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface 280.
[0036] In some embodiments, detector 230 is used to acquire signals from the external environment or to interact with the outside world. For example, detector 230 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.
[0037] In some embodiments, the display 260 includes display function components for presenting images and driving components for driving image display. The display 260 is used to receive and display image signals output from the controller 250. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces, etc.
[0038] In some embodiments, the communication device 220 is a component used to communicate with external devices or the server 400 according to various communication protocol types. The display device 200 may have multiple communication devices 220 depending on the supported communication methods. For example, when the display device 200 supports wireless network communication, it may have a communication device 220 with WiFi functionality. When the display device 200 supports Bluetooth connectivity, it needs to have a communication device 220 with Bluetooth functionality.
[0039] The communication device 220 enables the display device 200 to communicate with external devices or the server 400 via wireless or wired connections. Wired connections utilize data cables, interfaces, or other components to connect the display device 200 to external devices. Wireless connections utilize wireless signals or wireless networks. The display device 200 can directly establish a connection with external devices or indirectly through gateways, routers, or other connection devices.
[0040] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200.
[0041] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0042] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 260, and the user input interface 280 receives the user input commands through the graphical user interface (GUI).
[0043] In some embodiments, the audio output device 270 can be a built-in speaker of the display device 200 or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the display device 200 to output sound from the display device 200.
[0044] In some embodiments, the user input interface 280 can be used to receive instructions from user input.
[0045] To enable user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program used to manage and control the hardware and software resources of the display device 200. The operating system can control the display device to provide a user interface; for example, the operating system can directly control the display device to provide a user interface, or it can provide a user interface by running an application. The operating system also allows users to interact with the display device 200.
[0046] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for display devices.
[0047] An operating system can be divided into different modules or levels based on the functions it implements, for example... Figure 3 As shown, in some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the System Runtime Library layer, and the Kernel Layer.
[0048] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer may contain at least one application, which may be a built-in Windows program, system settings program, or clock program of the operating system; or it may be an application developed by a third-party developer. In specific implementations, the application packages in the application layer are not limited to the examples above.
[0049] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.
[0050] like Figure 3As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, text boxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0051] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.
[0052] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as the C / C++ instruction library, to implement the functions to be performed by the framework layer.
[0053] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 3 As shown, hardware drivers can be configured in the kernel layer. The kernel layer can contain at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.
[0054] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the display device 200 in this application embodiment. Depending on the function of the display device, the type of operating system, and other factors, the number of levels and the specific level type of the operating system may be expressed in other forms.
[0055] During operation, the display device 200 can generate a virtual avatar based on user needs and control the display screen 260 to display the virtual avatar. The virtual avatar refers to a visualized two-dimensional or three-dimensional model generated using digital technology.
[0056] Display device 200 can display virtual avatars statically or dynamically. Display device 200 can respond to control commands input by the user via voice, gestures, or control device 100, and execute specific animation sequences through the virtual avatar, such as lip-syncing, facial expression changes, and motion changes, thereby enhancing interactivity with the user.
[0057] Virtual avatars can serve as a user's identity in the digital world and can be widely used in various scenarios. For example, they can be used as visual representatives of participants in virtual meeting applications; as user-controlled characters in gaming applications; as material for generating personalized videos in social media applications; and as visual avatars of voice assistants in smart home applications.
[0058] In some embodiments, to achieve rapid virtual avatar generation, the display device 200 has a built-in virtual avatar template library, which includes multiple virtual avatar templates with different appearance attributes and style features. The appearance attributes include, but are not limited to, hairstyle, face shape, clothing, etc., and the style features include, but are not limited to, cartoon style, realistic style, ancient style, etc.
[0059] Users can select the appearance and style of the virtual avatar they want to generate from the virtual avatar template library. The display device 200 can load the corresponding template data from the virtual avatar template library according to the user's selection, generate the final virtual avatar by combination, and control the monitor 260 to display the virtual avatar. It can be seen that, limited by fixed templates, it is difficult to generate virtual avatars that meet the user's personalized needs.
[0060] Furthermore, the virtual avatar template library is statically pre-built. Virtual avatars generated based on fixed templates cannot be associated with the specific scenarios in which they are used, thus reducing user experience and immersion. For example, when display device 200 is playing a period drama, the user expects to generate a virtual avatar with the characteristics of the characters' traditional costumes. However, the display device 200's pre-built virtual avatar templates do not contain template data with the same characteristics of the characters' traditional costumes. As a result, the generated virtual avatar cannot establish a linkage effect with the currently playing drama, making it difficult to integrate into the user's current experience environment and reducing user experience.
[0061] To address the aforementioned issues, the display device 200 provided in this embodiment can execute a virtual avatar generation method. This method generates a virtual avatar based on a character in a current media asset video, and associates the virtual avatar with the specific scenario where it is applied, thereby improving user experience and immersion. The display device 200 capable of applying the virtual avatar generation method includes at least a display 260 and a controller 250. The display 260 is configured to display a user interface, which includes the media asset video. The controller 250 is configured to execute the specific steps of the virtual avatar generation method.
[0062] like Figure 4 The diagram shown is a flowchart illustrating the virtual avatar generation method provided in this application embodiment, which specifically includes the following steps: S401, in response to the first instruction to generate the virtual avatar, determines the first playback time of the media asset video upon receiving the first instruction.
[0063] When playing media asset videos, the display device 200 can respond to the first instruction to generate a virtual image and execute the corresponding virtual image generation process to generate a virtual image associated with the target character image in the media asset video.
[0064] The first playback time point refers to the current playback position of the media asset video when the first instruction is received. For example, if the media asset video starts playing at 00:00:00, and the time point at which the media asset video starts playing on the display device 200 is 08:00:00, and the time point at which the display device 200 receives the first instruction is 08:10:08, then the first playback time point is 00:10:08.
[0065] In some embodiments, the input method for the first instruction includes multiple methods. The display device 200 can control the display 260 to display functional controls in the user interface for instructing the generation of a virtual avatar, and generate the first instruction in response to the user's selection of the functional controls. The display device 200 can also receive key signals from target physical buttons sent by the control device 100, and generate the first instruction in response to the key signals.
[0066] In some embodiments, when the display device 200 is playing media video, to avoid user operations affecting the playback effect of the video, the display device 200 can provide a voice input function. The display device 200 includes a sound acquisition unit, which can acquire and parse the user's voice commands, confirm the user's intent, and generate a first command when the user's intent instructs the generation of a virtual avatar.
[0067] S402, extract the first video frame corresponding to the first playback time point from the media asset video.
[0068] After the display device 200 determines the first playback time point, it can extract the first video frame corresponding to the first playback time point from the media asset video for use in the subsequent virtual avatar generation process.
[0069] The media asset video consists of multiple video frames ordered in a time sequence. The first video frame refers to the video frame corresponding to the first playback time point in the media asset video. For example, if the first playback time point is 00:10:08, the display device 200 will locate the video frame at 00:10:08 in the media asset video as the first video frame and extract it for use in the subsequent virtual avatar generation process.
[0070] In some embodiments, the display device 200 can access the corresponding video source file based on the media asset information of the currently playing media asset video, such as video identifier ID and Uniform Resource Locator (URL), locate and decode the video frame corresponding to the first playback time point in the video source file, and use it as the first video frame.
[0071] In some embodiments, the display device 200 may, in response to a first instruction, directly perform a screenshot operation on the video frame currently being rendered by the display device 200 to obtain the first video frame.
[0072] In some embodiments, in order to ensure the accuracy of video frame screenshots, the display device 200 may respond to the first instruction, pause the playback of the media asset video, and then take a screenshot of the currently rendered video frame to obtain the first video frame.
[0073] S403, extract the non-facial features of the target person in the first video frame.
[0074] After the display device 200 extracts the first video frame, it can extract the non-facial features of the target person's image in the first video frame for use in the subsequent virtual image generation process.
[0075] The target person image refers to the person image appearing in the first video frame. A person image is a visual object that can be detected and identified within a video frame. Non-facial features are visual attributes of the target person image other than facial features, including but not limited to clothing features (such as texture, color, style, etc.), hairstyle features (such as outline, layering, hair color, hair accessories, etc.), makeup features (such as eyebrow shape, lip color, etc.), features of other body parts (such as hand features, neck features, etc.), and style features (such as cartoon style, traditional Chinese style, etc.).
[0076] In some embodiments, such as Figure 5 The diagram shown is a flowchart illustrating the extraction of non-facial features according to an embodiment of this application, specifically including the following steps: S501, Identify the target person's image in the first video frame and extract the target image region containing the target person's image.
[0077] After the display device 200 acquires the first video frame, it can identify the target person's image in the first video frame and extract the target image area corresponding to the target person's image to facilitate subsequent feature extraction.
[0078] In some embodiments, the first video frame includes multiple human figures. To avoid misjudgment, the display device 200 can identify the target human figure based on user-input description information. The display device 200 can obtain description information of the target human figure according to a first instruction, and identify the target human figure in the first video frame based on the description information.
[0079] For example, when the display device 200 is playing a period drama, the user's voice commands can be collected by a sound acquisition device, and semantic analysis can be performed on the voice commands to obtain descriptive information, such as "generate the image of the character on the left" or "generate the image of the girl in red". The display device 200 can then locate the corresponding target character image in the first video frame based on the descriptive information and image recognition technology.
[0080] In some embodiments, to improve user experience and avoid misjudgment, the display device 200 can identify the target person image based on the user's active selection. After acquiring the first video frame, the display device 200 can identify and extract all the person images in the first video frame, and generate a person image selection interface based on all the person images. The person image selection interface includes selection items corresponding to all person images. In response to the user's selection operation on the target selection item, the person image corresponding to the target selection item is determined as the target person image.
[0081] In some embodiments, to facilitate user selection, the display device 200 may pause the playback of the media video while displaying the character selection interface, control the display 260 to display the character selection interface in full screen, and after confirming the target character selected by the user, cancel the display of the character selection interface and continue playing the media video, ensuring that the user's selection is not affected by changes in the video content.
[0082] In some embodiments, in order to maintain the continuity of viewing, the display device 200 may not pause the playback of the media asset video when displaying the video frame selection interface, and control the display 260 to display the character image selection interface on top of the media asset video according to a preset size and preset position.
[0083] For example, when multiple figures appear in the first video frame, the display device 200 identifies and crops all the figures in the first video frame, and accordingly generates and controls the display 260 to display such figures. Figure 6 The character selection interface 600 shown lists all character options 601 in thumbnail format. Users can select a target character from all options using a remote control to move the focus or by direct touch. The display device 200 can then identify the character corresponding to the user's selected option as the target character.
[0084] In some embodiments, the display device 200 can identify the target person image based on focus detection. After acquiring the first video frame, the display device 200 can detect the person image that is the visual focus in the first video frame as the target person image.
[0085] In some embodiments, after the display device 200 acquires the first video frame, it can detect the human figure located in the center of the screen in the first video frame as the target human figure.
[0086] In some embodiments, after the display device 200 acquires the first video frame, it can detect the size of each character in the first video frame and select the character with the largest size as the target character.
[0087] In some embodiments, after the display device 200 acquires the first video frame, it can detect the clarity of each character image in the first video frame and select the character image with the highest clarity as the target character image.
[0088] S502, input the target image region into the non-facial feature extraction model to extract non-facial features and obtain non-facial features.
[0089] Among them, the non-facial feature extraction model is a neural network model that is pre-trained based on sample images labeled with non-facial features.
[0090] For example, the non-facial feature extraction model can be a pre-trained ResNet50 model. The display device 200 inputs the target image region into the ResNet50 model and extracts feature vectors (512-dimensional) such as clothing features (e.g., texture, color, style), hairstyle features (e.g., outline, layer, hair color), and makeup features (e.g., eyebrow shape, lip color) as non-facial features for subsequent virtual image generation.
[0091] In some embodiments, the display device 200 includes a second communication device configured to establish a communication connection with the server 400. The server 400 can perform the extraction of non-facial features. That is, the server 400 deploys a non-facial feature extraction model, and after the display device 200 acquires a first video frame, it can send the first video frame to the server. The server 400 can receive the first video frame sent by the display device 200, execute steps S501-S502 to obtain non-facial features, and send the non-facial features to the display device 200. The display device 200 can receive the non-facial features extracted from the first video frame sent by the server 400 for subsequent virtual avatar generation.
[0092] In some embodiments, there may be a time lag between the user issuing a command and the display device 200 receiving the command, which may result in the extracted video frame no longer representing the scene the user intended to view; for example, the character may have changed or the camera angle may have shifted. To address this, in order to improve the accuracy of virtual avatar generation, such as... Figure 7 The diagram shown is a flowchart illustrating the extraction of the second video frame according to an embodiment of this application, specifically including the following steps: S701, based on the first playback time point, trace back the target duration to obtain the second playback time point of the media asset video.
[0093] In some embodiments, the target duration is a preset duration. For example, if the preset duration is 3 seconds, and the display device 200 determines that the first playback time point when receiving the first instruction is 00:10:08, then it will backtrack the first playback time point by 3 seconds to determine the second playback time point as 00:10:05.
[0094] In some embodiments, the target duration is the duration required to receive the first instruction. The display device 200 continuously detects audio via a sound collector. When the sound collector detects a user's voice signal, it can start timing at the moment the sound collector detects the voice signal. It continuously collects and analyzes the user's voice to confirm the user's intent. When the user's intent instructs the generation of a virtual avatar, a first instruction is generated. In response to the first instruction, timing ends to measure the duration required to receive the first instruction. This timed duration is used as the target duration for retrospectively calculating the second playback time.
[0095] S702, extract the second video frame corresponding to the second playback time point from the media asset video.
[0096] For example, if the second playback time is 00:10:05, the display device 200 will locate the video frame at 00:10:05 in the media asset video as the second video frame, and extract the second video frame for use in the subsequent virtual avatar generation process.
[0097] The implementation details of step S702 can be found in step S402, and will not be repeated here.
[0098] S703, extract non-facial features of the target person in the second video frame.
[0099] The implementation details of step S703 can be found in step S403, and will not be repeated here.
[0100] In some embodiments, to improve the accuracy of virtual avatar generation, such as Figure 8 The diagram shown is a flowchart illustrating the extraction of the third video frame according to an embodiment of this application, specifically including the following steps: S801, in the media asset video, extract the first video segment within the first duration before the first playback time point, and the second video segment within the second duration after the first playback time point.
[0101] The first and second durations are preset durations. For example, if the first playback time is 00:10:08, the first duration is 3 seconds, and the second duration is 2 seconds, then the display device 200 will extract all video frames within the time period from 00:10:05 to 00:10:08 as the first video segment, and extract all video frames within the time period from 00:10:08 to 00:10:10 as the second video segment.
[0102] S802, extract the third video frame from the first video segment and the second video segment.
[0103] The third video frame is a video frame that includes the image of the target person. The display device 200 can traverse all video frames in the first and second video segments to identify whether the video frame includes the image of the target person. If it includes the image of the target person, the video frame is determined to be the third video frame; if it does not include the image of the target person, the video frame is determined not to be the third video frame.
[0104] Regarding the identification of the target person's image, please refer to the implementation content of step S501 above, which will not be repeated here.
[0105] In some embodiments, the display device 200 may perform frame-by-frame analysis on the acquired multiple third video frames to select the most suitable third video frame for use in the subsequent virtual avatar generation process.
[0106] In some embodiments, the display device 200 can detect whether the target person image in the third video frame is the visual focus. If the target person image is the visual focus, the third video frame is retained. If the target person image is not the visual focus, the third video frame is discarded.
[0107] In some embodiments, the display device 200 can detect the area proportion of the target person's image in the video frame of the third video frame. If the area proportion is greater than or equal to a preset proportion threshold, the third video frame is retained. If the area proportion is less than the preset proportion threshold, the third video frame is discarded.
[0108] In some embodiments, the display device 200 can detect the clarity of the target person's image in the third video frame. If the clarity is greater than or equal to a preset clarity threshold, the third video frame is retained. If the clarity is less than the preset clarity threshold, the third video frame is discarded.
[0109] S803, extract non-facial features of the target person in the third video frame.
[0110] The implementation details of step S803 can be found in step S403, and will not be repeated here.
[0111] In some embodiments, in order to ensure that the video frames are the video frames that the user actually wants, the display device 200 may extract a first video segment within a first duration (e.g., 3 seconds) before the first playback time point and / or a second video segment within a second duration (e.g., 2 seconds) after the first playback time point from the media asset video, generate a video frame selection interface based on the video frames in the first video segment and / or the video frames in the second video segment, and control the display 260 to display the video frame selection interface.
[0112] The video frame selection interface includes multiple selection options for different video frames, allowing the user to choose. The display device 200 can use the video frame corresponding to any of the user-selected options for non-facial feature extraction.
[0113] In some embodiments, to facilitate user selection, the display device 200 may pause the playback of the media video while displaying the video frame selection interface, control the display 260 to display the video frame selection interface in full screen, and after confirming the video frame selected by the user, cancel the display of the video frame selection interface and continue playing the media video, ensuring that the user's selection is not affected by changes in the video content.
[0114] In some embodiments, in order to maintain the continuity of viewing, the display device 200 may not pause the playback of the media asset video when displaying the video frame selection interface, and control the display 260 to display the video frame selection interface on top of the media asset video according to a preset size and preset position.
[0115] For example, such as Figure 9 As shown, display device 200 plays media video in full screen. A smaller video frame selection interface 900 is displayed on the right side of the monitor 260 screen. The video frame selection interface 900 includes selection items 901 for each video frame. Users can select the video frame that best matches their intentions via remote control or voice command. In response to the user's selection, display device 200 uses the video frame corresponding to the user's selected selection item 901 for non-facial feature extraction.
[0116] In some embodiments, to give users more direct and precise control, when the display device 200 plays media asset videos, it controls the display 260 to display playback progress controls. These playback progress controls are used to adjust the playback time points of the media asset videos. The display device 200 can respond to a first instruction, determine a third playback time point input by the user based on the playback progress controls, and extract the video frame corresponding to the third playback time point for use in the subsequent virtual avatar generation process.
[0117] For example, such as Figure 10 As shown, the display device 200 displays a playback progress control 1001 in the form of a slider, the length of which represents the total duration of the media asset video, and the position of the slider 1002 on it corresponds to the current playback time point. Users can drag the slider 1002 using the remote control's directional keys or by directly touching the screen of the display 260 to locate any playback position in the media asset video. To facilitate precise positioning, the display device 200 can dynamically display a thumbnail 1003 of the video frame corresponding to the real-time position of the slider 1002.
[0118] When the user stops dragging the slider and enters a confirmation operation (such as pressing the confirmation button on the remote control or pausing for more than a preset time), the display device 200 can respond to the confirmation operation and determine the playback time point corresponding to the current position of the slider 1002 (e.g., Figure 10 As shown in the figure (26:50), the video frame corresponding to this playback time point is extracted for the extraction of non-facial features.
[0119] S404, Obtain the user's facial features.
[0120] When playing media video, the display device 200 can obtain the user's facial features in response to a first instruction to generate a virtual avatar, thereby generating a virtual avatar with the user's features.
[0121] Among these, facial features include, but are not limited to, facial contours (such as round face, square face, oval face), facial proportions (distance between eyes, relative position of eyes and nose), skin color and texture (such as pores, spots), etc.
[0122] In some embodiments, such as Figure 11 The diagram shown is a flowchart illustrating the process of extracting facial features according to an embodiment of this application, specifically including the following steps: S1101, Obtain the user's facial information.
[0123] Facial information includes facial images and / or facial videos.
[0124] In some embodiments, the display device 200 can capture the user's facial information via a camera. For example, when it is necessary to obtain the user's facial information, the display device 200 can control the display 260 to display a guided interface with a facial outline and display prompts such as "Please face the camera" to guide the user to collect facial information. After ensuring that the user's face is within the facial outline, the display device 200 will take a picture of the user's face or record a 3-second facial video to obtain the user's facial information.
[0125] In some embodiments, the display device 200 may access a locally stored file library to locate stored user facial information.
[0126] In some embodiments, the display device 200 includes a first communication device configured to establish a communication connection with the mobile terminal 300. The display device 200 can send a first notification to the mobile terminal, which instructs the mobile terminal 300 to upload the user's facial information in order to receive the user's facial information sent by the mobile terminal 300.
[0127] In some embodiments, the first notification may be a push notification message. For example, display device 200 controls display 260 to display a temporary QR code. Mobile terminal 300 can scan the QR code using the built-in scanning function of the installed target application to trigger the establishment of a communication link between terminal device 300 and display device 200. After the communication link is established, when display device 200 needs the user's facial information, it can send a push notification to the target application through the communication link. The content may be "Generating a virtual avatar for you, please click to upload your photo." By clicking the push notification, the user can control terminal device 300 to display the upload interface of the target application. The user can interact with mobile terminal 300 through the upload interface, controlling terminal device 300 to obtain the user's facial information stored in the local photo album, or to use a camera to capture the user's facial information and send it to display device 200.
[0128] S1102, input facial information into the facial feature extraction model, perform facial feature extraction, and obtain facial features.
[0129] Among them, the facial feature extraction model is a neural network model that is pre-trained based on sample images labeled with facial features.
[0130] For example, the facial feature extraction model can be the open-source machine learning tool MediaPipe Face Mesh. The display device 200 inputs the user's facial information into MediaPipe Face Mesh to extract 468 facial feature points. These 468 facial feature points can represent the shape and position of the user's eyebrows, eyes, nose, lips, facial contours, and other facial features.
[0131] In some embodiments, the display device 200 includes a second communication device configured to establish a communication connection with the server 400. The server 400 can perform facial feature extraction. That is, the server 400 deploys a facial feature extraction model, and after the display device 200 obtains the user's facial information, it can send the facial information to the server. The server 400 can receive the facial information sent by the display device 200, execute steps S1101-S1102 to obtain facial features, and send the facial features to the display device 200. The display device 200 can receive the facial features extracted from the facial information sent by the server 400 for subsequent virtual avatar generation.
[0132] In some embodiments, the display device 200 can store the extracted facial features of the user in a memory, and when the user's facial features are needed, the corresponding facial features can be directly searched and extracted from the memory.
[0133] S405 generates a basic head model based on facial features.
[0134] After the display device 200 acquires the user's facial features, it can generate a basic head model with user-specific features based on the facial features, which can be used for the subsequent generation of virtual avatars.
[0135] In some embodiments, such as Figure 12 The diagram shown is a flowchart illustrating the process of generating a basic header model according to an embodiment of this application, which specifically includes the following steps: S1201, call the header model library.
[0136] The display device 200 creates and maintains a head model library containing multiple preset head models that may cover a variety of head features, including but not limited to different face shapes (such as round face, square face, long face, oval face), different skull structures, and different ethnic characteristics.
[0137] S1202 calculates the matching degree between facial features and each head model.
[0138] In some embodiments, the matching degree can be measured based on cosine similarity. The display device 200 can calculate the cosine similarity between the user's facial features and the facial features configured for each head model. If the cosine similarity is greater than or equal to a preset cosine similarity threshold, the head model is determined to be the target head model.
[0139] In some embodiments, the matching degree can be measured based on Euclidean distance. The display device 200 can calculate the Euclidean distance between the user's facial features and the facial features configured for each head model. If the Euclidean distance is greater than or equal to a preset distance threshold, the head model is determined to be the target head model.
[0140] S1203, based on the matching degree between facial features and each head model, determines the target head model.
[0141] In some embodiments, the target head model is a head model in a head model library whose matching degree with facial features is greater than a preset matching degree threshold. For example, the head model library includes five head models: head model 1, head model 2, head model 3, head model 4, and head model 5. The preset matching degree threshold is 0.85, and the matching degrees between the five head models and facial features are 0.92, 0.78, 0.45, 0.80, and 0.83, respectively. Since the matching degree between head model 1 and facial features is greater than the preset matching degree threshold, head model 1 is determined to be the target head model.
[0142] In some embodiments, the target head model is the head model in the head model library that has the highest matching degree with facial features. For example, the head model library includes three head models: head model 1, head model 2, and head model 3. The matching degrees between the three head models and facial features are 0.78, 0.80, and 0.83, respectively. Since head model 2 has the highest matching degree with facial features, head model 2 is determined to be the target head model.
[0143] S1204, facial features are fused onto the target head model to obtain the basic head model.
[0144] After determining the target head model, the display device 200 can fuse the user's facial features into the target head model to generate the final base head model.
[0145] In some embodiments, in order to improve response speed, the display device 200 may directly determine the head models in the head model library that match the user's facial features with a degree greater than a preset matching degree threshold as the basic head models.
[0146] In some embodiments, the display device 200 includes a second communication device configured to establish a communication connection with the server 400. The server 400 can perform the generation of a basic head model. That is, the server 400 maintains a head model library; after the display device 200 obtains the user's facial features, it can send the facial features to the server. The server 400 can receive the facial features sent by the display device 200, execute steps S1201-S1204 to obtain a basic head model, and send the basic head model to the display device 200. The display device 200 can receive the basic head model sent by the server 400 for subsequent generation of a virtual avatar.
[0147] S406 integrates non-facial features into a basic head model to obtain a virtual avatar.
[0148] After the display device 200 obtains the basic head model, it can integrate the non-facial features of the target person's image into the user's basic head model, thereby generating a virtual image that has both user characteristics and conforms to the characteristics of the target person's image.
[0149] In some embodiments, the display device 200 may employ a generative adversarial network algorithm to fuse non-facial features with a user’s base head model to generate a virtual avatar.
[0150] S407 controls the display 260 to show a virtual avatar on top of the media asset video.
[0151] For example, such as Figure 13The diagram shows the display effect of the virtual avatar provided in this embodiment. The display device 200 creates a floating window 1301 on top of the media asset video, and displays the virtual avatar through the floating window 1301. The default display position of the floating window 1301 is the lower right corner of the monitor 260 screen, and its display size can be automatically adapted based on the playback window of the media asset video or manually adjusted by the user to avoid obscuring the main content area of the video.
[0152] In some embodiments, when the display device 200 displays a virtual avatar, it can control the virtual avatar to display corresponding expressions and actions based on the content of the media asset video to enhance its vividness. For example, if laughter is detected in the audio content of the media asset video, the virtual avatar's expressions and actions will be controlled to present a preset happy posture. If horror content is detected in the visual content of the media asset video, the virtual avatar's expressions and actions will be controlled to present a preset terrified posture.
[0153] In some embodiments, the display device 200 provides a virtual avatar editing function, allowing users to further adjust the virtual avatar after it has been generated. Specifically, the display device 200 can respond to editing instructions for the virtual avatar by controlling the display 260 to show an editing interface for interaction with the user. The editing interface includes at least one editing control and a preview area, with each editing control corresponding to at least one attribute of the virtual avatar.
[0154] The display device 200 can adjust the attributes of the virtual avatar according to the attribute parameters input by the user based on the editing control, and control the display 260 to display the adjusted virtual avatar in the preview area.
[0155] In some embodiments, when the display device 200 displays the editing interface, it can pause the playback of the media asset video, control the display 260 to display the editing interface in full screen, and after confirming that the user has completed the virtual avatar editing, cancel the display of the editing interface and continue playing the media asset video, ensuring that the user's choices are not affected by changes in the video content.
[0156] In some embodiments, in order to maintain the continuity of viewing, the display device 200 may control the display 260 to display the editing interface on top of the media asset video without pausing the playback of the media asset video, according to a preset size and preset position.
[0157] For example, such as Figure 14The diagram shows a schematic of the editing interface provided in this embodiment. The display device 200 displays a small editing interface 1400 on the right side of the monitor 260 screen. The editing interface 1400 includes editing controls 1401 for attributes such as clothing color, hairstyle volume, and facial proportions corresponding to the virtual character, as well as a preview area 1402 for real-time display of the virtual character. The user can input attribute parameters through any of the editing controls 1401, and the display device 200 will re-render the virtual character based on the new attribute parameters and display it through the preview area 1402. Examples include changing the clothing color from red to blue, adding hair accessories, and adjusting the eye magnification ratio by 10%.
[0158] In some embodiments, the display device 200 can export the generated virtual image as media content in multiple formats. After generating the virtual image, the display device 200 can control the display 260 to display the export interface in response to the export command of the virtual image. The export interface includes at least one export control, and each export control corresponds to the export format of a virtual image.
[0159] The display device 200 can determine the target export format corresponding to the target export control selected by the user. The target export control is any export control in the export interface. Based on the virtual image, the display device 200 generates media content corresponding to the target export format and outputs the media content.
[0160] The media asset content includes virtual images. Outputting media asset content refers to sending the media asset content to the target location, including but not limited to saving the media asset content to the local storage of the display device 200; sending the media asset content to the mobile terminal 300 that is communicatively connected to the display device 200; and uploading the media asset content to cloud storage.
[0161] For example, such as Figure 15 The diagram shows an export interface provided in an embodiment of this application. The export interface 1500 is displayed as a floating window on the right side of the monitor 260 screen. The export interface 1500 includes multiple export controls 1501, such as high-resolution image controls, animated emoticon controls, and short video controls. The high-resolution image control is used to export high-resolution static images, suitable for setting as avatars or wallpapers. The animated emoticon control is used to export GIF animations containing looping animations (such as blinking or smiling), suitable for social chat. The short video control is used to export short videos of preset duration (such as 3 seconds or 5 seconds) with custom actions (such as waving or turning around), suitable for sharing on short video platforms.
[0162] If a user selects the animated emoticon control via remote control or touch, the display device 200 responds to the selection operation of the animated emoticon control, determines the target export format as GIF, and generates a GIF animation corresponding to the currently generated virtual image.
[0163] In some embodiments, the display device 200 can work with the mobile terminal 300 and the server 400 to generate a virtual avatar. For example... Figure 16 The diagram shown is a timing diagram for the generation of a virtual avatar provided in an embodiment of this application. When playing media video, the display device 200 can collect user voice and perform speech recognition using Automatic Speech Recognition (ASR) technology to convert the user's voice into text, and then send the recognized text to the server 400.
[0164] Server 400 performs semantic understanding on the text sent by display device 200 to analyze user intent. If the user intent indicates the generation of a virtual avatar, it generates a first instruction and sends the first instruction to display device 200.
[0165] In response to the first instruction, the display device 200 extracts the corresponding video frame from the media asset video and sends a first notification to the mobile terminal 300, which instructs the mobile terminal 300 to upload the user's facial information.
[0166] In response to the first notification, the mobile terminal 300 interacts with the user to obtain the user's facial information and sends the user's facial information to the display device 200.
[0167] Display device 200 sends the facial information sent by mobile terminal 300 to server 400.
[0168] Server 400 performs image processing on video frames extracted from media assets and user facial information. Non-facial features of the target person are extracted from the video frames, and facial features of the user are extracted from the facial information. Based on the facial features and a head model library, a basic head model is generated, and non-facial features are fused onto the basic head model to obtain a virtual avatar. Finally, the virtual avatar is sent to display device 200.
[0169] After receiving the virtual avatar, the display device 200 controls the monitor 260 to display the virtual avatar on top of the media video. It can also further display an editing interface for the virtual avatar, allowing users to adjust it (such as changing hairstyle or clothing). After the user completes the editing, the display device 200 can apply the edited virtual avatar to various scenarios, such as using it as a voice assistant avatar. It can also send it to the mobile terminal 300 in a specific format (such as a static image or a GIF animation).
[0170] In some embodiments, such as Figure 17 The diagram shows a module representation of the edge-cloud collaborative architecture provided in this embodiment. The display device 200 includes a speech recognition module, a video frame extraction module, a display module, and a data receiving and reporting module. The mobile terminal 300 includes an image acquisition module. The server 400 includes a semantic understanding module, an image processing module, and a header model library.
[0171] The system includes the following modules: a speech recognition module to receive and process user speech, converting the audio signal into digital text; a video frame extraction module to handle operations related to video footage, extracting corresponding video frames from the currently playing media asset video stream; a display module to handle the rendering and display of the user interface, such as displaying the video frame selection interface, character selection interface, virtual avatar editing interface, and export interface; and a data receiving and reporting module for data communication with the mobile terminal 300 and the server 400, such as receiving facial information uploaded by the mobile terminal 300; sending video frames and facial information to the server 400; and receiving virtual avatars from the server 400.
[0172] The image acquisition module is used to acquire facial images and / or facial videos of the user and send them to the display device 200.
[0173] The semantic understanding module is used to understand the text forwarded by the user's voice and determine the user's intent. The image processing module is used to perform image processing to generate a virtual avatar. The head model library is used to store head models.
[0174] Based on the aforementioned display device 200, this application embodiment also provides a method for generating a virtual image, the method including the following steps: In response to the first instruction to generate the virtual avatar, determine the first playback time of the media asset video upon receiving the first instruction.
[0175] Extract the first video frame corresponding to the first playback time point from the media asset video.
[0176] Extract the non-facial features of the target person in the first video frame. The target person is the person appearing in the first video frame.
[0177] Obtain the user's facial features.
[0178] A basic head model is generated based on facial features.
[0179] By fusing non-facial features onto a basic head model, a virtual avatar is obtained.
[0180] A virtual avatar is displayed on top of the media asset video.
[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0182] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.
Claims
1. A display device, characterized in that, include: The display is configured to display a user interface, the user interface including media asset video; The controller is configured as follows: In response to a first instruction to generate a virtual avatar, the first playback time point of the media asset video is determined upon receiving the first instruction; Extract the first video frame corresponding to the first playback time point from the media asset video; Extract the non-facial features of the target person in the first video frame; The target character is the character appearing in the first video frame; Obtain the user's facial features; Based on the facial features, a basic head model is generated; The non-facial features are fused onto the base head model to obtain a virtual image; The display is controlled to show the virtual image on top of the media asset video.
2. The display device according to claim 1, characterized in that, The controller extracts non-facial features of the target person in the first video frame and is configured as follows: According to the first instruction, obtain descriptive information about the target person's image; Based on the description information, the image of the target person is identified in the first video frame; Extract the target image region containing the image of the target person; The target image region is input into the non-facial feature extraction model to extract non-facial features and obtain the non-facial features; the non-facial feature extraction model is a neural network model pre-trained based on sample images labeled with non-facial features.
3. The display device according to claim 1, characterized in that, The controller acquires the user's facial features and is configured as follows: Obtain the user's facial information; the facial information includes facial images and / or facial videos; The facial information is input into the facial feature extraction model to extract facial features and obtain the facial features; the facial feature extraction model is a neural network model pre-trained based on sample images labeled with facial feature tags.
4. The display device according to claim 1, characterized in that, The controller generates a basic head model based on the facial features and is configured as follows: Call the header model library; the header model library includes multiple preset header models; Calculate the matching degree between the facial features and each of the head models; The target head model is determined based on the matching degree between the facial features and each head model; The target head model is a head model in the head model library whose matching degree with the facial features is greater than a preset matching degree threshold. The facial features are fused onto the target head model to obtain the base head model.
5. The display device according to claim 3, characterized in that, It also includes a first communication device and a second communication device; the first communication device is configured to establish a communication connection with a mobile terminal; the second communication device is configured to establish a communication connection with a server; After the controller extracts the first video frame corresponding to the first playback time point from the media asset video, it is configured to: Send a first notification to the mobile terminal, the first notification being used to instruct the mobile terminal to upload the user's facial information; Receive the user's facial information sent by the mobile terminal; The first video frame and the facial information are sent to the server so that the server generates the virtual image based on the first video frame and the facial information; Receive the virtual image sent by the server.
6. The display device according to claim 1, characterized in that, The controller, after fusing the non-facial features onto the base head model to obtain the virtual avatar, is configured as follows: In response to a user's editing command for the virtual avatar, the display is controlled to show an editing interface; the editing interface includes at least one editing control and a preview area, and each editing control corresponds to at least one attribute of the virtual avatar; The attributes of the virtual avatar are adjusted based on the attribute parameters input by the user using the editing control. Control the display to show the adjusted virtual image in the preview area.
7. The display device according to claim 1, characterized in that, The controller, after fusing the non-facial features onto the base head model to obtain the virtual avatar, is configured as follows: In response to the export command of the virtual avatar, control the display to show the export interface; The export interface includes at least one export control, and each export control corresponds to the export format of a virtual image; Determine the target export format corresponding to the target export control selected by the user; The target export control is any one of the export controls in the export interface; Based on the virtual avatar, media asset content corresponding to the target export format is generated; the media asset content includes the virtual avatar. Output the media asset content.
8. The display device according to claim 1, characterized in that, After determining the first playback time of the media asset video upon receiving the first instruction, the controller is further configured to: Based on the first playback time point, the target duration is traced back to obtain the second playback time point of the media asset video; the target duration is a preset duration or the duration required to receive the first instruction. Extract the second video frame corresponding to the second playback time point from the media asset video; Extract the non-facial features of the target person in the second video frame.
9. The display device according to claim 1, characterized in that, After determining the first playback time of the media asset video upon receiving the first instruction, the controller is further configured to: In the media video, a first video segment within a first duration before the first playback time point and a second video segment within a second duration after the first playback time point are extracted. Extract a third video frame from the first video segment and the second video segment. The third video frame is a video frame that includes the image of the target person. Extract the non-facial features of the target person in the third video frame.
10. A method for generating a virtual avatar, characterized in that, include: In response to the first instruction to generate a virtual avatar, determine the first playback time point of the media asset video upon receiving the first instruction; Extract the first video frame corresponding to the first playback time point from the media asset video; Extract the non-facial features of the target person in the first video frame; The target character is the character appearing in the first video frame; Obtain the user's facial features; Based on the facial features, a basic head model is generated; The non-facial features are fused onto the base head model to obtain a virtual image; The virtual avatar is displayed on top of the media asset video.