Display apparatus and display method of virtual object
By extracting user gesture features in smart TV, generating hand shadow images and obtaining virtual object images, the problem of difficulty in animation generation in smart TV is solved, and entertainment and interactivity are improved.
Patent Information
- Application Number
- CN202510158027.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-24
AI Technical Summary
How to reduce the difficulty of animation generation in smart TVs, improve entertainment and interactivity, and develop more functions through user gestures.
A display method for displaying a display device and a virtual object is provided, which captures images by shooting components, extracts gesture features, generates hand shadow images, and acquires virtual object images based on hand shadow images, and displays them on the screen of a display.
The user can not only see hand shadow images formed by light transmission, but also see virtual character images corresponding to gestures and hand shadow images, enhancing the entertainment and interactiveness of smart TVs.
Smart Images

Figure CN120201219A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer application technologies, and particularly to a display device and a method for displaying virtual objects. Background Art
[0002] Smart TVs can capture user gestures and support users to perform operation controls through gestures, such as switching channels, adjusting volume, playing and pausing, etc. As the market increasingly demands to enhance the entertainment of smart TVs, how to develop more functions for user gestures to enhance the entertainment and interactivity of smart TVs remains to be solved urgently. Summary of the Invention
[0003] Based on this, in view of the above technical problems, it is necessary to provide a display device and a method for displaying virtual objects that can reduce the difficulty of generating animations.
[0004] In a first aspect, this application provides a display device, including:
[0005] A display configured to display images and / or user interfaces;
[0006] A shooting component for shooting images;
[0007] A controller configured to:
[0008] For the image captured by the shooting component, according to the gesture features extracted from the image, generate a hand shadow image for representing the corresponding projection of the gesture; wherein, the gesture features at least include one of the shape of the gesture contour, the line curvature of the gesture contour, or the area of the closed region formed by the gesture contour, and the hand shadow image carries the gesture features;
[0009] Based on the hand shadow image, obtain a virtual object image; wherein, the contour features of the virtual object in the virtual object image are associated with the gesture features carried by the hand shadow image;
[0010] Display the hand shadow image on the display screen of the display, and display the virtual object image corresponding to the hand shadow image on the display screen.
[0011] The above technical solution has the following advantages or beneficial effects: When the image captured by the shooting component contains a gesture, gesture features are extracted, and a hand shadow image is generated based on the gesture features and the hand shadow image is displayed on the display screen. Further, a virtual object image is also obtained and displayed based on the hand shadow image. That is to say, through gesture interaction, the user can not only see a hand shadow image formed by light transmission on the display screen of the display, but also see a virtual character image corresponding to the gesture and the hand shadow image. Thus, this embodiment can develop a hand shadow display function and a virtual object display function for user gestures, increasing the entertainment and interactivity of the display device.
[0012] In a second aspect, the present application provides a method for displaying a virtual object, which is applied to the display device as described in the first aspect. The method includes:
[0013] For the image captured by the shooting component, according to the gesture features extracted from the image, a hand shadow image is generated for representing the corresponding projection of the gesture; wherein, the gesture features at least include one of the shape of the gesture contour, the line curvature of the gesture contour, or the area of the closed region formed by the gesture contour, and the hand shadow image carries the gesture features;
[0014] Based on the hand shadow image, a virtual object image is obtained; wherein, the contour features of the virtual object in the virtual object image are associated with the gesture features carried by the hand shadow image;
[0015] The hand shadow image is displayed on the display screen of the display, and the virtual object image corresponding to the hand shadow image is displayed on the display screen.
[0016] The above technical solution has the following advantages or beneficial effects: When the image captured by the shooting component contains a gesture, gesture features are extracted, and a hand shadow image is generated based on the gesture features and the hand shadow image is displayed on the display screen. Further, a virtual object image is also obtained and displayed based on the hand shadow image. That is to say, through gesture interaction, the user can not only see a hand shadow image formed by light transmission on the display screen of the display, but also see a virtual character image corresponding to the gesture and the hand shadow image. Thus, this embodiment can develop a hand shadow display function and a virtual object display function for user gestures, increasing the entertainment and interactivity of the display device. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the accompanying drawings required for the description in the embodiments of the present application or the related art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related accompanying drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 Schematic diagram of an operation scenario between a display device and a control device provided by some embodiments of the present application;
[0019] Figure 2 Schematic diagram of the hardware configuration of a display device provided by some embodiments of the present application;
[0020] Figure 3 Schematic diagram of the hardware configuration of a control device provided by some embodiments of the present application;
[0021] Figure 4 Schematic diagram of the software configuration of a display device provided by some embodiments of the present application;
[0022] Figure 5 Schematic diagram of the system architecture provided by some embodiments of the present application;
[0023] Figure 6 Schematic diagram of a hand shadow image displayed on the display screen provided by some embodiments of the present application;
[0024] Figure 7 Schematic diagram of a virtual object image displayed on the display screen provided by some embodiments of the present application;
[0025] Figure 8 Schematic diagram of the circumscribed rectangular area of the hand shadow image provided by some embodiments of the present application;
[0026] Figure 9 Flow chart of the display method of a virtual object in one embodiment;
[0027] Figure 10 Flow chart of the display method of a virtual object in another embodiment;
[0028] Figure 11 Flow chart of the display method of a virtual object in yet another embodiment;
[0029] Figure 12 Flow chart of the display method of a virtual object in another embodiment;
[0030] Figure 13 Interaction flow chart between a camera component, a controller, a display, and a cloud server in one embodiment;
[0031] Figure 14 It is a structural block diagram of a display device for a virtual object in an embodiment. Detailed implementation manners
[0032] The embodiments will be described in detail below, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present application. They are merely examples of systems and methods consistent with some aspects of the present application detailed in the claims.
[0033] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the following described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.
[0034] The terms "first", "second", "third", etc. in the specification, claims and the above drawings of the present application are used to distinguish similar or same-kind objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms used can be interchanged under appropriate circumstances.
[0035] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclusively include. For example, a product or device comprising a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0036] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or a combination of hardware or / and software code that can perform functions related to the element.
[0037] In the embodiments of the present application, the display device 200 generally refers to a device having the capabilities of screen display and data processing. For example, the display device 200 includes but is not limited to smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.
[0038] Figure 1 It is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. As Figure 1 shown, the user can operate the display device 200 through touch operations, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote control, a stylus, a gamepad, etc.
[0039] The mobile terminal 300 can be used as a control device to perform human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device to establish a communication connection with the display device 200 for data interaction. In some embodiments, software applications can be installed on the mobile terminal 300 and the display device 200, and the connection communication can be achieved through network communication protocols to achieve the purpose of one-to-one control operations and data communication. It is also possible to transmit the audio and video content displayed on the mobile terminal 300 to the display device 200 to achieve the synchronous display function.
[0040] As Figure 1 also shown in, the display device 200 also communicates with the server 400 through various communication methods. The display device 200 is allowed to establish a communication connection through a local area network (LAN), a wireless local area network (WLAN), and other networks.
[0041] The display device 200 can provide a broadcast receiving television function, and can also additionally provide an intelligent network television function with computer support functions, including but not limited to, network television, smart television, Internet Protocol Television (IPTV), etc.
[0042] Figure 2 For some embodiments of this application Figure 1 The hardware configuration block diagram of the display device 200 in.
[0043] In some embodiments, the display device 200 may include at least one of a tuner demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0044] In some embodiments, the detector 230 is used to collect signals from the external environment or interact with the outside. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of environmental light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures. Or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.
[0045] In some embodiments, the display 260 includes a display function component for presenting a picture and a driving component for driving image display. The display 260 is used to receive the image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, components of a menu manipulation interface, and a user manipulation UI interface, etc.
[0046] In some embodiments, the communication device 220 is a component for communicating with an external device or server 400 according to various communication protocol types. The display device 200 may be provided with a plurality of communication devices 220 according to different supported communication methods. For example, when the display device 200 supports wireless network communication, the display device 200 may be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.
[0047] The communication device 220 can enable the display device 200 to communicate with an external device or server 400 through a wireless or wired connection. Among them, the wired connection can connect the display device 200 with an external device through components such as data lines and interfaces. The wireless connection can connect the display device 200 with an external device through wireless signals or wireless networks. The display device 200 can directly establish a connection relationship with an external device, or can indirectly establish a connection relationship through a gateway, router, connection device, etc.
[0048] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and first to n interfaces for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.
[0049] In some embodiments, the controller 250 and the tuner demodulator 210 may be located in different split devices, that is, the tuner demodulator 210 may also be in an external device of the main device where the controller 250 is located, such as an external set-top box, etc.
[0050] In some embodiments, the user can input a user command on the graphical user interface (GUI) displayed on the display 260, and then the user input interface receives the user input command through the graphical user interface (GUI).
[0051] In some embodiments, the audio output device 270 may be a built-in speaker of the display device 200, or may be an external audio output device connected to the display device 200. Among them, for the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device can be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.
[0052] In some embodiments, the user input interface 280 can be used to receive instructions from user input.
[0053] Figure 3 For some embodiments of the present application Figure 1 The hardware configuration block diagram of the control device in. As Figure 3 shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.
[0054] The control device 100 is configured to control the display device 200, and can receive user input operation instructions, and convert the operation instructions into instructions recognizable and responsive by the display device 200, playing the role of an interaction intermediary between the user and the display device 200.
[0055] In some embodiments, the control device 100 may be an intelligent device. For example: the control device 100 can install various applications for controlling the display device 200 according to user needs.
[0056] In some embodiments, as Figure 1 shown, the mobile terminal 300 or other intelligent electronic devices, after installing the application for controlling the display device 200, can perform functions similar to those of the control device 100.
[0057] The controller 110 includes a processor 112, a RAM 113, a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between internal components and the data processing functions between the external and internal.
[0058] Under the control of the controller 110, the communication interface 130 realizes the communication of control signals and data signals with the display device 200. The communication interface 130 may include at least one of a WiFi chip 131, a Bluetooth module 132, an NFC module 133, and other near-field communication modules.
[0059] The user input / output interface 140, where the input interface includes at least one of a microphone 141, a touchpad 142, a sensor 143, a key 144, and other input interfaces.
[0060] In some embodiments, the control device 100 includes at least one of the communication interface 130 and the input / output interface 140. The communication interface 130 is configured in the control device 100. For example: modules such as WiFi, Bluetooth, and NFC can encode user input instructions through the WiFi protocol, or the Bluetooth protocol, or the NFC protocol, and send them to the display device 200.
[0061] The memory 190 is used to store various operating programs, data, and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.
[0062] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.
[0063] In order to perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system may provide a user interface (control the display device), allow the user to interact with the display device 200, and support the running of various application programs.
[0064] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.
[0065] The operating system can be divided into different modules or layers according to the functions implemented, for example Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (Applications) layer (referred to as "application layer"), the application framework layer (Application Framework) layer (referred to as "framework layer"), the system library layer and the kernel layer.
[0066] In some embodiments, the application layer is used to provide services and interfaces for applications so that the display device 200 can run applications and interact with users based on the applications. At least one application can be run in the application layer, and these applications can be window programs, system settings programs, clock programs, etc. that come with the operating system; they can also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.
[0067] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center that determines the actions that applications in the application layer take. Through the API interface, applications can access system resources and obtain system services during execution.
[0068] like Figure 4As shown, in the embodiments of the present application, the application framework layer includes a view system, managers, content providers, etc. Among them, the view system can design and implement the interfaces and interactions of application programs. The view system includes lists, grids, text boxes, buttons, etc. The managers include at least one of the following modules: The Activity Manager is used to interact with all the activities running in the system; the Location Manager is used to provide access to the system location service for system services or applications; the Package Manager is used to retrieve various information related to the application program packages currently installed on the device; the Notification Manager is used to control the display and clearing of notification messages; the Window Manager is used to manage icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0069] In some embodiments, the Activity Manager is used to manage the life cycles of various application programs and the usual navigation back functions, such as controlling the exit, opening, and backward movement of application programs. The Window Manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, taking screenshots, and controlling the changes in the display window. For example, shrinking the display window, jittering the display, distorting the display, etc.
[0070] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction libraries included in the system runtime layer, such as C / C++ instruction libraries, to implement the functions that the framework layer is to achieve.
[0071] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, as Figure 4 shown, hardware drivers can be configured in the kernel layer. The drivers included in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor drivers (such as fingerprint sensors, temperature sensors, pressure sensors, etc.), and power drivers, etc.
[0072] It should be noted that the above examples are only simple divisions of the operating system functions, and do not limit the specific form of the operating system of the display device 200 in the embodiments of the present application. Depending on factors such as the functions of the display device and the type of the operating system, the number of hierarchical levels and the specific hierarchical types included in the operating system can be in other forms.
[0073] Combined with the specific operating system of the display device 200 in the embodiments of the present application, the system architecture diagram of the display device 200 can be referred to Figure 5 . As Figure 5 shown, in some embodiments, the system is divided into four layers, from top to bottom are the Application layer (abbreviated as "application layer"), the Application Framework layer (abbreviated as "framework layer"), the core layer of the input subsystem, and the driver layer.
[0074] Among them, the application layer can provide user interfaces and user interaction functions, including camera startup, the main interface of the shadow puppetry game, virtual object display, control panel, display of gesture recognition results, playing virtual object animations, and processing user inputs, etc. Processing user inputs includes starting the game, adjusting sensitivity, and viewing the game status, etc. Combined with Figure 5 , in some embodiments, the application layer can run the application program "AI Shadow Puppetry Paintings". This application program can further include a shadow puppetry capture program, a camera program, an interaction interface program, an AI animation display program, etc.
[0075] The framework layer provides a running environment for core application logics, such as real-time analysis of image data, gesture recognition, animation control logic, etc. In addition, it can also manage the interaction between applications and system services, such as camera services, sensor data acquisition, and cloud communication modules, etc. Combined with Figure 5 , in some embodiments, the framework layer can further include a camera manager, a camera device, an alarm manager, an activity management service, a package management service, and a window management service.
[0076] The core layer of the input subsystem can provide system services accessible to applications, such as call interfaces for Camera services, network communication services, and image processing libraries, etc. Combined with Figure 5 , in some embodiments, the core layer can include an input core driver, such as "Driver / input / input.c".
[0077] The driver layer of the input subsystem can provide underlying hardware abstraction and driver support, including camera drivers, GPU acceleration, and image processing support, etc. The driver layer can support the hardware decoding and rendering tasks of the device, and optimize the processing speed and efficiency of image data. Combined with Figure 5, in some embodiments, the driver layer may include a touch screen driver (such as S3C2410TS.C) and a USB keyboard driver (such as USBKBD.C), etc.
[0078] Combined with the foregoing content, embodiments of the present application provide a display device, including: a display configured to display images and / or user interfaces; a shooting component for shooting images; a controller configured to: for the image captured by the shooting component, generate a hand shadow image for characterizing the corresponding projection of the gesture according to the gesture features extracted from the image; wherein, the gesture features at least include one of the shape of the gesture contour, the line curvature of the gesture contour, or the area of the closed region formed by the gesture contour, and the hand shadow image carries the gesture features; based on the hand shadow image, obtain a virtual object image; wherein, the contour features of the virtual object in the virtual object image are associated with the gesture features carried by the hand shadow image; display the hand shadow image on the display screen of the display, and display the virtual object image corresponding to the hand shadow image on the display screen.
[0079] Wherein, the shooting component may be built into the display device or externally connected to the display device. The shooting component may be a camera. The image captured by the shooting component includes the user's gesture. The image captured by the shooting component can be understood as each frame of image captured by the shooting component. Use the Camera2 API (Application Programming Interface, Chinese name Application Programming Interface) or the camera SDK (Software Development Kit) to capture a real-time video stream.
[0080] Wherein, the gesture features include the shape of the gesture contour, and the shape of the gesture contour may present various forms such as a circle, a square, an irregular polygon, etc. Wherein, the gesture features include the line curvature of the gesture contour. The line curvature refers to the rotation rate of the tangent direction angle of a point on the curve with respect to the arc length, which is defined by differentiation and is used to indicate the degree to which the curve deviates from a straight line. Mathematically, it is a specific numerical value used to quantify the degree of curvature of the curve at a certain point. The greater the line curvature, the greater the degree of curvature of the curve. The smaller the line curvature, the closer the curve is to a straight line. The line curvature is used to describe whether the line of the gesture contour is gently curved or sharply turning. Wherein, the gesture features include the area of the closed region formed by the gesture contour. The area of the closed region is the numerical value of the closed space enclosed by the lines of the gesture contour.
[0081] Among them, in the field of light propagation, a hand shadow image is a shadow image generated on a wall or other plane under the illumination of a light source through the shape and movement of a hand. In this embodiment, after an image is captured by the shooting component, the image is processed to form a hand shadow image. The generated hand shadow image is similar to the hand shadow image formed under the illumination of a light source, and the generated hand shadow image is used to represent the corresponding projection of a gesture.
[0082] According to the gesture features, the process of generating a hand shadow image can be to first determine the gesture contour in the image according to the gesture features, and then crop out the hand shadow image from the image based on the gesture contour. Among them, in the process of determining the gesture contour in the image, the image can be first processed into a grayscale image, then the gesture contour is determined based on the grayscale image, and finally the hand shadow image is cropped out from the grayscale image based on the gesture contour. Processing the image into a grayscale image first can reduce the image data to be processed and improve the generation speed of the hand shadow image.
[0083] Among them, the virtual object image includes a virtual object. The type of the virtual object can be an animal, a plant, a building, etc., which is not limited in this embodiment. The type of the virtual object displayed on the display screen can be one of an animal, a plant or a building. When the image captured by the shooting component contains at least two independent gestures, the virtual objects corresponding to different gestures can be of the same type of virtual object or different types of virtual objects. For example, the types of the virtual objects corresponding to two gestures are both animals. For example, the type of the virtual object corresponding to one of the two gestures is an animal, and the type of the virtual object corresponding to the other gesture is a plant.
[0084] Optionally, virtual animals can be generated based on common hand shadow shapes. For example, when the hand presents a contour similar to the flapping of a bird's wings, that is, the fingers are extended and slightly bent, simulating the posture of a bird's wings flapping, the intention of the hand movement is captured, and then a bird character flying freely in the air with vivid feathers is generated. The fingers are bent to simulate the shape of a rabbit's ears, the index finger and the middle finger are erected and bent inward, and the rest of the fingers are clenched into a fist, just like the lively erected ears of a rabbit, and a lively and lovely rabbit image jumping is correspondingly generated.
[0085] Optionally, virtual characters with fantasy colors can also be generated, such as elves with magic special effects triggered by complex gesture combinations. These complex gesture combinations may be that the ten fingers of both hands quickly cross and change positions, and at the same time, a faint light flashes in the palms. The system accurately analyzes these unique hand shadow features with advanced recognition technology, and then flexibly constructs elves with different shapes and surrounded by dazzling light. They either hold a magic wand or have wings on their backs, full of fantasy charm. The system flexibly constructs the virtual character form according to the recognized hand shadow features, aiming to meet the rich and diverse creative expression needs of users and enhance the interactive entertainment experience based on smart TVs.
[0086] It should be noted that in the case where the image captured by the shooting component contains at least two independent gestures, a hand shadow image and a virtual object image corresponding to the hand shadow image can be generated for each independent gesture, or the independent gestures can be associated as a combined gesture, and a hand shadow image and a virtual object image corresponding to the hand shadow image can be generated for the combined gesture. It should be noted that the at least two independent gestures described refer to that there is no contact point between at least two gestures. Correspondingly, there is also no connection point between the hand shadow images corresponding to each gesture.
[0087] Among them, the hand shadow image carries the gesture feature. Obtaining the virtual object image based on the hand shadow image is mainly to generate the virtual object image according to the gesture feature. The contour feature of the virtual object in the virtual object image is associated with the gesture feature carried by the hand shadow image. For example, if the user's hand presents a contour similar to a bird spreading its wings, then the virtual object in the virtual object image generated according to the gesture feature is a bird. Correspondingly, the contour feature of the virtual object is used to represent that the contour of the virtual object is the contour of a bird spreading its wings, and the contour feature of the virtual object is associated with the gesture feature. The contour feature of the virtual object may include one of the shape of the virtual object contour, the line curvature of the virtual object contour, or the area of the closed region formed by the virtual object contour. The gesture feature constitutes the basic information closely related to the construction of the virtual character form, ensuring that the generated virtual object can correspond to the hand shadow action in terms of appearance.
[0088] In one embodiment, a deep learning algorithm can be used to classify and recognize the gesture features, map the recognized gestures to predefined virtual objects, and further generate virtual object images. Correspondingly, before classifying and recognizing the gesture features, a database containing various gestures and corresponding virtual objects needs to be constructed.
[0089] Among them, the virtual object image can be a two-dimensional image or a three-dimensional image, which is not limited in this embodiment.
[0090] Among them, the display position of the hand shadow image can be fixed, or it can correspond to the position of the gesture in the image captured by the shooting component and move with the movement of the gesture position, which is not limited in this embodiment.
[0091] Among them, the display positions of the virtual object image and the corresponding hand shadow image can be the same or different, which is not limited in this embodiment.
[0092] Among them, in the case of gesture change, it is equivalent to the shooting component capturing a new image, and then a new hand shadow image and a virtual object image are generated for the new image.
[0093] Among them, the schematic diagram of displaying the hand shadow image on the display screen is as follows Figure 6 shown. The outline of the hand shadow image is consistent with the outline of the user's gesture. Among them, the schematic diagram of displaying the virtual object image on the display screen is as follows Figure 7 shown. It can be seen that the contour features of the virtual objects corresponding to different gestures are different.
[0094] For the display device provided in this embodiment, when the image captured by the shooting component contains a gesture, the gesture feature is extracted, a hand shadow image is generated based on the gesture feature, and the hand shadow image is displayed on the display screen. Further, a virtual object image is also obtained and displayed based on the hand shadow image. That is to say, through gesture interaction, the user can not only see a hand shadow image formed by light transmission on the display screen of the display, but also see a virtual character image corresponding to the gesture and the hand shadow image. In this way, this embodiment can develop a hand shadow display function and a virtual object display function for user gestures, increasing the entertainment and interactivity of the display device.
[0095] In some embodiments, the controller is configured to obtain a virtual object image based on the hand shadow image, and is configured to: input the hand shadow image carrying the gesture feature into a virtual object generation model, and obtain the virtual object image generated by the virtual object generation model; and / or, send the hand shadow image carrying the gesture feature to a cloud server, and obtain the virtual object image generated and sent by the cloud server based on the virtual object generation model.
[0096] Among them, the hand shadow image is processed by a deep learning algorithm. The virtual object generation model can be a Generative Adversarial Networks (GANs) or a Variational Autoencoders (VAE), or other models capable of generating virtual object images, which are not limited in this embodiment. The virtual object generation model can be set in the display device or the cloud server, or can be set in both the display device and the cloud server.
[0097] Correspondingly, the hand shadow image carrying the gesture feature can be input into the virtual object generation model set in the display device, and the virtual object image generated by the virtual object generation model can be obtained.
[0098] Correspondingly, the hand shadow image carrying the gesture feature can also be sent to the cloud server, and the virtual object image generated and sent by the cloud server based on the virtual object generation model can be obtained. Specifically, the cloud server inputs the hand shadow image into the virtual object generation model, obtains the virtual object image generated by the virtual object generation model, and sends the virtual object image to the display device. After receiving the virtual object image sent by the cloud server, the display device can load the virtual object image through a third-party library such as Glide or a tool such as BitmapFactory, and then display the virtual object image on the display screen.
[0099] Inputting the hand shadow image carrying the gesture feature into the virtual object generation model can obtain the virtual object image output by the virtual object generation model.
[0100] The display device provided in this embodiment can improve the generation efficiency of the virtual object image by utilizing the computing power of the cloud server through sending the hand shadow image to the cloud server for processing. In addition, compared with directly sending the image captured by the shooting component to the cloud server, in this embodiment, the processed hand shadow image is sent to the cloud server, and the data volume of the hand shadow image is significantly smaller than that of the image captured by the shooting component. Therefore, by combining local processing and cloud server computing, this embodiment can improve the system response speed and reduce the excessive dependence on cloud resources while ensuring high-quality image generation.
[0101] In some embodiments, the controller is configured to display the virtual object image corresponding to the hand shadow image on the display screen as follows: determining the area of the hand shadow region based on the number of pixel points included in the hand shadow region in the hand shadow image; determining the display ratio of the virtual object in the virtual object image based on the aspect ratio and the area of the circumscribed rectangular region of the hand shadow region; and displaying the virtual object image based on the display ratio of the virtual object.
[0102] Among them, the hand shadow region in the hand shadow image can be understood as the closed region formed by the gesture contour. The area of the hand shadow region can be determined according to the number of pixel points included in the hand shadow region.
[0103] Among them, please refer to Figure 8 , the circumscribed rectangular region of the hand shadow region (region A marked in the figure) can be understood as the circumscribed rectangular region of the gesture contour. The circumscribed rectangular region can be understood as the smallest rectangle that just contains all the points of the hand shadow region or the gesture contour.
[0104] Among them, the aspect ratio of the circumscribed rectangular area refers to the ratio of the length to the width of the circumscribed rectangular area, which is used to reflect the shape characteristics of the hand shadow area. Generally speaking, through this aspect ratio, it can be determined whether the hand shadow area is slender or short and stout.
[0105] Among them, according to the aspect ratio and the area of the hand shadow area, the display ratio (i.e., the scaling ratio) of the virtual object image can be determined, and it can be specified how much the virtual object image should be enlarged or reduced for display.
[0106] The display device provided in this embodiment displays the virtual object image based on the display ratio, ensuring that the display effect of the generated virtual object on the display screen can perfectly match the hand shadow image (or the amplitude of the hand shadow movement), and is also visually reasonable.
[0107] In some embodiments, the controller is configured to display the hand shadow image on the display screen of the display, and display the virtual object image corresponding to the hand shadow image on the display screen, and is configured to: determine the display position of the hand shadow image on the display screen based on the gesture position extracted from the image; based on the display position, display the hand shadow image on the display screen, and based on the display position, display the virtual object image on the display screen.
[0108] Among them, the gesture position can be understood as the position coordinates of the gesture in the image. The display position (i.e., the position coordinates) of the hand shadow image on the display screen can be the same as or close to the position coordinates of the gesture in the image.
[0109] Optionally, the mapping relationship between the gesture position and the display position of the hand shadow image can be set in advance, and after determining the gesture position, the display position of the hand shadow image is determined based on the mapping relationship.
[0110] Among them, after determining the display position of the hand shadow image, the hand shadow image is displayed based on the display position, and the virtual object image is displayed based on the display position. That is to say, the display position of the virtual object image is the same as that of the corresponding hand shadow image.
[0111] The display device provided in this embodiment determines the display position based on the gesture position extracted from the image, and then displays the hand shadow image and the virtual object image based on the display position. In this way, when the user waves the gesture, the hand shadow image and the virtual object image move with the wave of the gesture, giving the user a feeling of real-time projection, and enhancing the entertainment and interactivity of the display device.
[0112] In some embodiments, after the controller displays the virtual object image on the display screen based on the display position, it is further configured to: when the virtual object in the virtual object image is an animal, display the introduction information of the virtual object.
[0113] Among them, the introduction information can be displayed on the side of the virtual object image, or at the corner of the display screen (such as the upper left corner, lower left corner, upper right corner, lower right corner, etc.), or in other areas. This embodiment does not make a limitation. Preferably, the displayed introduction information does not block the virtual object image. Optionally, the displayed introduction information can also partially block or semi-transparently block the virtual object image.
[0114] Among them, the introduction information can include the name, subject, habits, etc. of the animal.
[0115] The display device provided in this embodiment can enhance the educational nature by displaying the introduction information, which is helpful for children's education assistance or family parent-child interactive entertainment.
[0116] In some embodiments, after the controller displays the virtual object image on the display screen based on the display position, it is further configured to: when there are multiple virtual objects displayed on the display screen, obtain the voting result reflecting the popularity of the virtual object and display it on the display screen.
[0117] Among them, the voting result corresponding to the virtual object can be displayed on the side of the virtual object. The voting results corresponding to multiple virtual objects can also be centrally displayed in one area. For example, it can be displayed at the corner of the display screen (such as the upper left corner, lower left corner, upper right corner, lower right corner, etc.), or in other areas. This embodiment does not make a limitation. Preferably, the voting result does not block the virtual object image. Optionally, the voting result can also partially block or semi-transparently block the virtual object image.
[0118] Among them, the process of obtaining the voting result can be to display a voting interface and display the information of all historical virtual objects and voting buttons on the voting interface. In response to the selected information of the voting button (such as the number of selections, selection duration, selection frequency, etc.), the voting result of the virtual object corresponding to the voting button is determined according to the selected information. For example, if the voting button of virtual object A is selected 5 times, the voting button of virtual object B is selected 3 times, and the voting button of virtual object C is selected 1 time, then the voting result of virtual object A is 5 votes, the voting result of virtual object B is 3 votes, and the voting result of virtual object C is 1 vote.
[0119] Among them, the voting result can be displayed in the form of a bar chart, pie chart, line chart or other graphics.
[0120] In one embodiment, the generation preference of virtual objects can be optimized according to the voting result. For example, if it is known from the voting result that the virtual object of a bird is the most popular, then in the subsequent generation process of virtual objects, the virtual object of a bird is preferentially generated.
[0121] The display device provided in this embodiment, when displaying multiple virtual objects on the display screen, displays the voting results of different virtual objects to reflect the popularity of different virtual objects, enhancing the interest of virtual object display.
[0122] In some embodiments, the controller executes the extraction process of the gesture feature and is configured to: convert the image into a grayscale image, perform binarization processing on the grayscale image to obtain a binarized image; based on a limb recognition algorithm, recognize the hand joint point information in the binarized image, and determine the gesture feature according to the hand joint point information.
[0123] Among them, the method of converting the image into a grayscale image can be the average method, weighted average method, maximum value method, minimum value method, brightness method, direct reading method or other methods, which are not limited in this embodiment. Among them, the average method is to use the average value of the red, green, and blue color components of each pixel as the grayscale value, and perform arithmetic averaging on the pixel values of the red, green, and blue channels, that is, Gray = (R + G + B) / 3, where R represents red, G represents green, and B represents blue, and the new grayscale value of each pixel point is determined through such calculation. Among them, the maximum value method is to use the largest color component in the pixel as the grayscale value, that is, directly take the maximum value among the red, green, and blue channels as the grayscale value of the pixel point, and the formula is expressed as Gray = max(R, G, B), where R represents red, G represents green, and B represents blue.
[0124] Among them, the minimum value method is to use the smallest color component in the pixel as the grayscale value.
[0125] In one embodiment, converting the image into a grayscale image can be for the pixel points in the image, and according to the color values of the pixel points in different color channels and the corresponding weights of the color channels, obtain the grayscale value of the pixel points. For example, different weights are assigned to the red (R), green (G), and blue (B) color channels according to the sensitivity of the human eye to different colors. Usually, the weight coefficients adopted are red channel weight 0.299, green channel weight 0.587, and blue channel weight 0.114. Then, calculate the grayscale value corresponding to each pixel point according to the formula Gray = 0.299×R + 0.587×G + 0.114×B, and assign the grayscale value to the original pixel point, thereby realizing the conversion from color to grayscale.
[0126] Through the above-mentioned grayscale operation, the image taken by the shooting component can remove redundant color information, and present it in a more concise data form on the basis of retaining the basic structure of the image and the key features of the hand, thereby laying a good data foundation for subsequent image processing processes such as binarization and feature extraction, and reducing problems such as excessive computational complexity and reduced processing efficiency that may arise in the subsequent processing process due to overly complex data.
[0127] Among them, binarization of grayscale images is to convert grayscale images into an image with only black and white colors. Binarization can remove noise in grayscale images and separate gestures from the background in the grayscale images for subsequent processing. The binarized images enable subsequent limb recognition algorithms to locate hand joints more efficiently and accurately, clearly identify gesture features such as finger distribution, and accurately determine gesture actions.
[0128] In one embodiment, during the binarization process of the grayscale image, the grayscale value of the pixel in the image is compared with a preset threshold value, and according to the comparison result, the grayscale image formed by the grayscale values of different pixels is converted into a binary image. Specifically, a specific threshold value is selected as the preset threshold value according to actual needs, and a corresponding binarization algorithm (such as a global threshold method) is used to compare the pixel value of each pixel in the grayscale image with the preset threshold value. If the pixel value is greater than the preset threshold value, the pixel point is assigned a pixel value representing white (in an 8-bit grayscale image, the pixel value is set to 255 to represent white). If the pixel value is less than or equal to the preset threshold value, it is assigned a pixel value representing black (in an 8-bit grayscale image, the pixel value is set to 0 to represent black). The grayscale image is simplified into a binary image containing only black and white pixel values, so that the complex grayscale level difference between the original hand body and the background is converted into a clear black and white contrast relationship, thereby greatly enhancing the contrast between the hand body and the background. In the final binary image, the shape and outline of the hand are clearly presented, allowing people to distinguish it at a glance, providing a more intuitive and effective image basis for subsequent related processing operations such as hand feature extraction and gesture recognition.
[0129] In actual use scenarios, since the lighting conditions of the shooting environment are often difficult to achieve ideal conditions, hand images are prone to shadow gradients caused by uneven lighting, as well as complex textures of the hand skin itself. These unnecessary details can be effectively removed after binarization. After such processing, a pure and well-defined binary image is obtained, which provides an ideal data basis for the limb recognition algorithm.
[0130] Among them, the limb recognition algorithm realizes the real-time interpretation and understanding of human movements, gestures and postures by detecting and tracking the body parts and joints of the human body. The limb recognition algorithm can be a bottom-up algorithm (also known as the Part-Based method). The limb recognition algorithm first detects the key points of the human body in the image or video, and then matches different key points, and connects the key points belonging to one hand. After connecting the key points belonging to one hand, a hand contour is formed, and the gesture features can be determined based on the hand contour. By using the limb recognition algorithm to accurately recognize gestures, the key node information of the gestures can be deeply extracted, such as the extension state of each finger, fully extended, slightly bent, and the bending angle is accurate to a specific degree, and the orientation of the palm, whether the palm is facing up, down, left or right, etc. These features accurately reflect the intention behind the hand movement.
[0131] It should also be noted that the display device provided by any of the above embodiments generates a hand shadow image and a virtual object image based on the recognition of gesture actions. From the perspective of action recognition accuracy, gesture actions are relatively more refined and concentrated than limb actions. The hand joints have high flexibility and can present a rich variety of subtle changes. For example, when simulating the shape of an animal, the slight bending and extension angle adjustment of the fingers can accurately simulate details such as the opening and closing of a bird's beak and the shaking of a rabbit's ears. The hand shadow features formed by these fine actions can be quickly and accurately captured by a high-precision image recognition algorithm, effectively reducing the risk of misjudgment. In contrast, limb actions have a wide range and complex joint linkages, are easily interfered by factors such as body posture and occlusion, are more difficult to identify, and it takes longer to accurately locate the key action features.
[0132] Again, in terms of data processing efficiency, gesture generation focuses on the hand area, and the amount of data is significantly reduced compared to limb generation. In the image acquisition stage, the shooting component only needs to capture the image of the hand range, reducing the data load for transmission and processing, and can speed up the overall processing process. For example, in a real-time interaction scenario, the smart TV can process gesture images faster and with lower latency, ensuring that the user can almost instantly see the response of the virtual character. However, limb generation involves full-body movements, and the amount of high-definition image data is large, which requires strict requirements for transmission bandwidth and processing computing power, and is prone to bottlenecks in the data transmission and processing links, affecting the real-time nature of the interaction.
[0133] Furthermore, considering the richness of creative expression, gestures, with their dexterity and variability, can create a vast number of unique hand shadow forms. One hand can transform into a variety of animal and object contours, and the combination of both hands can unlock even more fantastic and complex shapes, inspiring users' infinite creative imagination. For example, simple finger crossing and flipping can switch from a flying bird to a sailboat to meet the needs of different scenarios. Limb generation is limited by the body structure and movement amplitude, and the creative expression is relatively limited, mostly expanding around common human actions, and it is difficult to switch and combine as quickly and arbitrarily as gestures, bringing users endless novel experiences.
[0134] In addition, in terms of adaptability and convenience, gesture generation does not rely on a spacious space. Users can easily perform it in narrow areas such as on the living room sofa or beside the bedroom bed, enabling interaction anytime and anywhere. Body generation, on the other hand, often requires a larger activity space to ensure the complete capture of full-body movements. Its usage scenarios are limited, and frequent large-scale body movements can easily make people tired, which is not conducive to long-term immersive interaction. Gesture generation is more casual and convenient in daily usage scenarios.
[0135] The scenarios where this display device can be applied include family parent-child interactive entertainment, children's education assistance, social gathering interaction, rehabilitation training assistance, etc.
[0136] In the family parent-child interactive entertainment scenario, parents and children can play shadow puppetry games together using the display device. By making various shadow puppetry poses, children can instantly see vivid virtual animals or fantasy characters generated on the screen, which can not only stimulate children's imagination and creativity but also enhance the interaction and communication between parents and children, making family entertainment time more colorful.
[0137] In the children's education assistance scenario, knowledge learning is combined with interesting shadow puppetry interaction. Teachers or parents can design shadow puppetry courses according to the teaching content with the help of the system of this display device. For example, when learning about animal knowledge, let children create the corresponding animals with shadow puppetry, and at the same time, the screen shows popular science materials such as the detailed information and living habits of the animal, so as to deepen children's understanding and memory of knowledge in an intuitive and interesting way and improve the learning effect.
[0138] In the social gathering interaction scenario, the display device becomes the focus of entertainment. Participants show their best shadow puppetry one after another, and various funny and cool virtual characters are generated on the screen, creating a happy atmosphere for the gathering. Everyone can also compete to see whose virtual character generated by shadow puppetry is the most popular, becoming a new way of social interaction and breaking the monotony of traditional gatherings.
[0139] In the rehabilitation training assistance scenario, for patients with hand rehabilitation, medical staff set a sequence of shadow puppetry actions for rehabilitation training. Patients perform shadow puppetry as required, and the system recognizes the completion degree of the actions and generates virtual characters to give encouraging feedback. For example, when a patient successfully completes the shadow puppetry action of clenching and stretching the fist, a virtual little person applauding appears on the screen, motivating the patient to continue training, improving the enthusiasm and compliance of rehabilitation training, and helping to restore hand function.
[0140] The above content mainly describes the display device. In an exemplary embodiment, a method for displaying virtual objects is also provided, which is applied to the display device mentioned above. Refer to Figure 9 , this method includes:
[0141] Step 902: For the image captured by the shooting component, generate a shadow image of a hand for representing the corresponding projection of the gesture according to the gesture features extracted from the image.
[0142] Among them, the gesture features at least include one of the shape of the gesture contour, the line curvature of the gesture contour, or the area of the closed region formed by the gesture contour, and the shadow image of the hand carries the gesture features.
[0143] Step 904: Obtain a virtual object image based on the shadow image of the hand; among them, the contour features of the virtual object in the virtual object image are associated with the gesture features carried by the shadow image of the hand.
[0144] Step 906: Display the shadow image of the hand on the display screen of the display, and display the virtual object image corresponding to the shadow image of the hand on the display screen.
[0145] Among them, the shooting component can be built into the display device or externally connected to the display device. The shooting component can be a camera. The image captured by the shooting component contains the user's gesture. The image captured by the shooting component can be understood as each frame of image captured by the shooting component. Use the Camera2 API (Application Programming Interface, Chinese name Application Programming Interface) or the camera SDK (Software Development Kit) to capture the real-time video stream.
[0146] Among them, the gesture features include the shape of the gesture contour, and the shape of the gesture contour may present various forms such as a circle, a square, an irregular polygon, etc. Among them, the gesture features include the line curvature of the gesture contour. The line curvature refers to the rotation rate of the tangent direction angle of a point on the curve with respect to the arc length, which is defined by differentiation and is used to indicate the degree of deviation of the curve from a straight line. Mathematically, it is a specific numerical value used to quantify the degree of bending of the curve at a certain point. The greater the line curvature, the greater the degree of bending of the curve. The smaller the line curvature, the closer the curve is to a straight line. The line curvature is used to describe whether the line of the gesture contour is gently curved or sharply turning. Among them, the gesture features include the area of the closed region formed by the gesture contour. The area of the closed region is the numerical value of the closed space enclosed by the lines of the gesture contour.
[0147] Among them, in the field of light propagation, the shadow image of the hand is a shadow image generated on a wall or other plane under the illumination of a light source through the shape and movement of the hand. In this embodiment, after the shooting component captures an image, the image is processed to form a shadow image of the hand. The generated shadow image of the hand is similar to the shadow image formed under the illumination of a light source, and the generated shadow image of the hand is used to represent the corresponding projection of the gesture.
[0148] According to the gesture features, the process of generating a hand shadow image can be to first determine the gesture contour in the image based on the gesture features, and then crop out the hand shadow image from the image based on the gesture contour. Among them, in the process of determining the gesture contour in the image, the image can be first processed into a grayscale image, then the gesture contour is determined based on the grayscale image, and finally the hand shadow image is cropped out from the grayscale image based on the gesture contour. Processing the image into a grayscale image first can reduce the image data to be processed and improve the generation speed of the hand shadow image.
[0149] Among them, the virtual object image includes a virtual object. The type of the virtual object can be an animal, a plant, a building, etc., which is not limited in this embodiment. The type of the virtual object displayed on the display screen can be one of an animal, a plant or a building. When the image captured by the shooting component includes at least two independent gestures, the virtual objects corresponding to different gestures can be of the same type of virtual object or different types of virtual objects. For example, the types of the virtual objects corresponding to two gestures are both animals. For example, the type of the virtual object corresponding to one of the two gestures is an animal, and the type of the virtual object corresponding to the other gesture is a plant.
[0150] Optionally, virtual animals can be generated based on common hand shadow shapes. For example, when the hand presents a contour similar to the flapping of a bird's wings, that is, the fingers are extended and slightly bent, simulating the gesture of a bird's wings flapping, capturing the intention of the hand movement, and then generating a bird character flying freely in the air with vivid feathers. The fingers are bent to simulate the shape of a rabbit's ears, the index finger and the middle finger are erected and bent inward, and the rest of the fingers are clenched into a fist, just like the lively erected ears of a rabbit, corresponding to generating a lively and lovely rabbit image that jumps.
[0151] Optionally, virtual characters with fantasy colors can also be generated, such as little elves with magic special effects triggered by complex gesture combinations. These complex gesture combinations may be that the ten fingers of both hands quickly cross and change positions, and at the same time, a faint light flashes in the palms. The system accurately analyzes these unique hand shadow features with advanced recognition technology, and then flexibly constructs little elves with different shapes and surrounded by dazzling light. They either hold a magic wand or have wings on their backs, full of fantasy charm. The system flexibly constructs the virtual character form according to the recognized hand shadow features, aiming to meet the diverse creative expression needs of users and enhance the interactive entertainment experience based on smart TVs.
[0152] It should be noted that when the images captured by the shooting component contain at least two independent gestures, a hand shadow image and a virtual object image corresponding to the hand shadow image can be generated for each independent gesture, or the independent gestures can be associated as a combined gesture, and a hand shadow image and a virtual object image corresponding to the hand shadow image can be generated for the combined gesture. It should be noted that the at least two independent gestures described refer to that there is no contact point between at least two gestures. Correspondingly, there is also no connection point between the hand shadow images corresponding to each gesture.
[0153] Among them, the hand shadow image carries the gesture feature. Obtaining the virtual object image based on the hand shadow image is mainly to generate the virtual object image according to the gesture feature. The contour feature of the virtual object in the virtual object image is associated with the gesture feature carried by the hand shadow image. For example, if the user's hand presents a contour similar to a bird spreading its wings, then the virtual object in the virtual object image generated according to the gesture feature is a bird. Correspondingly, the contour feature of the virtual object is used to represent that the contour of the virtual object is the contour of a bird spreading its wings, and the contour feature of the virtual object is associated with the gesture feature. The contour feature of the virtual object may include one of the shape of the virtual object contour, the line curvature of the virtual object contour, or the area of the closed region formed by the virtual object contour. The gesture feature constitutes the basic information closely related to the construction of the virtual character form, ensuring that the generated virtual object can correspond to the hand shadow action in terms of appearance.
[0154] In one embodiment, a deep learning algorithm can be used to classify and recognize gesture features, map the recognized gestures to predefined virtual objects, and further generate virtual object images. Correspondingly, before classifying and recognizing gesture features, a database containing various gestures and corresponding virtual objects needs to be constructed.
[0155] Among them, the virtual object image can be a two-dimensional image or a three-dimensional image, which is not limited in this embodiment.
[0156] Among them, the display position of the hand shadow image can be fixed, or it can correspond to the position of the gesture in the image captured by the shooting component and move with the movement of the gesture position, which is not limited in this embodiment.
[0157] Among them, the display positions of the virtual object image and the corresponding hand shadow image can be the same or different, which is not limited in this embodiment.
[0158] Among them, when the gesture changes, it is equivalent to the shooting component capturing a new image, and then a new hand shadow image and a virtual object image are generated for the new image.
[0159] Among them, the schematic diagram of displaying the hand shadow image on the display screen is asFigure 6 As shown, the outline of the hand shadow image is consistent with the outline of the user's gesture. Among them, the schematic diagram of displaying the virtual object image on the display screen is as Figure 7 shown. It can be seen that the outline features of the virtual objects corresponding to different gestures are different.
[0160] In the method provided in this embodiment, when the image captured by the shooting component contains a gesture, the gesture feature is extracted, and a hand shadow image is generated based on the gesture feature and the hand shadow image is displayed on the display screen. Further, a virtual object image is also obtained and displayed based on the hand shadow image. That is to say, through gesture interaction, the user can not only see a hand shadow image formed by light transmission on the display screen of the display, but also see a virtual character image corresponding to the gesture and the hand shadow image. In this way, this embodiment can develop a hand shadow display function and a virtual object display function for the user's gesture, increasing the entertainment and interactivity of the display device.
[0161] In some embodiments, step 904 obtains a virtual object image based on the hand shadow image, including:
[0162] Step 1, input the hand shadow image carrying the gesture feature into the virtual object generation model, and obtain the virtual object image generated by the virtual object generation model; and / or, send the hand shadow image carrying the gesture feature to the cloud server, and obtain the virtual object image generated and sent by the cloud server based on the virtual object generation model.
[0163] Among them, the hand shadow image is processed by a deep learning algorithm. The virtual object generation model can be a Generative Adversarial Networks (GANs) or a Variational Autoencoders (VAE), or other models capable of generating virtual object images, which are not limited in this embodiment. The virtual object generation model can be set in the display device or the cloud server, or can be set in both the display device and the cloud server.
[0164] Correspondingly, the hand shadow image carrying the gesture feature can be input into the virtual object generation model set in the display device, and the virtual object image generated by the virtual object generation model can be obtained.
[0165] Correspondingly, the hand shadow image carrying the gesture feature can also be sent to the cloud server, and the virtual object image generated and sent by the cloud server based on the virtual object generation model can be obtained. Specifically, the cloud server inputs the hand shadow image into the virtual object generation model, obtains the virtual object image generated by the virtual object generation model, and sends the virtual object image to the display device. After receiving the virtual object image sent by the cloud server, the display device can load the virtual object image through a third-party library such as Glide or a tool such as BitmapFactory, and then display the virtual object image on the display screen.
[0166] Inputting the hand shadow image carrying the gesture feature into the virtual object generation model can obtain the virtual object image output by the virtual object generation model.
[0167] The method provided in this embodiment can improve the generation efficiency of the virtual object image by utilizing the computing power of the cloud server through sending the hand shadow image to the cloud server for processing. In addition, compared with directly sending the image captured by the shooting component to the cloud server, in this embodiment, the processed hand shadow image is sent to the cloud server, and the data volume of the hand shadow image is significantly smaller than that of the image captured by the shooting component. Therefore, by combining local processing and cloud server computing, this embodiment can improve the system response speed and reduce the excessive dependence on cloud resources while ensuring the generation of high-quality images.
[0168] In some embodiments, please refer to Figure 10 , step 906 of displaying the virtual object image corresponding to the hand shadow image on the display screen includes:
[0169] Step 1002, determining the area of the hand shadow region based on the number of pixel points included in the hand shadow region in the hand shadow image.
[0170] Step 1004, determining the display ratio of the virtual object in the virtual object image based on the aspect ratio of the circumscribed rectangle region of the hand shadow region and the area of the hand shadow region.
[0171] Step 1006, displaying the virtual object image based on the display ratio of the virtual object.
[0172] Among them, the hand shadow region in the hand shadow image can be understood as the closed region formed by the gesture contour. According to the number of pixel points included in the hand shadow region, the area of the hand shadow region can be determined.
[0173] Among them, please refer to Figure 8 , the circumscribed rectangle region of the hand shadow region can be understood as the circumscribed rectangle region of the gesture contour. The circumscribed rectangle region can be understood as the smallest rectangle that just contains all the points of the hand shadow region or the gesture contour.
[0174] Among them, the aspect ratio of the circumscribed rectangular area refers to the ratio of the length to the width of the circumscribed rectangular area, which is used to reflect the shape characteristics of the hand shadow area. Generally speaking, it can be determined whether the hand shadow area is slender or short and fat through this aspect ratio.
[0175] Among them, according to the aspect ratio and the area of the hand shadow area, the display ratio (i.e., the scaling ratio) of the virtual object image can be determined, and it can be clarified how much the virtual object image should be enlarged or reduced for display.
[0176] The method provided in this embodiment displays the virtual object image based on this display ratio, ensuring that the display effect of the generated virtual object on the display screen can not only perfectly match the hand shadow image (or the amplitude of the hand shadow movement), but also be fully reasonable in terms of visual perception.
[0177] In some embodiments, please refer to Figure 11 , step 906 includes:
[0178] Step 1102, based on the gesture position extracted from the image, determine the display position of the hand shadow image on the display screen.
[0179] Step 1104, based on the display position, display the hand shadow image on the display screen, and based on the display position, display the virtual object image on the display screen.
[0180] Among them, the gesture position can be understood as the position coordinates of the gesture in the image. The display position (i.e., the position coordinates) of the hand shadow image on the display screen can be the same as or close to the position coordinates of the gesture in the image.
[0181] Optionally, the mapping relationship between the gesture position and the display position of the hand shadow image can be set in advance. After determining the gesture position, the display position of the hand shadow image is determined based on this mapping relationship.
[0182] Among them, after determining the display position of the hand shadow image, the hand shadow image is displayed based on this display position, and the virtual object image is displayed based on this display position. That is to say, the display position of the virtual object image is the same as that of the corresponding hand shadow image.
[0183] The method provided in this embodiment determines the display position based on the gesture position extracted from the image, and then displays the hand shadow image and the virtual object image based on the display position. In this way, when the user waves the gesture, the hand shadow image and the virtual object image move with the wave of the gesture, giving the user a feeling of real-time projection and enhancing the entertainment and interactivity of the display device.
[0184] In some embodiments, after step 1104 displays the virtual object image on the display screen based on the display position, it further includes:
[0185] Step 1, when the virtual object in the virtual object image is an animal, display the introduction information of the virtual object.
[0186] Among them, the introduction information can be displayed on the side of the virtual object image, or at the corner of the display screen (such as the upper left corner, lower left corner, upper right corner, lower right corner, etc.), or in other areas. This embodiment does not make a limitation. Preferably, the displayed introduction information does not block the virtual object image. Optionally, the displayed introduction information can also partially block or semi-transparently block the virtual object image.
[0187] Among them, the introduction information can include the name, subject, habits, etc. of the animal.
[0188] The method provided in this embodiment can improve the educational nature by displaying the introduction information, which is helpful for children's educational assistance or family parent-child interactive entertainment.
[0189] In some embodiments, after step 1104 displays the virtual object image on the display screen based on the display position, it is further configured to:
[0190] Step 1, when there are multiple virtual objects displayed on the display screen, obtain the voting result reflecting the popularity of the virtual objects and display it on the display screen.
[0191] Among them, the voting result corresponding to the virtual object can be displayed on the side of the virtual object. The voting results corresponding to multiple virtual objects can also be centrally displayed in an area. For example, it can be displayed at the corner of the display screen (such as the upper left corner, lower left corner, upper right corner, lower right corner, etc.), or in other areas. This embodiment does not make a limitation. Preferably, the voting result does not block the virtual object image. Optionally, the voting result can also partially block or semi-transparently block the virtual object image.
[0192] Among them, the process of obtaining the voting result can be to display a voting interface and display the information of all historical virtual objects and voting buttons on the voting interface. Responding to the selection information of the voting button (such as the number of selections, selection duration, selection frequency, etc.), the voting result of the virtual object corresponding to the voting button is determined according to the selection information. For example, if the voting button of virtual object A is selected 5 times, the voting button of virtual object B is selected 3 times, and the voting button of virtual object C is selected 1 time, then the voting result of virtual object A is 5 votes, the voting result of virtual object B is 3 votes, and the voting result of virtual object C is 1 vote.
[0193] Among them, the voting results can be displayed in the form of a bar chart, a pie chart, a line chart or other graphs.
[0194] In one embodiment, the generation preference of virtual objects can be optimized according to the voting results. For example, if it is known from the voting results that the virtual objects of birds are the most popular, then in the subsequent generation process of virtual objects, virtual objects of birds are preferentially generated.
[0195] The method provided in this embodiment, when displaying multiple virtual objects on the display screen, displays the voting results of different virtual objects to reflect the popularity of different virtual objects, enhancing the interest of virtual object display.
[0196] In some embodiments, please refer to Figure 12 , the process of extracting the gesture features includes:
[0197] Step 1202: Convert the image into a grayscale image and perform binarization processing on the grayscale image to obtain a binarized image.
[0198] Step 1204: Based on the limb recognition algorithm, recognize the hand joint point information in the binarized image, and determine the gesture features according to the hand joint point information.
[0199] Among them, the method of converting the image into a grayscale image can be the average method, the weighted average method, the maximum value method, the minimum value method, the brightness method, the direct reading method or other methods, which are not limited in this embodiment. Among them, the average method is to use the average value of the red, green and blue color components of each pixel as the grayscale value, and perform arithmetic averaging on the pixel values of the red, green and blue channels, that is, Gray = (R + G + B) / 3, where R represents red, G represents green, and B represents blue, and the new grayscale value of each pixel point is determined through such calculation. Among them, the maximum value method is to use the largest color component in the pixel as the grayscale value, that is, directly take the maximum value in the red, green and blue channels as the grayscale value of the pixel point, and the formula is expressed as Gray = max(R, G, B), where R represents red, G represents green, and B represents blue.
[0200] Among them, the minimum value method is to use the smallest color component in the pixel as the grayscale value.
[0201] In one embodiment, converting the image into a grayscale image may involve, for each pixel point in the image, obtaining the grayscale value of the pixel point based on the color value of the pixel point in different color channels and the corresponding weights of the color channels. For example, different weights are assigned to the three color channels of red (R), green (G), and blue (B) according to the sensitivity of the human eye to different colors. The commonly used weight coefficients are 0.299 for the red channel, 0.587 for the green channel, and 0.114 for the blue channel. Then, the grayscale value corresponding to each pixel point is calculated according to the formula Gray = 0.299×R + 0.587×G + 0.114×B, and this grayscale value is assigned to the original pixel point, thus achieving the conversion from color to grayscale.
[0202] Through the above grayscale operation, the image captured by the shooting component can remove redundant color information and be presented in a more concise data form while retaining the basic structure of the image and the key features of the hand, laying a good data foundation for subsequent further image processing processes such as binarization and feature extraction, and reducing problems such as excessive computation and reduced processing efficiency that may occur due to overly complex data in the subsequent processing.
[0203] Among them, binarizing the grayscale image means converting the grayscale image into an image with only black and white colors. Through binarization, noise in the grayscale image can be removed, and the gesture in the grayscale image can be separated from the background, facilitating subsequent processing. The image after binarization enables subsequent limb recognition algorithms to more efficiently and accurately locate hand joint points, clearly identify gesture features such as finger distribution, and then accurately determine the gesture action.
[0204] In one embodiment, during the binarization process of the grayscale image, the grayscale value of the pixel in the image is compared with a preset threshold value, and according to the comparison result, the grayscale image formed by the grayscale values of different pixels is converted into a binary image. Specifically, a specific threshold value is selected as the preset threshold value according to actual needs, and a corresponding binarization algorithm (such as a global threshold method) is used to compare the pixel value of each pixel in the grayscale image with the preset threshold value. If the pixel value is greater than the preset threshold value, the pixel point is assigned a pixel value representing white (in an 8-bit grayscale image, the pixel value is set to 255 to represent white). If the pixel value is less than or equal to the preset threshold value, it is assigned a pixel value representing black (in an 8-bit grayscale image, the pixel value is set to 0 to represent black). The grayscale image is simplified into a binary image containing only black and white pixel values, so that the complex grayscale level difference between the original hand body and the background is converted into a clear black and white contrast relationship, thereby greatly enhancing the contrast between the hand body and the background. In the final binary image, the shape and outline of the hand are clearly presented, allowing people to distinguish it at a glance, providing a more intuitive and effective image basis for subsequent related processing operations such as hand feature extraction and gesture recognition.
[0205] In actual use scenarios, since the lighting conditions of the shooting environment are often difficult to achieve ideal conditions, hand images are prone to shadow gradients caused by uneven lighting, as well as complex textures of the hand skin itself. These unnecessary details can be effectively removed after binarization. After such processing, a pure and well-defined binary image is obtained, which provides an ideal data basis for the limb recognition algorithm.
[0206] Among them, the limb recognition algorithm realizes real-time interpretation and understanding of human movements, gestures and postures by detecting and tracking the body parts and joints of the human body. The limb recognition algorithm can be a bottom-up algorithm (also known as a Part-Based method). The limb recognition algorithm first detects the key points of the human body in the image or video, and then matches different key points to connect the key points belonging to a hand. After connecting the key points belonging to a hand, a hand contour is formed, and the gesture features can be determined based on the hand contour. Using the limb recognition algorithm to accurately identify gestures, the key node information of the gesture can be deeply extracted, such as the extension state of each finger, fully straightened, slightly bent, the bending angle is accurate to a specific degree, and the direction of the palm, whether the palm is facing up, down, left, right, etc. These features accurately reflect the intention behind the hand movement.
[0207] In some embodiments, in combination with the above content, the interaction process between the camera assembly, the controller and the display can be referred to as follows: Figure 13 .exist Figure 13In it, the imaging component can capture images. The controller can generate a hand shadow image used to represent the corresponding projection of the gesture based on the gesture features extracted from the image, and the hand shadow image is displayed on the display screen of the display. The controller sends the hand shadow image to the cloud server. The cloud server generates a virtual object image based on the hand shadow image and sends the virtual object image to the controller. The controller receives the virtual object image, and the virtual object image is then displayed on the display screen of the display.
[0208] It should also be noted that for the method provided in any of the above embodiments, the hand shadow image and the virtual object image are generated based on the recognition of gesture actions. In terms of the accuracy of action recognition, gesture actions are relatively more delicate and concentrated than limb actions. The hand joints have high flexibility and can present a rich variety of subtle changes. For example, when simulating the shape of an animal, the slight bending and stretching angle adjustment of the fingers can accurately simulate details such as the opening and closing of a bird's beak and the shaking of a rabbit's ears. The hand shadow features formed by these fine actions can be quickly and accurately captured by a high-precision image recognition algorithm, effectively reducing the risk of misjudgment. In contrast, limb actions have a wide range and complex joint linkages, are easily interfered by factors such as body posture and occlusion, have a higher recognition difficulty, and it takes longer to accurately locate the key action features.
[0209] Furthermore, in terms of data processing efficiency, gesture generation focuses on the hand area, and the amount of data is significantly reduced compared to limb generation. In the image acquisition stage, the shooting component only needs to capture the image within the range of the hand, reducing the data load for transmission and processing, and can speed up the overall processing process. For example, in a real-time interaction scenario, a smart TV can process gesture images faster with lower latency, ensuring that users can almost instantly see the response of the virtual character. However, limb generation involves full-body actions, with a large amount of high-definition image data, which requires strict requirements for transmission bandwidth and processing computing power, and is prone to bottlenecks in the data transmission and processing links, affecting the real-time nature of the interaction.
[0210] Moreover, considering the richness of creative expression, gestures, with their dexterity and variability, can create a vast number of unique hand shadow forms. One hand can transform into various animal and object outlines, and the combination of both hands can unlock even more fantastic and complex shapes, inspiring users' infinite creative imagination. For example, simple finger crossing and flipping can switch from a flying bird to a sailboat, meeting the needs of different scenarios. Limb generation is limited by the body structure and action amplitude, and the creative expression is relatively limited, mostly expanding around common human actions, and it is difficult to switch and combine as quickly and randomly as gestures, bringing users a continuous stream of novel experiences.
[0211] In addition, in terms of adaptability and convenience, gesture generation does not rely on a spacious space. Users can easily perform it in narrow areas such as on the living room sofa or beside the bedroom bed, enabling interaction at any time and anywhere. Body generation, on the other hand, often requires a large activity space to ensure the complete capture of full-body movements. Its usage scenarios are limited, and frequent large-scale body movements can easily make people tired, which is not conducive to long-term immersive interaction. Gesture generation is more casual and convenient in daily usage scenarios.
[0212] The scenarios to which the method applied to the display device can be applied include family parent-child interactive entertainment, children's education assistance, social gathering interaction, rehabilitation training assistance, etc.
[0213] In the scenario of family parent-child interactive entertainment, parents and children can play shadow puppetry games together using the display device. By making various shadow puppetry poses, children can immediately see vivid virtual animals or fantasy characters generated on the screen, which can not only stimulate children's imagination and creativity but also enhance the interactive communication between parents and children, making family entertainment time more colorful.
[0214] In the scenario of children's education assistance, knowledge learning is combined with interesting shadow puppetry interaction. Teachers or parents can design shadow puppetry courses according to the teaching content with the help of the system of the display device. For example, when learning about animal knowledge, let children create the corresponding animals with shadow puppetry, and at the same time, the display on the screen shows detailed information about the animal, its living habits and other popular science materials, so as to deepen children's understanding and memory of knowledge in an intuitive and interesting way and improve the learning effect.
[0215] In the scenario of social gathering interaction, the display device becomes the focus of entertainment. Participants show their best shadow puppetry one after another, and various funny and cool virtual characters are generated on the screen, creating a happy atmosphere for the gathering. Everyone can also compete to see whose virtual character generated by shadow puppetry is the most popular, becoming a new way of social interaction and breaking the monotony of traditional gatherings.
[0216] In the scenario of rehabilitation training assistance, for hand rehabilitation patients, medical staff set a sequence of shadow puppetry movements for rehabilitation training. Patients perform shadow puppetry as required, and the system recognizes the completion degree of the movements and generates virtual characters to give encouraging feedback. For example, when a patient successfully completes the shadow puppetry movement of clenching and stretching the fist, a virtual little person applauding appears on the screen, inspiring the patient to continue training and improving the enthusiasm and compliance of rehabilitation training to assist in the recovery of hand function.
[0217] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0218] Based on the same inventive concept, an embodiment of the present application further provides a virtual object display device for implementing the virtual object display method described above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the virtual object display device provided below can refer to the limitations on the virtual object display method in the above text, and will not be repeated here.
[0219] In an exemplary embodiment, as Figure 14 shown, a virtual object display device is provided, including: an image generation module 1402 and a display module 1404, where:
[0220] The image generation module 1402 is configured to generate a hand shadow image for characterizing the corresponding projection of the gesture according to the gesture features extracted from the image for the image captured by the shooting component; wherein, the gesture features at least include one of the shape of the gesture contour, the line curvature of the gesture contour, or the area of the closed region formed by the gesture contour, and the hand shadow image carries the gesture features.
[0221] The image generation module 1402 is further configured to obtain a virtual object image based on the hand shadow image; wherein, the contour features of the virtual object in the virtual object image are associated with the gesture features carried by the hand shadow image.
[0222] The display module 1404 is configured to display the hand shadow image on the display screen of the display, and display the virtual object image corresponding to the hand shadow image on the display screen.
[0223] In an exemplary embodiment, the image generation module 1402 is specifically configured to input the hand shadow image carrying the gesture features into a virtual object generation model and obtain the virtual object image generated by the virtual object generation model; and / or send the hand shadow image carrying the gesture features to a cloud server and obtain the virtual object image generated and sent by the cloud server based on the virtual object generation model.
[0224] In an exemplary embodiment, the display module 1404 is specifically configured to determine the area of the hand shadow region based on the number of pixel points included in the hand shadow region in the hand shadow image; determine the display ratio of the virtual object in the virtual object image based on the aspect ratio of the circumscribed rectangular region of the hand shadow region and the area of the hand shadow region; and display the virtual object image based on the display ratio of the virtual object.
[0225] In an exemplary embodiment, the display module 1404 is specifically configured to determine the display position of the hand shadow image on the display screen based on the gesture position extracted from the image; display the hand shadow image on the display screen based on the display position, and display the virtual object image on the display screen based on the display position.
[0226] In an exemplary embodiment, the display module 1404 is further configured to display the introduction information of the virtual object when the virtual object in the virtual object image is an animal.
[0227] In an exemplary embodiment, the display module 1404 is further configured to, when there are multiple virtual objects displayed on the display screen, obtain the voting result reflecting the popularity of the virtual objects and display it on the display screen.
[0228] In an exemplary embodiment, the image generation module 1402 is specifically configured to convert the image into a grayscale image, perform binarization processing on the grayscale image to obtain a binarized image; identify the hand joint point information in the binarized image based on a limb recognition algorithm, and determine the gesture feature according to the hand joint point information.
[0229] In an exemplary embodiment, the image generation module 1402 is specifically configured to, for the pixel points in the image, obtain the grayscale value of the pixel point according to the color value of the pixel point in different color channels and the corresponding weight of the color channel; compare the grayscale value of the pixel point in the image with a preset threshold, and convert the grayscale image formed by the grayscale values of different pixel points into a binarized image according to the comparison result.
[0230] Each module in the above display device of the virtual object can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or be stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.
[0231] In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the possible implementation manners provided by the above various methods are realized.
[0232] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the possible implementation manners provided by the above various methods are realized.
[0233] In an embodiment, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the possible implementation manners provided by the above various methods are realized.
[0234] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0235] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0236] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.
[0237] The above embodiments only represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A display device, characterized in that: include: a display configured to display an image and / or a user interface; A shooting component, used for shooting images; The controller is configured as: For the image captured by the shooting component, a hand shadow image for representing the corresponding projection of the gesture is generated according to the gesture features extracted from the image; wherein the gesture features include at least one of the shape of the gesture contour, the line curvature of the gesture contour, or the area of the closed area formed by the gesture contour, and the hand shadow image carries the gesture features; Based on the hand shadow image, a virtual object image is acquired; wherein the contour features of the virtual object in the virtual object image are associated with the gesture features carried by the hand shadow image; The hand shadow image is displayed on a display screen of a display, and a virtual object image corresponding to the hand shadow image is displayed on the display screen.
2. The display device according to claim 1, characterized in that The controller executes acquiring a virtual object image based on the hand shadow image, and is configured to: Inputting the hand shadow image carrying the gesture feature into a virtual object generation model, and obtaining the virtual object image generated by the virtual object generation model; and / or, The hand shadow image carrying the gesture features is sent to the cloud server, and the virtual object image generated and sent by the cloud server based on the virtual object generation model is obtained.
3. The display device according to claim 1 or 2, characterized in that: The controller is configured to: display a virtual object image corresponding to the hand shadow image on the display screen; Determine the area of the hand shadow region based on the number of pixels included in the hand shadow region in the hand shadow image; determining a display ratio of the virtual object in the virtual object image based on the aspect ratio of a circumscribed rectangular area of the hand shadow area and the area of the hand shadow area; The virtual object image is displayed based on the display ratio of the virtual object.
4. The display device according to claim 1 or 2, characterized in that: The controller executes displaying the hand shadow image on a display screen of the display, and displays a virtual object image corresponding to the hand shadow image on the display screen, and is configured as follows: Determining a display position of the hand shadow image on the display screen based on the gesture position extracted from the image; Based on the display position, displaying the hand shadow image on the display screen, and, The virtual object image is displayed on the display screen based on the display position.
5. The display device according to claim 4, characterized in that After the controller executes displaying the virtual object image on the display screen based on the display position, the controller is further configured to: In a case where the virtual object in the virtual object image is an animal, introduction information of the virtual object is displayed.
6. The display device according to claim 4, characterized in that After the controller executes displaying the virtual object image on the display screen based on the display position, the controller is further configured to: When a plurality of virtual objects are displayed on the display screen, a voting result reflecting the popularity of the virtual objects is obtained and displayed on the display screen.
7. The display device according to claim 1 or 2, characterized in that: The controller performs the gesture feature extraction process and is configured to: Converting the image into a grayscale image, and performing binarization processing on the grayscale image to obtain a binarized image; Based on a limb recognition algorithm, the hand joint point information in the binary image is identified, and the gesture feature is determined according to the hand joint point information.
8. The display device according to claim 7, characterized in that The controller converts the image into a grayscale image and performs binarization processing on the grayscale image to obtain a binarized image, and is configured as follows: For a pixel in the image, obtaining a grayscale value of the pixel according to the color value of the pixel in different color channels and the corresponding weight of the color channel; The grayscale values of the pixels in the image are compared with a preset threshold, and according to the comparison result, the grayscale image formed by the grayscale values of different pixels is converted into a binary image.
9. A method for displaying a virtual object, characterized in that: Applied to the display device according to any one of claims 1 to 8, the method comprises: For the image captured by the shooting component, a hand shadow image for representing the corresponding projection of the gesture is generated according to the gesture features extracted from the image; wherein the gesture features include at least one of the shape of the gesture contour, the line curvature of the gesture contour, or the area of the closed area formed by the gesture contour, and the hand shadow image carries the gesture features; Based on the hand shadow image, a virtual object image is acquired; wherein the contour features of the virtual object in the virtual object image are associated with the gesture features carried by the hand shadow image; The hand shadow image is displayed on a display screen of a display, and a virtual object image corresponding to the hand shadow image is displayed on the display screen.
10. The method according to claim 9, characterized in that The step of displaying a virtual object image corresponding to the hand shadow image on the display screen comprises: Determine the area of the hand shadow region based on the number of pixels included in the hand shadow region in the hand shadow image; determining a display ratio of the virtual object in the virtual object image based on the aspect ratio of a circumscribed rectangular area of the hand shadow area and the area of the hand shadow area; The virtual object image is displayed based on the display ratio of the virtual object.