Gesture interaction method and system for controlling smart television, smart television and readable storage medium
By collecting hand image data on smart TVs and using gesture recognition models to identify custom gestures, combined with a preset gesture interaction database, the problem of low flexibility in gesture recognition interaction mode in the prior art is solved, and personalized and flexible gesture control is achieved.
Patent Information
- Application Number
- CN202510243874.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-01
AI Technical Summary
The gesture recognition interaction method of existing smart TVs is low in flexibility and cannot achieve personalized gesture recognition interaction.
By collecting image data of the hand, input it into the target gesture recognition model, determine the target gesture in the preset gesture interaction database that matches the target gesture recognition result, and obtain the associated interaction method to control the smart TV.
Personalized gesture recognition interaction is realized, improving the flexibility of gesture recognition interaction, and allowing users to modify or expand the gesture interaction database according to actual needs.
Smart Images

Figure CN120233873A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of human-computer interaction technology, and in particular to a gesture interaction method and system for controlling a smart television, a smart television, and a readable storage medium. Background Art
[0002] Smart TV refers to a TV that has a fully open platform and is equipped with an operating system. Users can install and uninstall software, games and other programs provided by third-party service providers. Such programs can continuously expand the functions of the TV and the TV can surf the Internet through network cables and wireless networks.
[0003] Nowadays, smart TVs are favored by more and more users. On smart TVs, users can not only enjoy ordinary TV content, but also install and uninstall various application software by themselves, and can continuously expand and upgrade the TV functions; smart TVs bring users a rich and personalized experience different from traditional TVs. With the continuous increase in the functions of smart TVs, remote control operation has become more and more cumbersome and difficult, which has seriously affected the control experience of smart TVs. In this context, gestures, as an intuitive and natural input method, allow people to operate smart TVs in a more natural way. Therefore, most smart TVs now have added gesture recognition interaction functions.
[0004] However, most traditional gesture recognition interaction methods are based on fixed gesture sets preset by the manufacturer before the smart TV leaves the factory. These gesture sets are fixed in the software system of the smart TV when it leaves the factory, and users cannot modify or expand them during use, resulting in low flexibility of gesture recognition interaction methods and inability to achieve personalized gesture recognition interaction.
[0005] Therefore, how to improve the flexibility of gesture recognition interaction and realize personalized gesture recognition interaction is a technical problem to be solved urgently in the field of this technology. Summary of the invention
[0006] The main purpose of the present application is to provide a gesture interaction method, system, smart TV and readable storage medium for controlling a smart TV, aiming to solve the technical problem of how to improve the flexibility of gesture recognition interaction mode and realize personalized gesture recognition interaction.
[0007] To achieve the above object, the present application provides a gesture interaction method for controlling a smart TV, and the gesture interaction method for controlling a smart TV comprises the following steps:
[0008] Collecting image data of the hand, and inputting the image data into a target gesture recognition model to obtain a target gesture recognition result;
[0009] Determine a target gesture in a preset gesture interaction database that matches the target gesture recognition result, where the gesture interaction database includes at least one custom gesture and an interaction method associated with each custom gesture;
[0010] Obtain the target interaction method associated with the target gesture in the gesture interaction database, and control the smart TV with the target interaction method.
[0011] In one embodiment, the step of determining a target gesture in a preset gesture interaction database that matches the target gesture recognition result includes:
[0012] Perform vectorization processing on the target gesture recognition result to obtain a gesture result vector, and calculate the similarity between the gesture result vector and each gesture feature vector in the preset gesture interaction database;
[0013] Determine the highest similarity among all the similarities. If the highest similarity is greater than or equal to a preset similarity threshold, determine the gesture feature vector corresponding to the highest similarity as the target gesture feature vector;
[0014] Determine the gesture associated with the target gesture feature vector in the gesture interaction database as the target gesture.
[0015] In one embodiment, the step of inputting the image data into a target gesture recognition model to obtain a target gesture recognition result includes:
[0016] Obtain the spectral data of the hand, where the spectral data includes infrared information and / or visible light information;
[0017] Fuse the image data and the spectral data to obtain input data, and input the input data into the target gesture recognition model to obtain a target gesture recognition result.
[0018] In one embodiment, the step of fusing the image data and the spectral data to obtain input data includes:
[0019] If the data formats of the spectral data and the image data are inconsistent, convert the data format of the spectral data to obtain the spectral data after format conversion, where the data formats of the image data and the spectral data after format conversion are consistent, and the data format is the color format of the data;
[0020] If the resolutions of the spectral data after format conversion and the image data are inconsistent, adjust the resolution of the spectral data after format conversion to obtain target spectral data, where the resolution of the target spectral data is consistent with that of the image data;
[0021] Merge the image data and the target spectral data in the channel dimension to obtain input data.
[0022] In one embodiment, before the step of collecting the image data of the hand, the method further includes:
[0023] In response to a gesture input request, obtain the input gesture data according to the gesture input request, where the gesture data at least includes the image data of the hand;
[0024] Detect whether the gesture data is qualified;
[0025] If it is qualified, train a preset gesture recognition model according to the gesture data to obtain the target gesture recognition model, and store the custom gesture corresponding to the gesture data in the preset gesture interaction database.
[0026] In one embodiment, the step of detecting whether the gesture data is qualified includes:
[0027] Input the gesture data into a pre-trained gesture data evaluation model to obtain an evaluation result, where the gesture data evaluation model is a classification model based on machine learning;
[0028] If the evaluation result indicates that the data is qualified, determine that the gesture data is qualified;
[0029] If the evaluation result indicates that the data is unqualified, determine that the gesture data is unqualified.
[0030] In one embodiment, the step of storing the custom gesture corresponding to the gesture data in the preset gesture interaction database includes:
[0031] Generate a gesture recognition result corresponding to the gesture data according to the target gesture recognition model, and obtain a feedback result for the gesture recognition result;
[0032] If the feedback result indicates that the gesture recognition result is correct, obtain the interaction method configured by the user;
[0033] Obtain the custom gesture corresponding to the gesture data, and store the configured interaction method and the input custom gesture in an associated manner in the preset gesture interaction database.
[0034] In addition, to achieve the above object, the present application further provides a gesture interaction system, where the gesture interaction system includes a controller and a smart TV, the controller is connected to the smart TV, and the controller is used to execute the steps of the gesture interaction method for controlling the smart TV as described above.
[0035] In addition, to achieve the above-mentioned purpose, the present application also provides a smart TV, which includes: a memory, a processor, and a gesture interaction program for controlling the smart TV stored in the memory and runnable on the processor, and the gesture interaction program for controlling the smart TV, when executed by the processor, implements the steps of the gesture interaction method for controlling the smart TV as described above.
[0036] In addition, to achieve the above-mentioned purpose, the present application also provides a readable storage medium, which is a computer-readable storage medium, and a program for implementing the gesture interaction method is stored on the computer-readable storage medium. The program for implementing the gesture interaction method is executed by a processor to implement the steps of the gesture interaction method for controlling a smart TV as described above.
[0037] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the gesture interaction method for controlling a smart TV as described above.
[0038] One or more technical solutions proposed in this application have at least the following technical effects:
[0039] Collect image data of the hand, input the image data into the target gesture recognition model to obtain the target gesture recognition result; determine the target gesture matching the target gesture recognition result in the preset gesture interaction database, wherein the gesture interaction database includes at least one custom gesture and an interaction mode associated with each custom gesture; obtain the target interaction mode associated with the target gesture in the gesture interaction database, and control the smart TV with the target interaction mode. In this way, the embodiment of the present application recognizes the user's custom gesture by using the gesture recognition model, and at the same time, the interaction mode of each custom gesture is pre-associated in the gesture interaction database, so that when interacting with the smart TV through gesture recognition, the user's custom gesture can be recognized, and the smart TV can be controlled in the interaction mode associated with the custom gesture to achieve interaction with the smart TV. During use, the user can modify or expand the gesture interaction database based on actual needs, and use the custom gesture in the gesture interaction database to interact with the smart TV, rather than being limited to the gesture set that is solidified in the smart TV at the factory for gesture recognition interaction, thereby improving the flexibility of gesture recognition interaction and achieving personalized gesture recognition interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0041] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 It is a schematic flowchart of the first embodiment of the gesture interaction method for controlling a smart TV in the present application;
[0043] Figure 2 It is a schematic diagram of the gesture interaction process involved in an embodiment of the gesture interaction method for controlling a smart TV in the present application;
[0044] Figure 3 It is a schematic flowchart of the gesture input and model training process involved in an embodiment of the gesture interaction method for controlling a smart TV in the present application;
[0045] Figure 4 It is a schematic diagram of the system structure of the gesture interaction system in the present application;
[0046] Figure 5 It is a schematic diagram of the device structure of the hardware operating environment involved in the gesture interaction device for controlling a smart TV in the embodiments of the present application.
[0047] The realization of the purpose, functional characteristics, and advantages of the present application will be further described with reference to the embodiments and the drawings. Detailed Embodiments
[0048] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0049] In the existing TV interaction technologies, users usually rely on remote controls or voice commands to operate the TV. However, these traditional interaction methods have certain limitations in terms of operation complexity and personalized experience. In particular, they fail to make full use of the intuitive and expressive interaction method of user gestures. In addition, the existing gesture recognition technologies are often limited to a preset set of gestures, lacking flexibility and customization capabilities.
[0050] The existing TV interaction solutions have at least the following disadvantages:
[0051] Remote control operation: Currently, most TV interactions on the market still rely on traditional remote control operation. Although this technology is mature, the user's operating experience is limited by the limited buttons and complex menu navigation of the remote control, resulting in less intuitive and efficient operation.
[0052] Voice command recognition: With the popularity of smart TVs, voice command recognition has become a popular way of interaction. Although this method simplifies the operation process to a certain extent, it may be affected by environmental noise and has limited processing capabilities for non-standard pronunciations or specific accents.
[0053] Gesture recognition: Existing gesture recognition technology operates the TV by analyzing hand movements through video streams, but it has some defects. For example, it is limited to a fixed set of gestures preset by the manufacturer before the smart TV leaves the factory, and users cannot modify or expand it during use. This results in low flexibility in gesture recognition interaction and inability to achieve personalized gesture recognition interaction.
[0054] Based on this, the main solution of the present application is: collect image data of the hand, input the image data into the target gesture recognition model to obtain the target gesture recognition result; determine the target gesture in a preset gesture interaction database that matches the target gesture recognition result, wherein the gesture interaction database includes at least one custom gesture and an interaction method associated with each custom gesture; obtain the target interaction method associated with the target gesture in the gesture interaction database, and control the smart TV with the target interaction method.
[0055] The present application recognizes user-defined gestures by using a gesture recognition model, and at the same time pre-associates the interaction mode of each customized gesture in a gesture interaction database, so that when interacting with a smart TV through gesture recognition, the user's customized gestures can be recognized, and the smart TV can be controlled in the interaction mode associated with the customized gestures to achieve interaction with the smart TV. During use, the user can modify or expand the gesture interaction database based on actual needs, and use the customized gestures in the gesture interaction database to interact with the smart TV, rather than being limited to gesture recognition interaction based on the gesture set solidified in the smart TV at factory time, thereby improving the flexibility of gesture recognition interaction and achieving personalized gesture recognition interaction.
[0056] It should be noted that the executor of each embodiment of the gesture interaction method for controlling a smart TV in the present application may be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a smart TV that can realize the above functions. The embodiments of the gesture interaction method for controlling a smart TV in the present application do not impose any specific restrictions on this.
[0057] Based on this, the present application proposes a gesture interaction method for controlling a smart TV in the first embodiment. Please refer to Figure 1 , and the gesture interaction method for controlling the smart TV includes steps S10 to S30:
[0058] Step S10, collect image data of the hand, and input the image data into a target gesture recognition model to obtain a target gesture recognition result;
[0059] The image data of the hand can be collected by an image acquisition sensor. Specifically, the image data can be an image sequence composed of at least one image.
[0060] The image acquisition sensor is used to collect images, and can specifically be a camera, a depth sensor, etc. This embodiment does not make specific limitations thereto. Considering that the depth sensor can provide three-dimensional information about the scene, which includes the shape, position, and distance of objects, and helps to more accurately understand the posture and movement of the hand; compared with a camera that relies on visible light, the depth sensor is usually not affected by environmental lighting conditions; the depth information can help the system distinguish the foreground (such as the hand) and the background, thereby reducing background interference and improving the accuracy of gesture recognition; in a complex environment, gestures may be blocked by other objects, and the depth sensor can better handle the situation of partial occlusion because it can identify the occluder and the occluded part of the gesture; since the depth image directly provides distance information, this can simplify the image processing algorithm and reduce the dependence on complex image processing techniques, and the depth sensor can directly provide key gesture features, which can reduce subsequent processing steps and thus improve the speed of gesture recognition. Based on this, in a preferred embodiment, the image data is collected at least by the depth sensor.
[0061] After collecting the image data of the hand, input the collected image data into the target gesture recognition model to obtain a target gesture recognition result. The target gesture recognition model can specifically be any AI (Artificial Intelligence) model for gesture recognition that has been pre-trained, such as a convolutional neural network model, a recurrent neural network model, an autoencoder model, etc. This embodiment does not make specific limitations thereto.
[0062] Step S20, determine a target gesture in the preset gesture interaction database that matches the target gesture recognition result, where the gesture interaction database includes at least one custom gesture and an interaction method associated with each custom gesture;
[0063] The preset gesture interaction database can specifically be a pre-set database, which is used to store gestures and interaction methods associated with each gesture, wherein the gestures stored in the gesture interaction database include at least one custom gesture, and the custom gesture refers to a user-defined gesture, such as "waving left hand", "waving right hand", "like after making a fist", etc. The interaction method refers to a method of controlling a smart TV, such as "previous channel", "next channel", "play / pause", "increase volume", "lower volume", etc.
[0064] Furthermore, the gestures stored in the gesture interaction database may also include system gestures. System gestures refer to gestures that are solidified in the smart TV before it leaves the factory, so that both system gestures and custom gestures can realize gesture recognition interaction by matching a unified database, thereby unifying the gesture interaction process. For existing system gestures, if the user does not need to change them, they do not need to re-enter them, which reduces the preparation cost of the gesture interaction database and improves the user experience.
[0065] It should be noted that the gestures stored in the gesture interaction database may be text data or data containing both text and images. For example, a custom gesture stored in the gesture interaction database may be the text "wave left hand" or an image containing both the text "wave left hand" and the image of waving left hand. Furthermore, each gesture stored in the gesture interaction database may be a gesture action (such as "wave left hand") or a combination of multiple gesture actions (such as "like after making a fist").
[0066] Furthermore, the target gesture recognition model is a model trained based on the gestures stored in the gesture interaction database. It should be noted that, specifically, the target gesture recognition model can be a model trained based on the image data of each gesture in the gesture interaction database, so that the target gesture recognition model can recognize the gestures in the gesture interaction database based on the image data. The gesture recognition model can be first trained using a public gesture data set. During use, when a new custom gesture is added, the gesture recognition model is incrementally trained with the newly added custom gesture to obtain a trained target gesture recognition model.
[0067] Furthermore, in order to improve the pre-training effect of the model, the gesture recognition model can be trained using a public gesture data set on a device with higher computing power (such as a server) (hereinafter referred to as initial training), and the trained gesture recognition model can be deployed on the smart TV, so that the subsequent smart TV can directly call the gesture recognition model locally for gesture recognition, and incrementally train the gesture recognition model on the smart TV.
[0068] Furthermore, both the initial training and incremental training of the gesture recognition model can use a preset loss function to iteratively optimize the model, and end the model training when the preset training end condition is met. Among them, the loss function can be a pre-set loss function, such as the cross-entropy loss function, and the training end condition can be a pre-set condition, such as reaching a predetermined number of iterations, the value of the loss function being lower than a predetermined threshold, the computing resources being exhausted, reaching the time limit, the accuracy rate reaching a predetermined threshold, etc. This embodiment does not make specific limitations on this.
[0069] After acquiring the image data of the hand, input the image data of the hand into the target gesture recognition model to obtain the target gesture recognition result, and match the target gesture that is consistent with the target gesture recognition result in the gesture interaction database. That is, the target gesture refers to the gesture in the gesture interaction database that matches the target gesture recognition result.
[0070] It should be noted that if the target gesture is not matched in the gesture interaction database, the current gesture interaction process can be ended, or a prompt message can be output to remind the user that the correct gesture has not been recognized. Relevant personnel can also pre-set other processing methods based on actual needs. This embodiment does not make specific limitations on this.
[0071] Step S30: Obtain the target interaction method associated with the target gesture in the gesture interaction database, and control the smart TV in the target interaction method.
[0072] After matching the target gesture in the gesture interaction database, obtain the interaction method associated with the target gesture, that is, the target interaction method, and control the smart TV in the target interaction method. The target interaction method can specifically be the interaction method stored in association with the target gesture in the gesture interaction database.
[0073] A corresponding control instruction can be generated according to the target interaction method, and the smart TV is controlled by the control instruction to perform corresponding actions to complete the current gesture interaction. For example, assume that the matched target gesture is "wave the left hand", and the target interaction method associated with "wave the left hand" is "previous channel", then a control instruction indicating to switch to the previous channel can be generated to the smart TV. After receiving this control instruction, the smart TV executes it to switch to the previous channel and complete the current gesture interaction.
[0074] This embodiment collects image data of the hand, inputs the image data into the target gesture recognition model to obtain the target gesture recognition result; determines the target gesture matching the target gesture recognition result in the preset gesture interaction database, wherein the gesture interaction database includes at least one custom gesture and an interaction mode associated with each custom gesture; obtains the target interaction mode associated with the target gesture in the gesture interaction database, and controls the smart TV in the target interaction mode. In this way, this embodiment recognizes the user's custom gesture by using the gesture recognition model, and at the same time, the interaction mode of each custom gesture is pre-associated in the gesture interaction database, so that when interacting with the smart TV through gesture recognition, the user's custom gesture can be recognized, and the smart TV can be controlled in the interaction mode associated with the custom gesture to achieve interaction with the smart TV. During use, the user can modify or expand the gesture interaction database based on actual needs, and use the custom gesture in the gesture interaction database to interact with the smart TV, rather than being limited to the gesture set that is solidified in the smart TV at the factory for gesture recognition interaction, thereby improving the flexibility of gesture recognition interaction and achieving personalized gesture recognition interaction.
[0075] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above description and will not be described in detail later. On this basis, the step of determining the target gesture in the preset gesture interaction database that matches the target gesture recognition result includes:
[0076] Step A10, performing vectorization processing on the target gesture recognition result to obtain a gesture result vector, and calculating the similarity between the gesture result vector and each gesture representation vector in a preset gesture interaction database;
[0077] The gesture result vector refers to the quantized target gesture recognition result. Vectorization refers to the process of converting data into a numerical vector. The specific processing method of vectorization processing can be preset. For example, the vectorization processing of data can be completed by embedding (such as word embedding). Embedding processing generally refers to converting data from its original form into a low-dimensional, continuous vector representation that can capture the intrinsic characteristics and structure of the data.
[0078] Optionally, the target gesture recognition result may be subjected to word embedding processing to vectorize the target gesture recognition result, such as inputting the target gesture recognition result into a Word2Vec (Word to Vector) model to output a gesture result vector.
[0079] After obtaining the gesture result vector, calculate the similarity between the gesture result vector and each gesture representation vector in the gesture interaction database. Each gesture representation vector in the gesture interaction database is used to represent a gesture. Among them, the similarity between two vectors is a measure for measuring the similarity degree between two vectors. Specifically, the similarity between two vectors can be the cosine value of the included angle between the two vectors, which can be expressed by the formula: cosθ = (A·B) / (||A||*||B||), where cosθ represents the cosine value between vector A and vector B, that is, the similarity between vector A and vector B, A·B represents the dot product of vector A and vector B, and ||A|| and ||B|| respectively represent the vector lengths (norms) of vector A and vector B.
[0080] It should be noted that when storing a gesture into the gesture interaction database, the stored gesture can be vectorized to obtain the gesture representation vector of the gesture, and the gesture representation vector is stored in association with the gesture, so as to facilitate subsequent searching for the target gesture by means of vector matching.
[0081] Step A20, determine the highest similarity among all the similarities. If the highest similarity is greater than or equal to the preset similarity threshold, determine the gesture representation vector corresponding to the highest similarity as the target gesture representation vector;
[0082] It should be noted that if the highest similarity among all the similarities is less than the preset similarity threshold, it is determined that no target gesture is matched in the gesture interaction database.
[0083] Step A30, determine the gesture associated with the target gesture representation vector in the gesture interaction database as the target gesture.
[0084] After finding the target gesture vector, determine the gesture associated with the target gesture representation vector in the gesture interaction database as the target gesture, thereby completing the matching search for the target gesture.
[0085] In a possible implementation manner, the step of inputting the image data into the target gesture recognition model to obtain the target gesture recognition result includes:
[0086] Step B10, obtain the spectral data of the hand, where the spectral data includes infrared information and / or visible light information;
[0087] Visible light information refers to the image information of the hand in the visible light band (400 - 700 nanometers), which can usually be obtained through a camera. It reflects the characteristics of the hand such as color, texture, and shape. Among them, the hand color refers to the skin color of the hand and the colors of possible ornaments (such as rings, watches). The hand texture refers to the texture details on the skin surface, such as fingerprints, veins, etc. The hand shape refers to the appearance contour, size, finger length and width, and the shape of the palm of the hand and other characteristics.
[0088] Thermal energy information, also known as thermal radiation information, refers to the infrared radiation energy emitted by an object due to its own temperature. In this embodiment, the thermal energy information of the hand can specifically be an infrared imaging diagram of the hand. An infrared camera or a thermal imager can be used to collect the thermal energy information of the hand. The infrared camera or the thermal imager can detect the infrared radiation emitted by an object and convert it into a visible temperature image.
[0089] Considering that the visible light information and the thermal energy information are usually complementary in gesture recognition, the thermal energy information can provide information about the thermal state of the hand, and these information are particularly useful in an environment with insufficient light because they do not depend on the ambient light. While the visible light information provides richer visual details, which helps in accurate gesture recognition under good lighting conditions. Based on this, environmental parameters that can characterize whether the light in the environment is good, such as illuminance, luminous flux, luminance, etc., can be collected first. If the collected environmental parameters indicate good light, only visible light information can be collected. If the collected environmental parameters indicate poor light, only thermal energy information can be collected. In this way, information that is more helpful for improving the accuracy of gesture recognition can be collected as needed based on the environmental light conditions, reducing the amount of information that needs to be collected while improving the accuracy of the target gesture recognition result, and thus reducing the unnecessary power consumption of the collection device.
[0090] It should be noted that the acquisition of spectral data can specifically be simultaneous with the acquisition of image data, or can be ahead of or behind the acquisition of image data. This embodiment does not make specific restrictions on this.
[0091] Step B20, fuse the image data and the spectral data to obtain input data, and input the input data into the target gesture recognition model to obtain the target gesture recognition result.
[0092] It should be noted that the data format of this input data is the same as that of the image data, such as data in the RGB three-channel image format.
[0093] In this embodiment, gesture recognition is performed by combining the image data and the spectral data of the hand. In this way, the spectral data can be used to make up for the problem of low-quality image data caused by factors such as poor environmental light or too fast gestures, thereby improving the accuracy of the target gesture recognition result.
[0094] In a possible implementation manner, the step of fusing the image data and the spectral data to obtain input data includes:
[0095] Step C10, if the data formats of the spectral data and the image data are inconsistent, convert the data format of the spectral data to obtain the spectral data after format conversion, where the data formats of the image data and the spectral data after format conversion are consistent, and the data format is the color format of the data;
[0096] The color format refers to the color representation mode of the data, such as RGB format, grayscale image format, etc.
[0097] When the data formats of the spectral data and the image data are inconsistent, convert the data format of the spectral data to unify the data formats of the image data and the spectral data. Specifically, the data format of the spectral data can be converted by means of channel expansion. For example, in a specific implementation manner, assume that the data format of the image data is RGB format, and the spectral data includes infrared information and visible light information, where the data format of the visible light information is RGB format and the data format of the infrared information is grayscale image format. Then, the data format of the infrared information can be converted. Specifically, the single-channel grayscale image can be copied to the three channels (R, G, B) of the RGB image through channel expansion operation so that the pixel values of each channel are the same.
[0098] Furthermore, if the data formats of the spectral data and the image data are consistent, determine that the spectral data is the spectral data after format conversion, and continue with subsequent processing.
[0099] Step C20, if the resolutions of the spectral data after format conversion and the image data are inconsistent, adjust the resolution of the spectral data after format conversion to obtain target spectral data, where the resolutions of the new spectral data and the image data are consistent;
[0100] If the resolutions of the spectral data after format conversion and the image data are different, the resolution of the spectral data after format conversion can be adjusted to unify the resolutions of the spectral data and the image data, facilitating subsequent merging operations. Specifically, the resolution can be adjusted by means of interpolation.
[0101] Furthermore, if the resolutions of the spectral data after format conversion and the image data are the same, determine that the spectral data after format conversion is the target spectral data, and continue with subsequent processing.
[0102] Step C30, merge the image data and the spectral data after format conversion in the channel dimension to obtain input data.
[0103] Merge the input data of the image data and the target spectral data in the channel dimension. Specifically, the two data can be merged by taking the average value or weighted sum of each pixel point in the channel dimension.
[0104] Exemplarily, in a specific embodiment, assume that for a certain pixel point, the pixel value of the image data at this pixel point is (R1, G1, B1), the pixel value of the infrared information in the target spectral data at this pixel point is (R2, G2, B2), and the pixel value of the visible light information in the target spectral data at this pixel point is (R3, G3, B3). Then, merge this pixel point in the channel dimension and set the pixel value of this pixel point to [(R1 + R2 + R3) / 3, (G1 + G2 + G3) / 3, (B1 + B2 + B3) / 3]. In this way, merge each pixel point in turn to obtain the input data.
[0105] Exemplarily, to help understand the technical concept or principle of the gesture interaction process for controlling a smart TV after combining this embodiment with the above-mentioned Embodiment 1, a specific embodiment is listed. Refer to Figure 2 As shown, in this specific embodiment, the gesture interaction process for controlling a smart TV includes:
[0106] 1. The user makes a gesture action for operating the TV, and the TV captures the user's gesture action in real time through the integrated camera and depth sensor to collect the image data of the user's hand and perform real-time image preprocessing.
[0107] 2. According to the trained AI gesture recognition model (the AI recognition model shown in Figure 2 ), recognize the image data to obtain the target gesture recognition result, that is, determine the user's specified gesture, and then match the custom gesture in the preset gesture interaction database, which can be an action or an action combination.
[0108] 3. Command execution: According to the recognized custom gesture, obtain the interaction method associated with this gesture and generate the corresponding operation instruction, that is, map the recognized custom gesture to a specific operation instruction.
[0109] 4. The operation command is transmitted to the TV operating system without delay through the control interface.
[0110] 5. After receiving the command, the TV operating system immediately executes the corresponding operation, such as switching channels, adjusting the volume, starting an application program, etc., to achieve a quick response to the user's intention.
[0111] It should be noted that the above specific embodiment is only used to understand the present application and does not constitute a limitation to the gesture interaction process for controlling a smart TV in the present application. Based on this technical concept, more forms of simple transformations are within the protection scope of the present application.
[0112] Based on the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the content that is the same as or similar to the above-mentioned first embodiment and second embodiment can be referred to the above introduction and will not be elaborated hereinafter. On this basis, before the step of collecting the image data of the hand, the method further includes:
[0113] Step D10, in response to a gesture input request, obtain the input gesture data according to the gesture input request, where the gesture data at least includes the image data of the hand;
[0114] The gesture data refers to the relevant data collected when the user inputs a gesture, including but not limited to the image data of the hand, such as spectral data may be included.
[0115] Further, after receiving the gesture input request, a preset gesture example template can also be obtained, and the gesture example template can be displayed, so that the user can refer to these examples as templates or benchmarks to help the user more accurately define the gesture that they hope to create.
[0116] Further, during the gesture input process, a prompt message can also be output to prompt the user to move the gesture to a specified position for accurate input.
[0117] Further, in an environment with insufficient light, a prompt message can also be output to remind the user to turn on the light source to ensure the input effect.
[0118] Further, during the gesture input process, a prompt message can also be output to remind the user to change the position of the gesture, so as to remind the user to gradually change the position of the gesture according to the indication, ensure that the gesture is recorded from multiple angles, and capture the complete details of the custom gesture action.
[0119] Step D20, detect whether the gesture data is qualified;
[0120] Step D30, if it is qualified, train a preset gesture recognition model according to the gesture data to obtain the target gesture recognition model, and store the custom gesture corresponding to the gesture data in the preset gesture interaction database.
[0121] If the input gesture data is qualified, train a preset gesture recognition model according to the gesture data. Specifically, incrementally train the gesture recognition model to obtain a trained gesture recognition model, and store the input custom gesture in the gesture interaction database, and store the newly input gesture in the gesture interaction database in a timely manner for subsequent use. Among them, the preset gesture recognition model refers to the gesture recognition model deployed in the smart TV.
[0122] It should be noted that if the gesture data includes image data and spectral data, when incrementally training the gesture recognition model, the image data and spectral data can be fused, and the fused data is used as the input of the gesture recognition model for training. The specific fusion method of the image data and spectral data can refer to the above embodiments, and will not be elaborated in this embodiment.
[0123] Before training the preset gesture recognition model, the gesture data can also be preprocessed to improve the quality of the gesture data, and the preprocessed gesture data is used to train the recognition model, thereby improving the training effect of the gesture recognition model. Among them, the data preprocessing includes, but is not limited to, one or more of data denoising, data normalization, and data augmentation. Relevant personnel can also set the data preprocessing method based on actual needs, and this embodiment does not make specific restrictions on this.
[0124] It should be noted that when storing the newly entered custom gesture, the interaction method configured by the user for this newly entered custom gesture can also be obtained, and the configured interaction method is associated with the entered custom gesture and stored in the gesture interaction database.
[0125] In addition, before storing this entered custom gesture in the gesture interaction database, the similarity between this newly entered custom gesture and the existing gestures in the gesture interaction database can be compared. If there is no gesture in the gesture interaction database whose similarity with this newly entered custom gesture is greater than the preset similarity threshold, the newly entered custom gesture is stored in the gesture interaction database; if there is a gesture in the gesture interaction database whose similarity with this newly entered custom gesture is greater than the preset similarity threshold, corresponding prompt information (such as "This gesture already exists") can be output to prevent the user from forgetting the custom gesture entered previously and entering the same or similar gesture again, thereby avoiding the situation where the same or similar gestures are associated with different interaction methods and improving the robustness of gesture interaction in the subsequent use process.
[0126] In addition, if the entered gesture data is unqualified, a gesture re-entry request can be initiated to enable the user to re-enter the gesture, and corresponding prompt information (such as "The hand image is not completely entered", "The gesture is too simple") can also be output to prompt the user about the reason for the unqualified.
[0127] In this embodiment, before training the gesture recognition model, it is evaluated whether the entered gesture data is qualified. When the entered gesture data is qualified, the gesture recognition model is trained and the custom gesture is stored, thus avoiding the ineffective training of the gesture recognition model, reducing the phenomenon of unqualified gestures being entered, and ensuring the effective progress of subsequent gesture interaction.
[0128] In a possible implementation manner, the step of detecting whether the gesture data is qualified includes:
[0129] Step E10: Input the gesture data into a pre-trained gesture data evaluation model to obtain an evaluation result, where the gesture data evaluation model is a classification model based on machine learning;
[0130] The gesture evaluation model refers to a machine learning model used to evaluate whether gesture data is qualified. Its goal is to classify the input data into one of the qualified or unqualified categories. Specifically, it can be a support vector machine model, a random forest model, etc. This embodiment does not make specific limitations on this.
[0131] Step E20: If the evaluation result indicates that the data is qualified, determine that the gesture data is qualified;
[0132] Step E30: If the evaluation result indicates that the data is unqualified, determine that the gesture data is unqualified.
[0133] The gesture data evaluation model evaluates whether the gesture data is qualified based on certain judgment criteria and outputs a gesture evaluation result. For example, in a specific implementation manner, the gesture data evaluation model evaluates whether the gesture data contains a complete hand image. If the gesture data contains a complete hand image, the data is determined to be qualified; if the gesture data does not contain a complete hand image, the data is determined to be unqualified. Another example is that in another specific implementation manner, the gesture evaluation model evaluates the complexity of the gesture in the gesture data and the probability of incorrect operation of the gesture. If the complexity of the gesture is less than a certain threshold or the probability of incorrect operation of the gesture is greater than a certain threshold, the data is determined to be unqualified; if the complexity of the gesture is greater than or equal to a certain threshold and the probability of incorrect operation of the gesture is less than or equal to a certain threshold, the data is determined to be qualified.
[0134] In a possible implementation manner, the step of storing the custom gesture corresponding to the gesture data into the preset gesture interaction database includes:
[0135] Step F10: Generate a gesture recognition result corresponding to the gesture data according to the target gesture recognition model, and obtain a feedback result for the gesture recognition result;
[0136] The text description of the gesture result can be the text description result output by the target gesture recognition model after inputting the gesture data into the target gesture recognition model.
[0137] After obtaining the gesture recognition result corresponding to the gesture data, the gesture recognition result can be output for the user to verify whether the gesture recognition result is accurate and submit a feedback result.
[0138] Step F20: If the feedback result indicates that the gesture recognition result is correct, obtain the interaction method configured by the user.
[0139] When the feedback result indicates that the gesture recognition result is correct, the interaction method configuration interface can be switched to (or the user can also enter the interaction method configuration interface by themselves), so that the user can configure the interaction method in the interaction method configuration interface. Then, after the user configures the interaction method, the smart TV can obtain the interaction method configured by the user.
[0140] Step F30: Obtain the custom gesture corresponding to the gesture data, and store the configured interaction method and the entered custom gesture in association with the preset gesture interaction database.
[0141] It should be noted that the entered custom gesture refers to the text description content of the gesture input by the user when performing gesture input, such as "clenching a fist", "giving a thumbs up", etc.
[0142] Furthermore, if the feedback result indicates that the gesture recognition result is incorrect, a gesture re-entry request can be initiated to re-enter the gesture when there is a deviation between the recognition result of the model and the actual result, so as to avoid entering a custom gesture with a deviation in model recognition.
[0143] In this embodiment, after the gesture recognition model is trained, the corresponding gesture recognition result is generated according to the target gesture recognition model, and the feedback result input by the user for the gesture recognition result is received. If the feedback result indicates that the gesture recognition result is correct, the entered custom gesture and the configured interaction method are stored to store the gestures that the model can accurately recognize in the gesture interaction database in a timely manner.
[0144] Exemplarily, to help understand the technical concept or technical principle of this embodiment, a specific embodiment is listed for reference Figure 3 As shown, in this specific embodiment, the process of the gesture input and model training stage includes:
[0145] Gesture input: The user enters a custom gesture operation through the gesture input interface.
[0146] (1) The user enters a specially designed gesture input interface, where the gesture actions that the user hopes to customize can be conveniently displayed and input.
[0147] (2) Before input, a set of predefined gesture examples and corresponding operation commands (i.e., interaction methods) are provided. The user can refer to these examples as templates or benchmarks to help the user more accurately define the gestures they hope to create.
[0148] (3) Start the input operation. During the input process, the system will intelligently prompt the user to move the gesture to the specified position and make a custom gesture, which can also be a set of gesture actions. In an environment with insufficient light, the system will also kindly remind the user to turn on the light source and move to an environment with good lighting for input work to ensure the input effect. During the entire input process, the user gradually changes the gesture position according to the instructions to ensure that the gesture is recorded from multiple angles to record the complete details of the custom gesture action.
[0149] (4) After recording the gesture, call the gesture evaluation model through the API (Application Programming Interface) to evaluate the gesture. The gesture evaluation model evaluates whether the gesture is too simple and what the probability of misoperation is. If it is too simple or the probability is greater than 60%, a prompt message will be output, prompting the user to re-enter a more complex gesture action to avoid misoperation problems.
[0150] (5) After the input is completed, preprocessing will be performed on the image data, such as data cleaning and removing invalid data (such as blurred gestures and occluded data) to improve the effect of model training.
[0151] (6) Identify the user's gesture through the AI gesture recognition model and match the corresponding gesture action. When the AI recognition model completes the gesture recognition, a text description of the custom gesture action just shown by the user will be automatically generated. If the user believes that the recognition result is inaccurate, they can choose to cancel the operation and start re-entering the gesture.
[0152] Model training: Package and use the collected gesture data to request the relevant API for training the AI gesture recognition model.
[0153] (1) Using the collected gesture data, first, the data will be cleaned to remove invalid data, such as problems with blurred gestures and occlusions.
[0154] (2) Extract the key features of the gesture data as the input of the AI model (i.e., the target gesture recognition model). It can be understood that feature extraction can also be regarded as part of the AI model. At this time, the gesture data is the input of the AI model.
[0155] (3) Initialize the machine learning model and use the preprocessed data to train the model. During the model training process, cross-validation is used to adjust the model parameters.
[0156] (4) Use the test set to evaluate the model performance, adjust the model parameters according to the test results, and save the finally optimized AI gesture recognition model.
[0157] (5) The AI gesture recognition model generates a text describing the custom gesture and feeds it back to the user.
[0158] Gesture configuration: After the gesture is input, let the user select the corresponding TV operation command (i.e., the interaction method). After the user selects the command, the custom gesture will be bound to this TV command to associate the input gesture with the TV operation command.
[0159] The user can, according to personal preferences, configure the one-to-one or one-to-many association between the input gesture and the TV operation command. For example, the user can configure the gesture of "waving left" as "previous channel", and the gesture of "waving right" as "next channel", or a set of action combinations, such as setting the gesture of "fist clenching" followed by "thumbs up" as "play / pause".
[0160] It should be noted that the above specific embodiments are only used to understand the present application and do not constitute a limitation on the gesture input and model training processes of the present application. Based on this technical concept, more forms of simple transformations are within the protection scope of the present application.
[0161] In addition, as shown in Figure 4 the embodiments of the present application also provide a gesture interaction system, which includes a controller and a smart TV. The controller is connected to the smart TV, and the controller is used to execute the steps of the gesture interaction method for controlling the smart TV as described above.
[0162] In addition, the embodiments of the present application also propose a smart TV, which includes a memory, a processor, and a gesture interaction program for controlling the smart TV stored on the memory and executable on the processor. When the gesture interaction program for controlling the smart TV is executed by the processor, the steps of the gesture interaction method for controlling the smart TV as described above are implemented.
[0163] As Figure 5As shown, the smart TV may include a processing system 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage system 1003 into the random access memory (RAM) 1004. In the RAM 1004, various programs and data required for the operation of the smart TV are also stored. The processing system 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input system 1007 including, for example, a touch screen, a touch pad, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output system 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage system 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication system 1009. The communication system 1009 may allow the smart TV to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a smart TV with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or had alternatively.
[0164] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through a communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0165] The smart TV provided by the embodiments of the present application adopts the gesture interaction method for controlling the smart TV in the above embodiments, and can solve the technical problem of gesture interaction for controlling the smart TV. Compared with the prior art, the beneficial effects of the smart TV provided by the present application are the same as those of the gesture interaction method for controlling the smart TV provided by the above embodiments, and the other technical features in this smart TV are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0166] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0167] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0168] In addition, to achieve the above object, an embodiment of this application also provides a readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the gesture interaction method for controlling a smart TV in the above embodiments.
[0169] The computer-readable storage medium provided by the embodiment of this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0170] The above computer-readable storage medium can be included in a smart TV; it can also exist alone without being assembled into a smart TV.
[0171] The above computer-readable storage medium carries one or more programs, which, when executed by a smart TV, cause the smart TV to: collect image data of a hand, input the image data into a target gesture recognition model to obtain a target gesture recognition result; determine a target gesture in a preset gesture interaction database that matches the target gesture recognition result, where the gesture interaction database includes at least one custom gesture and an interaction method associated with each custom gesture; obtain the target interaction method associated with the target gesture in the gesture interaction database, and control the smart TV in the target interaction method.
[0172] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0173] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0174] The modules involved in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0175] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned gesture interaction method for controlling a smart TV, which can solve the technical problem of gesture interaction for controlling a smart TV. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the gesture interaction method for controlling a smart TV provided in the above embodiments, and will not be elaborated here.
[0176] In addition, an embodiment of the present application also proposes a computer program product, including a gesture interaction program for controlling a smart TV. When the gesture interaction program for controlling a smart TV is executed by a processor, the steps of the gesture interaction method for controlling a smart TV as described above are implemented.
[0177] The specific implementation manners of the computer program product of the present application are basically the same as those of the embodiments of the above-mentioned gesture interaction method for controlling a smart TV, and will not be elaborated here.
[0178] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or system including the element.
[0179] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0180] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software sensor. The computer software sensor is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.
[0181] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A gesture interaction method for controlling a smart TV, characterized in that: The gesture interaction method for controlling a smart TV comprises the following steps: Collecting image data of the hand, and inputting the image data into a target gesture recognition model to obtain a target gesture recognition result; Determine a target gesture in a preset gesture interaction database that matches the target gesture recognition result, wherein the gesture interaction database includes at least one custom gesture and an interaction mode associated with each custom gesture; A target interaction mode associated with the target gesture in the gesture interaction database is obtained, and the smart TV is controlled in the target interaction mode.
2. The method according to claim 1, characterized in that The step of determining a target gesture in a preset gesture interaction database that matches the target gesture recognition result comprises: Performing vectorization processing on the target gesture recognition result to obtain a gesture result vector, and calculating the similarity between the gesture result vector and each gesture representation vector in a preset gesture interaction database; Determine the highest similarity among all the similarities, and if the highest similarity is greater than or equal to a preset similarity threshold, determine the gesture representation vector corresponding to the highest similarity as the target gesture representation vector; A gesture associated with the target gesture representation vector in the gesture interaction database is determined as a target gesture.
3. The method according to claim 1, characterized in that The step of inputting the image data into a target gesture recognition model to obtain a target gesture recognition result comprises: Acquiring spectral data of a hand, wherein the spectral data includes infrared information and / or visible light information; The image data and the spectral data are fused to obtain input data, and the input data is input into a target gesture recognition model to obtain a target gesture recognition result.
4. The method according to claim 3, characterized in that The step of fusing the image data with the spectral data to obtain input data comprises: If the data formats of the spectral data and the image data are inconsistent, converting the data format of the spectral data to obtain the spectral data after format conversion, wherein the data formats of the image data and the spectral data after format conversion are consistent, wherein the data format is a color format of the data; If the resolution of the spectral data after format conversion is inconsistent with that of the image data, adjusting the resolution of the spectral data after format conversion to obtain target spectral data, wherein the resolution of the target spectral data is consistent with that of the image data; The image data and the target spectrum data are combined in the channel dimension to obtain input data.
5. The method according to any one of claims 1 to 4, characterized in that: Before the step of collecting image data of the hand, the method further includes: In response to a gesture entry request, acquiring entered gesture data according to the gesture entry request, wherein the gesture data at least includes image data of a hand; Detecting whether the gesture data is qualified; If qualified, a preset gesture recognition model is trained according to the gesture data to obtain the target gesture recognition model, and the custom gesture entered corresponding to the gesture data is stored in the preset gesture interaction database.
6. The method according to claim 5, characterized in that The step of detecting whether the gesture data is qualified comprises: Inputting the gesture data into a pre-trained gesture data evaluation model to obtain an evaluation result, wherein the gesture data evaluation model is a classification model based on machine learning; If the evaluation result indicates that the data is qualified, determining that the gesture data is qualified; If the evaluation result indicates that the data is unqualified, it is determined that the gesture data is unqualified.
7. The method according to claim 5, characterized in that The step of storing the custom gesture corresponding to the gesture data entered into the preset gesture interaction database includes: generating a gesture recognition result corresponding to the gesture data according to the target gesture recognition model, and obtaining a feedback result for the gesture recognition result; If the feedback result indicates that the gesture recognition result is correct, obtaining the interaction mode configured by the user; The custom gesture entered corresponding to the gesture data is obtained, and the configured interaction mode is associated with the entered custom gesture and stored in the preset gesture interaction database.
8. A gesture interaction system, characterized in that: The gesture interaction system includes a controller and a smart TV, wherein the controller is connected to the smart TV, and the controller is used to execute the steps of the gesture interaction method for controlling the smart TV according to any one of claims 1 to 7.
9. A smart TV, characterized in that: The smart TV comprises: a memory, a processor, and a gesture interaction program stored in the memory and executable on the processor. When the gesture interaction program is executed by the processor, the steps of the gesture interaction method for controlling the smart TV as described in any one of claims 1 to 7 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium is a computer-readable storage medium, on which is stored a program for implementing the gesture interaction method, and the program for implementing the gesture interaction method is executed by a processor to implement the steps of the gesture interaction method for controlling a smart TV as described in any one of claims 1 to 7.