Non-contact interaction method, system and device of LED screen and medium
Through the combination of millimeter radar wave array and convolutional neural network model, a dynamic gesture recognition algorithm and multi-user conflict processing strategy are designed, which solves the problems of precise control and multi-user operation in non-contact interaction of LED large screens, and realizes efficient gesture detection and smooth interaction with all-round blind spots.
Patent Information
- Application Number
- CN202510359174.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-08
AI Technical Summary
The existing contactless interaction technology of LED large-screen large-screen cannot achieve accurate interactive control, especially in the simultaneous operation of multiple users and complex dynamic trajectory recognition, and the existing gesture recognition technology cannot meet the special needs of LED large-screens.
The millimeter radar wave array is used to collect gesture operation data, combine gesture recognition algorithms and convolutional neural network models for gesture recognition, design dynamic gesture recognition algorithms and multi-user conflict processing strategies, and achieve all-round blind spot-free gesture detection and smoothness of multi-user operation through millimeter wave radar technology.
It realizes contactless gesture interaction with large LED screens, has strong anti-interference, adapts to multiple environments, can handle complex gesture trajectories, and ensures the smoothness of interaction between multiple users at the same time.
Smart Images

Figure CN120276597A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction, and particularly relates to a non-contact interaction method, system, device and medium for an LED screen. Background Art
[0002] With the development of display technology, LED large screens have been widely used in fields such as commercial display, information release, and conference demonstration. Traditional interaction methods for LED large screens mainly include physical buttons, infrared remote controls, touch screens, camera vision recognition, etc. However, physical buttons and touch screens require contact operations, which have hygiene and durability problems. Infrared remote control requires the installation of complex sensing devices on the screen surface, which not only increases costs but is also easily physically damaged, and the detection distance is limited. Camera vision recognition has high requirements for light conditions. In strong light or low light environments, the recognition effect will be greatly reduced, and there are also problems such as occlusion and privacy restrictions.
[0003] In recent years, non-contact interaction technology has gradually become a research hotspot. Among them, vision-based gesture recognition technology is a relatively common one, such as non-contact gesture interaction with an LED large screen by capturing human gestures through a camera.
[0004] However, most of the existing gesture recognition technologies are designed for static gestures or small-screen devices, and for large-size and high-resolution display devices such as LED large screens, accurate interaction control cannot be achieved. Although millimeter-wave radar technology has been applied in some fields, such as the vehicle environment, it mainly focuses on general gesture recognition and does not fully consider the special requirements of LED large screens, such as multi-user simultaneous operation and complex dynamic trajectory recognition. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a non-contact interaction method, system, device and medium for an LED screen. In view of the characteristics of the LED large screen, a dynamic gesture recognition algorithm is designed, which can process complex gesture trajectories, and through the proposed multi-user conflict handling strategy, the interaction fluency when multiple users operate simultaneously is ensured.
[0006] To achieve the above purpose, the embodiments of the present invention provide a non-contact interaction method for an LED screen, including: Collecting gesture operation data of users in front of the LED screen by using a millimeter radar wave array; Adopting a gesture recognition algorithm to extract the time-frequency features of the gesture operation data, and inputting the extracted time-frequency features into a pre-trained convolutional neural network model for gesture recognition to obtain the gesture type and whether it is a multi-user operation; When it is determined that it is a multi - user operation gesture, obtain the position of the user relative to the LED screen, and assign priorities to each user according to this position and / or gesture type; Map the gesture type and user priorities to the control instructions of the LED screen to achieve real - time response of the LED screen to user gestures.
[0007] Optionally, after collecting the gesture operation data of the users in front of the LED screen using a millimeter - wave radar array, the non - contact interaction method of the LED screen further includes: Perform noise reduction and filtering processing on the gesture operation data to remove the interference data in the gesture operation data and obtain target gesture operation data.
[0008] Optionally, adopt a gesture recognition algorithm to extract the time - frequency features of the gesture operation data, including: Segment the gesture operation data by time period to obtain multiple time - period data; Perform Fourier transform on each time - period data to obtain the spectrum varying with time; Collect the peak value, spectrum bandwidth, and spectrum moment parameters in the spectrum based on the spectrum, and extract the time - frequency features of the gesture operation data based on the peak value, spectrum bandwidth, and spectrum moment parameters. Among them, the peak value in the spectrum corresponds to the vibration or motion mode generated during the execution of the gesture, which is used to distinguish different gestures, the spectrum bandwidth is used to distinguish the complexity and dispersion degree of the gesture, and the spectrum moment is used to characterize the spectral shape features of the gesture signal.
[0009] Optionally, the training process of the convolutional neural network model includes: Design multiple convolutional layers to extract the local features of the spectrum data after Fourier transform of the input historical single - user and multi - user gesture data. Among them, each convolutional layer contains multiple convolutional kernels, which are used to slide on the input spectrum data and calculate the convolution result; Add a pooling layer after the convolutional layer to reduce the dimension of the spectrum data; Design a first classifier and a second classifier after the pooling layer. Among them, the first classifier is used for classifying the gesture type, and the second classifier is used for judging multi - user operations; Set a loss function to determine the difference between the gesture category output by the first classifier and the actual gesture category, and the difference between the multi - user operation classification result output by the second classifier and the actual user operation classification result. Then, select an optimizer to update the weights of the convolutional neural network model according to this difference to minimize the loss function and obtain a trained convolutional neural network model.
[0010] Optionally, input the extracted time-frequency features into a pre-trained convolutional neural network model for gesture recognition to obtain the gesture type and whether it is a multi-user operation, including: Input the extracted time-frequency features into the second classifier of the pre-trained convolutional neural network model, so that the second classifier extracts the running trajectory of the gesture from the time-frequency features; Judge whether it is a multi-user operation according to the number and distribution characteristics of the intersection and separation feature points between the running trajectories.
[0011] Optionally, when it is determined that it is a multi-user operation gesture, obtain the position of the user relative to the LED screen, and assign priorities to each user according to the position and / or gesture type, including: Regarding the position of the user relative to the LED screen, the priority of the user closer to the LED screen is higher than that of the user farther from the LED screen; Regarding the gesture type, the priority of the user with a more complex gesture is higher than that of the user with a simpler gesture; Regarding the position of the user relative to the LED screen and the gesture type, set weights for the user position and gesture type respectively, and calculate a weighted score according to the priority values of the user position and gesture type. The higher the weighted score, the higher the priority.
[0012] Optionally, the non-contact interaction method of the LED screen further includes: Install millimeter-wave radars at preset positions on the LED screen to form a rectangular array. Among them, the setting requirements of the preset positions include but are not limited to that the detection ranges of adjacent millimeter-wave radars overlap, and the coverage range of the rectangular array meets the preset requirements.
[0013] Through the above technical solutions, non-contact gesture interaction with the LED large screen is realized based on millimeter-wave radar technology. The interaction technology has strong anti-interference ability (not affected by light conditions) and strong penetration ability (can adapt to environments such as rain, fog, and dust); the collaborative working mechanism of multiple radar arrays realizes all-round and blind-zone-free gesture detection; aiming at the characteristics of the LED large screen, a dynamic gesture recognition algorithm is designed, which can process complex gesture trajectories; a multi-user conflict handling strategy is proposed to ensure the interaction fluency when multiple users operate simultaneously.
[0014] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. Description of the Drawings
[0015] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification. Together with the following specific implementation manners, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings: Figure 1It is a flowchart of a non-contact interaction method for an LED screen provided by an embodiment of the present invention; Figure 2 It is a flowchart of extracting time-frequency features of gesture operation data by using a gesture recognition algorithm provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of a non-contact interaction system for an LED screen provided by an embodiment of the present invention; Figure 4 It is a flowchart of the execution process of a non-contact interaction system for an LED screen provided by an embodiment of the present invention; Figure 5 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0016] In the following detailed description, various embodiments of the present disclosure will be described more fully. The present disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents, and / or alternative solutions falling within the spirit and scope of the various embodiments of the present disclosure.
[0017] In the following, the term "comprising" or "may comprise" that may be used in various embodiments of the present disclosure indicates the presence of the disclosed function or operation, and does not limit the addition of one or more functions or operations. Further, as used in various embodiments of the present disclosure, the terms "comprising", "having" and their cognates are only intended to represent a specific feature, number, step, operation, or combination of the foregoing items, and should not be construed as first excluding the existence or addition of the possibility of one or more other features, numbers, steps, operations, or combinations of the foregoing items.
[0018] In various embodiments of the present disclosure, the expression "or" or "at least one of A or / and B" includes any combination or all combinations of the listed words. For example, the expression "A or B" or "at least one of A or / and B" may include A, may include B, or may include both A and B.
[0019] Next, the technical solutions in the embodiments of the present invention will be described clearly and completely with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] Refer to Figure 1The following is a flowchart of a non-contact interaction method for an LED screen in a specific embodiment, including the following execution steps: Step 100: Use a millimeter-wave radar array to collect gesture operation data of the user in front of the LED screen.
[0021] Exemplarily, the millimeter-wave radar can adopt an integrated multi-channel radar that supports the transmission and reception of high-frequency signals (such as 60 GHz or 77 GHz), or a frequency-modulated continuous-wave (FMCW) radar with a working frequency of 60 - 64 GHz and an integrated multi-input multi-output (MIMO) antenna array for real-time collection of gesture reflection signals. These millimeter-wave radars transmit and receive millimeter-wave signals at a frequency of 50 frames per second to obtain the original data of the gestures.
[0022] In some embodiments, before performing step 100, millimeter-wave radars are installed at preset positions on the LED screen to form a rectangular array. Among them, the setting requirements of the preset positions include, but are not limited to, that the detection ranges of adjacent millimeter-wave radars overlap, and the coverage range of the rectangular array meets the preset requirements.
[0023] Exemplarily, a millimeter-wave radar is installed at each of the four corners of the large LED screen to form a rectangular array, which can achieve full coverage of a certain area in front of the large screen. The detection range of each millimeter-wave radar covers an area from 2 meters to 5 meters in front of the large screen, and the detection ranges of adjacent radars have a certain overlap to ensure there are no blind spots. It is used to capture the original data of the user's gestures in real time, such as collecting "swipe" gestures, "zoom" gestures, etc. made by the user three meters away from the screen.
[0024] It should be understood that the detection range of the millimeter-wave radar is determined and set according to the specific model of the millimeter-wave radar, and is not limited here.
[0025] In some embodiments, after performing step 100, the following steps are also performed: perform noise reduction and filtering processing on the gesture operation data to remove the interference data in the gesture operation data and obtain the target gesture operation data.
[0026] Exemplarily, configure an FPGA (Field Programmable Gate Array) or a dedicated DSP (Digital Signal Processor) chip to be able to process the millimeter-wave signals transmitted and received by the millimeter-wave radar at a frequency of 50 frames per second. Then, use the median filtering algorithm to remove the interference of noise, perform noise reduction and filtering processing on the original data, and extract the effective dynamic features of the gestures, such as the speed, acceleration, trajectory, etc. of the "swipe" and "zoom" gestures made by the user.
[0027] Step 101: Use a gesture recognition algorithm to extract the time-frequency features of the gesture operation data, and input the extracted time-frequency features into a pre-trained convolutional neural network model for gesture recognition to obtain the gesture type and whether it is a multi-user operation.
[0028] Specifically, referring to Figure 2 As shown, when performing Step 101 to use a gesture recognition algorithm to extract the time-frequency features of the gesture operation data, the following steps can be specifically executed: S1010: Segment the gesture operation data by time period to obtain multiple time period data.
[0029] Specifically, according to a preset segmentation standard, calculate the start point and end point of each time period. For example, when segmenting the data by minute, calculate the corresponding start time and end time for each minute. According to the calculated segmentation points, segment the original gesture operation data into multiple time period data. This can be achieved through programming, such as using loop and conditional statements in programming languages like Python to traverse the data and perform segmentation. Store the segmented data segments in an appropriate storage medium, such as a file, database, etc. At the same time, for the convenience of subsequent processing and analysis, time tags or indexes can be added to each data segment to identify the time period to which it belongs.
[0030] S1011: Perform Fourier transform on each time period data to obtain a spectrum that changes with time.
[0031] Specifically, perform Fourier transform on the data of each time period, convert it from the time domain to the frequency domain. The result of the transform is a complex number array, representing the amplitude and phase of different frequency components. Extract the amplitude information of each frequency component from the result of the Fourier transform, and combine the spectrum amplitude information of different time periods to construct a spectrum diagram that changes with time. Among them, the horizontal axis of the spectrum diagram represents time, the vertical axis represents frequency, and the color or grayscale represents the spectrum amplitude.
[0032] S1012: Based on the spectrum, collect the peak value, spectrum bandwidth, and spectrum moment parameters in the spectrum, and extract the time-frequency features of the gesture operation data based on the peak value, spectrum bandwidth, and spectrum moment parameters.
[0033] Among them, the peak value in the spectrum corresponds to the vibration or motion pattern generated during the execution of the gesture, which is used to distinguish different gestures. The spectrum bandwidth is used to distinguish the complexity and dispersion degree of the gesture, and the spectrum moment is used to characterize the spectrum shape feature of the gesture signal.
[0034] It should be understood that the peaks in the spectrum usually represent a strong concentration of signal energy at specific frequencies. In gesture analysis, these peaks may correspond to specific vibrations or motion patterns generated during the execution of gestures. Therefore, the position and intensity of the peaks can be used to distinguish different gestures, as they may produce different energy distributions at different frequencies. The spectral bandwidth refers to the frequency range occupied by the signal in the spectrum. In gesture analysis, the spectral bandwidth reflects the degree of dispersion of the gesture signal. A wider spectral bandwidth may mean that the gesture contains multiple frequency components, indicating a more complex or diverse motion pattern of the gesture. On the contrary, a narrower spectral bandwidth may indicate that the gesture is relatively simple or single. Therefore, the spectral bandwidth can be used as an indicator to evaluate the complexity and dispersion degree of gestures. The spectral moment is a quantitative description of the spectral shape. Specifically, the first moment (mean) of the spectrum reflects the central frequency of the spectrum, that is, the main frequency position where the signal energy is concentrated; while the second moment (variance or standard deviation) reflects the degree of dispersion of the spectrum, that is, the distribution width of the signal energy in the frequency domain. In higher-order moments, more information about the spectral shape, such as skewness and kurtosis, can also be included.
[0035] In gesture analysis, the spectral moments can be used to characterize the spectral shape features of gesture signals. For example, the first moment can indicate the main frequency components of the gesture, while the second moment can reflect the spectral bandwidth or dispersion degree of the gesture signal. These characteristic information is of great significance for the accurate classification and recognition of gestures, as they can provide information about the overall structure and distribution of gesture signals in the frequency domain.
[0036] Exemplarily, different gestures such as swiping, clicking, making a fist, rotating, zooming, etc.
[0037] In some embodiments, the training process of the convolutional neural network model includes the following steps: S1: Design multiple convolutional layers to extract local features of the spectrum data after performing Fourier transform on the input historical single-user and multi-user gesture data.
[0038] Exemplarily, taking 2 convolutional layers as an example, where the first convolutional layer uses a (3, 3) convolutional kernel, the number of convolutional kernels is 32, and the stride of the convolutional kernel sliding on the input data is set to (1, 1) to capture local time-frequency features. The second convolutional layer uses a (5, 5) convolutional kernel and 128 convolutional kernels for further feature extraction.
[0039] S2: Add a pooling layer after the convolutional layer to reduce the dimension of the spectrum data.
[0040] Optionally, the pooling layer can select a maximum pooling layer or an average pooling layer according to the actual application scenario, wherein the maximum value in the pooling window is selected as the output in the maximum pooling layer, and the average value of all values in the pooling window is calculated as the output in the average pooling layer, and a 2x2 pooling window is used to process 2x2 spectrum data blocks each time. The pooling layer slides on the input data according to the 2x2 pooling window and the preset step size, and performs a pooling operation in each window to obtain an output information with reduced dimensionality in the spatial dimension. Among them, if the pooling window slides in the time dimension, the time resolution of the output data will be reduced, and if the pooling window slides in the frequency dimension, the frequency resolution of the output data will be reduced.
[0041] S3: Design a first classifier and a second classifier after the pooling layer, wherein the first classifier is used to classify gesture types, and the second classifier is used to determine multi-user operations.
[0042] Specifically, the first classifier is a gesture type classifier, whose input layer receives feature maps from the pooling layer. These feature maps have been feature extracted and reduced in dimension by the previous convolution and pooling layers. Fully connected layer One or more fully connected layers are used to further integrate and extract features. These layers flatten the feature maps into one-dimensional vectors, perform linear transformations through weight matrices, and then apply activation functions (such as ReLU) to increase nonlinearity. The output layer uses a fully connected layer with a softmax activation function to output the probability distribution of each gesture type. The number of neurons in the output layer is equal to the number of gesture types. Classification process: The input feature map is forward propagated through the fully connected layer, and the output layer calculates the probability of each gesture type, and selects the gesture type with the highest probability as the classification result.
[0043] The second classifier is a multi-user operation judgment classifier. The independent input layer directly extracts information from the original data or features that have been preprocessed differently. The feature extraction layer uses convolutional layers and pooling layers to extract features. If features are shared, they are directly connected to one or more fully connected layers. The fully connected layer and output layer are similar to the gesture type classifier, but the output layer has only one neuron and uses the sigmoid activation function to output a probability value between 0 and 1, indicating the possibility of multi-user operation.
[0044] S4: Setting a loss function to determine the difference between the gesture category output by the first classifier and the actual gesture category, and the difference between the multi-user operation classification result output by the second classifier and the actual user operation classification result, and selecting an optimizer to update the weights of the convolutional neural network model based on the difference to minimize the loss function and obtain a trained convolutional neural network model.
[0045] Specifically, in step 101, the extracted time-frequency features are input into a pre-trained convolutional neural network model for gesture recognition to obtain the gesture type and whether it is a multi-user operation, including: The extracted time-frequency features are input into the second classifier of the pre-trained convolutional neural network model, so that the second classifier extracts the running trajectory of the gesture from the time-frequency features; according to the number and distribution characteristics of the intersection and separation feature points between the running trajectories, it is judged whether it is a multi-user operation.
[0046] Exemplarily, in the second classifier, structures such as convolutional layers and pooling layers are used to map and transform the input time-frequency features to extract local and global features of the gesture. By connecting the outputs of convolutional layers or fully connected layers and combining time series analysis, the motion trajectory of the gesture is reconstructed, and the reconstructed gesture trajectory is optimized to remove noise and outliers, improving the accuracy and reliability of the trajectory. Geometric shape features of the trajectory, such as length, width, curvature, etc., are extracted to describe the contour and shape of the gesture. Dynamic features such as speed and acceleration of the trajectory are extracted to describe the motion state and change trend of the gesture. In a multi-user operation scenario, interaction features such as intersection and separation between trajectories are extracted to determine whether it is a multi-user operation. After extracting the gesture running trajectory, a threshold or rule set is set, and it is judged whether it is a multi-user operation according to parameters such as the number, density, and duration of intersection and separation feature points. For example, if the number of intersection points between trajectories exceeds a certain threshold, or the distribution of separation feature points is too wide, it may be judged as a multi-user operation.
[0047] Step 102: If it is determined that the gesture is a multi-user operation, obtain the position of the user relative to the LED screen, and assign priorities to each user according to this position and / or gesture type.
[0048] Specifically, when executing step 102, the following steps can be specifically executed: S1020: For the position of the user relative to the LED screen, the priority of the user closer to the LED screen is higher than that of the user farther from the LED screen.
[0049] S1021: For the gesture type, the priority of the user with complex gestures is higher than that of the user with simple gestures.
[0050] Exemplarily, complex gestures include but are not limited to multi-finger zooming, rotation gestures, and multi-touch; simple gestures include but are not limited to single-click and swipe.
[0051] S1022: For the position of the user relative to the LED screen and the gesture type, weights are set for the user position and gesture type respectively, and a weighted score is calculated according to the priority values of the position of the user relative to the LED screen and the gesture type. The higher the weighted score, the higher the priority.
[0052] Exemplarily, set the user position weight. Among them, for a short distance (such as within 1 meter): set the weight to 0.7. Because the user is close to the screen, their interaction actions have a greater impact on the screen and are more likely to be attracted by the screen content; for a medium distance (such as 1 to 3 meters): set the weight to 0.5. The user can clearly see the screen content at this distance, but the interaction actions have a relatively small impact on the screen; for a long distance (such as more than 3 meters): set the weight to 0.3. The user may not be able to clearly see the screen details at this distance, and the interaction actions also have a small impact on the screen. Set the gesture type weight. Among them, for simple gestures (such as single click, swipe), set the weight to 0.4; for complex gestures (such as multi-touch, rotation, zoom), set the weight to 0.6. Suppose there are two users, user A and user B, located at different positions on the LED screen and using different gesture types for interaction. The weighted scores of them can be calculated according to the following steps: User A, position: short distance (within 1 meter), gesture type: complex gesture (multi-touch), weighted score calculation: 0.7 (position weight) * 0.6 (gesture type weight) = 0.42; User B, position: medium distance (1 to 3 meters), gesture type: simple gesture (single click), weighted score calculation: 0.5 (position weight) * 0.4 (gesture type weight) = 0.20. The weighted score of user A (0.42) is higher than the weighted score of user B (0.20), so the interaction priority of user A is higher.
[0053] Step 103: Map the gesture type and user priority to the control instructions of the LED screen to achieve real-time response of the LED screen to the user's gestures.
[0054] Specifically, map the extracted and processed gesture signals to the corresponding LED large screen control instructions, such as "swipe", "zoom", etc. Then, send these instructions to the LED large screen control system to achieve real-time response to gestures. For example, if the "swipe" gesture is recognized, it is mapped to the "switch screen" instruction; if the "zoom" gesture is recognized, it is mapped to the "adjust image size" instruction. So that the LED display terminal can correctly receive the instructions and update the display content in real time.
[0055] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0056] Through the above technical solutions, non-contact gesture interaction with the LED large screen is realized based on millimeter-wave radar technology. The interaction technology has strong anti-interference ability (not affected by light conditions) and strong penetration ability (can adapt to environments such as rain, fog, and dust); the cooperative working mechanism of multiple radar arrays realizes all-round and blind-zone-free gesture detection; aiming at the characteristics of the LED large screen, a dynamic gesture recognition algorithm is designed, which can process complex gesture trajectories; a multi-user conflict handling strategy is proposed to ensure the interaction fluency when multiple users operate simultaneously.
[0057] As Figure 3 shown, the following is an embodiment of the non-contact interaction system of the LED screen provided by the embodiments of the present disclosure, which belongs to the same inventive concept as the non-contact interaction method of the LED screen in the above embodiments. For the details not described in detail in the embodiment of the non-contact interaction system of the LED screen, reference can be made to the embodiments of the non-contact interaction method of the LED screen above.
[0058] The non-contact interaction system of the LED screen includes: A millimeter-wave radar unit for collecting gesture operation data of users in front of the LED screen; A gesture recognition unit for using a gesture recognition algorithm to extract the time-frequency features of the gesture operation data, and inputting the extracted time-frequency features into a pre-trained convolutional neural network model for gesture recognition to obtain the gesture type and whether it is a multi-user operation; A multi-user processing unit for, when it is determined that the gesture is a multi-user operation gesture, obtaining the positions of the users relative to the LED screen, and assigning priorities to each user according to the positions and gesture types; An instruction mapping and execution unit for mapping the gesture type and user priorities to the control instructions of the LED screen to realize the real-time response of the LED screen to the user gestures.
[0059] Refer to Figure 4 shown, which is a flowchart of the execution process of a non-contact interaction system of an LED screen provided by an embodiment of the present invention. The millimeter-wave radar array is used to capture user gesture data, and the user gesture data is input into the signal processing unit. Then, it is judged by the gesture recognition unit whether there are multiple users. If not, it enters the instruction mapping and execution unit for output display. If so, it enters the multi-user management unit, then enters the instruction mapping and execution unit, and finally outputs the display.
[0060] Figure 5 It is a schematic hardware structure diagram of an electronic device for implementing various embodiments of the present invention.
[0061] The non-contact interaction method of the LED screen provided by the embodiments of the present application can be applied to an electronic device. Those skilled in the art can understand that the structure of the electronic device involved in the embodiments of the present invention does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.
[0062] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, a key, a camera, a display screen, and a SIM card interface, etc.
[0063] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0064] The processor may include one or more processing units. For example, the processor may include a central processing unit (CPU), etc., an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0065] Among them, the processor can be the nerve center and command center of the electronic device. The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0066] A memory can also be set in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can save the instructions or data that the processor has just used or recycled. If the processor needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor, and thus improves the system efficiency.
[0067] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to implement the storage capacity expansion of the electronic device. The external memory card communicates with the processor through the external memory interface to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.
[0068] The internal memory can be used to store computer-executable program codes, and the computer-executable program codes include instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. The internal memory can include a program storage area and a data storage area. The internal memory can include a high-speed random access memory and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0069] The wireless communication function of the electronic device can be implemented through an antenna, a wireless communication module, a modulation and demodulation processor, a baseband processor, etc.
[0070] The wireless communication module can provide wireless communication solutions applied to the electronic device, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0071] The electronic device can implement audio functions, etc. through an audio module, a speaker, a receiver, a microphone, a headphone interface, an application processor, etc.
[0072] An electronic device can implement a shooting function through an ISP, a camera, a video codec, a GPU, a display screen, an application processor, etc.
[0073] An electronic device can implement a display function through a GPU, a display screen, an application processor, etc.
[0074] The GPU is a microprocessor for image processing, connecting the display screen and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor may include one or more GPUs, which execute program instructions to generate or change display information.
[0075] The display screen is used to display images, videos, etc. The display screen includes a display panel.
[0076] In the storage medium provided by this application, there is a program product that can implement the non-contact interaction method of the LED screen.
[0077] The non-contact interaction method of the LED screen includes: using a millimeter radar wave array to collect gesture operation data of users in front of the LED screen; adopting a gesture recognition algorithm to extract the time-frequency characteristics of the gesture operation data, and inputting the extracted time-frequency characteristics into a pre-trained convolutional neural network model for gesture recognition to obtain the gesture type and whether it is a multi-user operation; if it is determined to be a multi-user operation gesture, obtain the position of the user relative to the LED screen, and assign priorities to each user according to this position and / or gesture type; map the gesture type and user priorities to the control instructions of the LED screen to achieve the real-time response of the LED screen to the user's gestures.
[0078] In some possible implementation manners, the subject matter name of the present disclosure, the non-contact interaction method and system of the LED screen, can be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to make the terminal device execute the steps according to various exemplary embodiments of the present disclosure described in the above "exemplary method" section of this specification.
[0079] The storage medium of the present disclosure can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0080] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A non-contact interaction method for an LED screen, characterized in that, Including: Using a millimeter radar wave array to collect gesture operation data of users in front of the LED screen; Adopting a gesture recognition algorithm to extract the time-frequency features of the gesture operation data, and inputting the extracted time-frequency features into a pre-trained convolutional neural network model for gesture recognition to obtain the gesture type and whether it is a multi-user operation; If it is determined to be a multi-user operation gesture, obtain the position of the user relative to the LED screen, and assign priorities to each user according to this position and / or gesture type; Map the gesture type and user priorities to the control instructions of the LED screen to achieve the real-time response of the LED screen to the user's gestures.
2. The non-contact interaction method of the LED screen according to claim 1, wherein After using a millimeter radar wave array to collect gesture operation data of users in front of the LED screen, the non-contact interaction method of the LED screen further includes: Performing denoising and filtering processing on the gesture operation data to remove the interference data in the gesture operation data and obtain the target gesture operation data.
3. The non-contact interaction method of the LED screen according to claim 1, characterized in that, Adopting a gesture recognition algorithm to extract the time-frequency features of the gesture operation data, including: Segmenting the gesture operation data by time period to obtain multiple time period data; Performing Fourier transform on each time period data to obtain the spectrum that changes with time; Collecting the peak value, spectrum bandwidth, and spectrum moment parameters in the spectrum based on the spectrum, and extracting the time-frequency features of the gesture operation data based on the peak value, spectrum bandwidth, and spectrum moment parameters. Among them, the peak value in the spectrum corresponds to the vibration or motion mode generated during the execution of the gesture, which is used to distinguish different gestures, the spectrum bandwidth is used to distinguish the complexity and dispersion degree of the gesture, and the spectrum moment is used to characterize the spectrum shape feature of the gesture signal.
4. The non-contact interaction method of the LED screen according to claim 1, characterized in that, The training process of the convolutional neural network model includes: Designing multiple convolutional layers to extract the local features of the spectrum data after Fourier transform of the input historical single-user and multi-user gesture data. Among them, each convolutional layer contains multiple convolutional kernels, which are used to slide on the input spectrum data and calculate the convolution result; Adding a pooling layer after the convolutional layer to reduce the dimension of the spectrum data; Designing a first classifier and a second classifier after the pooling layer, where the first classifier is used for classifying the gesture type, and the second classifier is used for judging multi-user operations; Setting a loss function to determine the difference between the gesture category output by the first classifier and the actual gesture category, and the difference between the multi-user operation classification result output by the second classifier and the actual user operation classification result, and selecting an optimizer to update the weights of the convolutional neural network model according to this difference to minimize the loss function and obtain the trained convolutional neural network model.
5. The non-contact interaction method of the LED screen according to claim 4, wherein Inputting the extracted time-frequency features into a pre-trained convolutional neural network model for gesture recognition to obtain the gesture type and whether it is a multi-user operation, including: Inputting the extracted time-frequency features into the second classifier of the pre-trained convolutional neural network model so that the second classifier extracts the running trajectory of the gesture from the time-frequency features; Judging whether it is a multi-user operation according to the number and distribution characteristics of the intersection and separation feature points between the running trajectories.
6. The non-contact interaction method of the LED screen according to claim 1, characterized in that, When it is determined that it is a multi - user operation gesture, obtain the position of the user relative to the LED screen, and assign priorities to each user according to this position and / or gesture type, including: Regarding the position of the user relative to the LED screen, the priority of the user closer to the LED screen is higher than that of the user farther from the LED screen; Regarding the gesture type, the priority of the user with a complex gesture is higher than that of the user with a simple gesture; Regarding the position of the user relative to the LED screen and the gesture type, set weights for the user position and gesture type respectively, and calculate a weighted score according to the priority values of the position of the user relative to the LED screen and the gesture type. The higher the weighted score, the higher the priority.
7. The non-contact interaction method of the LED screen according to claim 1, wherein, The non - contact interaction method of the LED screen further includes: Install millimeter - wave radars at preset positions on the LED screen to form a rectangular array. Among them, the setting requirements for the preset positions include, but are not limited to, that the detection ranges of adjacent millimeter - wave radars overlap, and the coverage range of the rectangular array meets the preset requirements.
8. A non-contact interaction system for an LED screen, characterized in that, Including: A millimeter - wave radar unit for collecting gesture operation data of users in front of the LED screen; A gesture recognition unit for using a gesture recognition algorithm to extract the time - frequency features of the gesture operation data, and inputting the extracted time - frequency features into a pre - trained convolutional neural network model for gesture recognition to obtain the gesture type and whether it is a multi - user operation; A multi - user processing unit for, when it is determined that it is a multi - user operation gesture, obtaining the position of the user relative to the LED screen, and assigning priorities to each user according to this position and the gesture type; An instruction mapping and execution unit for mapping the gesture type and user priority to the control instructions of the LED screen to achieve real - time response of the LED screen to user gestures.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the non - contact interaction method of the LED screen according to any one of claims 1 to 7.
10. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the steps of the non - contact interaction method of the LED screen according to any one of claims 1 to 7.