Remote control method for intelligent sound equipment and sound equipment
The gesture recognition model is built through the motion sensor and neural network technology of external intelligent devices, and combined with multi-layer filtering and verification mechanisms, the contactless and precise control problem of remote control of audio equipment is solved, and efficient gesture recognition and response in complex environments is achieved.
Patent Information
- Application Number
- CN202510863606.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-26
AI Technical Summary
The existing remote control methods of audio equipment are difficult to achieve contactless, efficient and precise gesture control, and the ambient noise is sensitive to speech recognition interference, resulting in false triggering, invalid operation and delayed response.
The motion sensor of external intelligent devices is used to capture gesture data, combine convolutional neural networks and long-term memory network technology, and introduce direction deviation correction factors and adaptive direction gain factors to build a gesture recognition model, and accurately identify and verify through multi-layer filtering and verification mechanisms, and realize contactless control with voice commands.
Significantly improve gesture recognition accuracy in complex environments, optimize interaction experience and response speed, avoid mistriggering and operation delays, and ensure accurate contactless control.
Smart Images

Figure CN120406745A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of device control remote control, and in particular to a remote control method and a speaker for an intelligent audio device. Background Art
[0002] A remote control method and speaker for smart audio devices are designed to enhance the user interaction experience and increase the accuracy of device control. By combining voice commands with gesture data captured by motion sensors of external smart devices, a multi-layer filtering and verification mechanism is used to parse and optimize gestures in real time, control functions such as volume adjustment, music playback, and Bluetooth connection, realize intelligent remote control without physical contact, and ensure the accurate recognition and effective execution of user gestures.
[0003] Existing remote control methods for audio equipment usually make it difficult to achieve contactless, efficient and precise gesture control. In addition, since traditional remote controls rely on a single physical input method, environmental noise is sensitive to interference with voice recognition, and they are highly dependent on voice ambiguity, which can lead to false triggering, invalid operations and delayed response to user input. Therefore, a remote control method and audio for intelligent audio equipment are provided. Summary of the Invention
[0004] The object of the present invention is to provide a remote control method and a speaker for an intelligent audio device to solve the problems raised in the above-mentioned background art, namely, that traditional remote controls rely on a single physical input method, are sensitive to interference from ambient noise on voice recognition, and are highly dependent on voice ambiguity, resulting in false triggering, invalid operations, and delayed response to user input.
[0005] To achieve the above object, the present invention provides a remote control method for an intelligent audio device, comprising: capturing movement data information of the external smart device through a built-in motion sensor of the external smart device, wherein the movement data information includes duration, movement path, speed, and direction of the external smart device; Based on convolutional neural network and long short-term memory network technology, and introducing direction deviation correction factor and adaptive direction gain factor, a gesture recognition model is constructed; Analyzing gestures in the movement data information in real time using the gesture recognition model; Using a multi-layer filtering and verification mechanism to filter the gestures parsed by the gesture recognition model to obtain valid gestures; Transmitting the valid gesture to the smart audio device; and The smart audio device converts the received valid gesture into a control instruction and executes the control instruction.
[0006] As a further improvement of the present technical solution, the construction of the gesture recognition model includes: Inputting the mobile data information into a convolutional layer to obtain a feature map; and Correcting the feature map according to the direction deviation correction factor to obtain a direction deviation corrected feature map, where the direction deviation correction factor is used to reduce the weight of invalid gestures that are inconsistent with a preset target gesture.
[0007] As a further improvement of the present technical solution, the construction of the gesture recognition model further includes: Inputting the direction deviation corrected feature map into the LSTM layer of a long short-term memory network and introducing the adaptive direction gain factor to construct the gesture recognition model, where the adaptive direction gain factor is used to dynamically adjust the gain according to the direction deviation and speed in the mobile data information.
[0008] As a further improvement of the present technical solution, the multi-layer filtering and verification mechanism includes: Duration filtering, which is used to filter gestures according to a preset duration range; Speed filtering, which is used to filter gestures according to a preset speed range; Direction verification, which is used to verify the gesture direction according to a preset direction angle and allowable deviation degree; and Comprehensive verification, which is used to combine the results of the duration filtering, speed filtering, and direction verification to determine the valid gesture.
[0009] As a further improvement of the present technical solution, the gesture mapped by the gesture recognition model is a function operation of a smart audio device, and the function operation includes but is not limited to: playing music, stopping music, increasing the volume, turning on Bluetooth, or turning off the device.
[0010] As a further improvement of the present technical solution, after the method further includes: Constructing a visual interface on the smart audio device, where the visual interface real-time displays the current state of the smart audio device, and the current state includes a playing state, track information, volume, and device connection state.
[0011] As a further improvement of the present technical solution, the method further includes the following steps: Activating the smart audio device through a preset wake-up word; and Converting a voice command issued by a user into an operation command through an automatic speech recognition engine and natural language processing technology to control the function operation of the smart audio device.
[0012] As a further improvement of the technical solution, the voice commands include "play music", "stop music", "increase volume", "turn on Bluetooth" or "turn off the device".
[0013] As a further improvement of the technical solution, before the method, the following steps are further included: Using NFC technology to quickly pair the external intelligent device with the intelligent audio device, and the pairing process includes exchanging the information required for Bluetooth pairing through NFC and completing Bluetooth pairing through the SSP handshake protocol.
[0014] On the other hand, the present invention provides an intelligent audio device, including: A processor; A memory for storing program instructions; A communication module for performing data communication with an external intelligent device; An NFC module for realizing near-field communication pairing with the external intelligent device; and A display module for displaying the current state of the intelligent audio device; The processor, by executing the program instructions in the memory, is configured to: Receive valid gestures in the mobile data information transmitted by an external intelligent device; Convert the valid gestures into control instructions and execute the control instructions; And, control the intelligent audio device according to the received voice commands or the pairing completed through the NFC module.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. In the remote control method and audio for the intelligent audio device, based on the convolutional neural network and long short-term memory network technologies, combined with the direction deviation correction factor and the adaptive direction gain factor, accurate recognition and real-time feedback of user gestures are realized. Especially in complex environments and dynamically changing conditions, the recognition accuracy of gesture input can be significantly improved, and the interaction experience and response speed of the intelligent audio device can be optimized.
[0016] 2. In the remote control method and audio for the intelligent audio device, through a multi-layer filtering and verification mechanism, dynamic analysis of gesture data is performed to realize accurate gesture recognition and elimination of invalid gestures, so as to ensure that only valid gestures can be received by the audio device and converted into control instructions, effectively avoiding mis-triggering and operation delay problems. Description of the Drawings
[0017] Figure 1 It is the overall method flow chart of the present invention. Detailed Embodiments
[0018] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0019] Embodiment 1: Please refer to Figure 1 As shown, this embodiment provides a remote control method for a smart audio device, including the following steps: S1. Activate the smart audio device through a preset wake-up word, and the user controls the volume, playback, Bluetooth, and on / off status of the smart audio device through voice commands.
[0020] In step S1 of this embodiment, the smart audio device is activated through a preset wake-up word, and the user controls the volume, playback, Bluetooth, and on / off status of the smart audio device through voice commands. The specific method is as follows: S1.1. Incorporate voice control technology into the smart audio device, and set the wake-up word and voice commands. Among them, the wake-up word includes "Hey, speaker"; the voice commands include "Play music", "Stop music", "Increase volume", "Turn on Bluetooth", "Turn off the device", and more custom voice commands.
[0021] S1.2. Adopt an automatic speech recognition engine to convert the captured voice signal into a text signal, and then use natural language processing technology to parse the wake-up word and voice commands issued by the user from the text signal.
[0022] S1.3. Map the wake-up word and voice commands issued by the user to the operation commands of the smart audio device.
[0023] S1.4. According to the mapped instructions, the smart audio device executes the corresponding functional operations.
[0024] In this embodiment, the automatic speech recognition engine is a system based on speech signal processing and deep learning technology, used to convert voice signals into text information that can be understood and processed by machines; the role of the automatic speech recognition engine is to convert the voice signals input by the user through the microphone into text information; it first recognizes the wake-up word in the voice signal to activate the device, and then the automatic speech recognition engine recognizes the subsequent voice commands and converts these voice commands into a text form that can be understood and executed by machines.
[0025] Natural language processing technology is a technology that combines computer science and linguistics and is used to understand, analyze, generate, and process human language. In speech recognition, the role of natural language processing technology is to semantically parse the text signal output by the automatic speech recognition engine and convert the text form of these speech instructions into commands that the device can execute through technologies such as intent recognition, lexical analysis, and context understanding.
[0026] In this embodiment, the mapping method for mapping wake-up words and voice instructions to the operation commands of the smart audio device is as follows: "Play music" is mapped to: start playing the music in the current playlist; "Stop music" is mapped to: pause the currently playing music; "Turn up the volume" is mapped to: increase the volume of the audio device; "Turn on Bluetooth" is mapped to: enable the Bluetooth function of the audio device; "Turn off the device" is mapped to: turn off the power of the smart audio device.
[0027] S2. Use NFC technology to quickly pair an external smart device to the smart audio device.
[0028] In step S2 of this embodiment, the method for quickly pairing an external smart device to the smart audio device using NFC technology is as follows: S2.1. Integrate an NFC chip in the smart audio device.
[0029] S2.2. The user brings the external smart device close to the NFC area of the smart audio device to trigger the pairing process.
[0030] S2.3. The NFC chip in the smart audio device detects the nearby external smart device, and the smart audio device automatically enters the pairing mode, exchanges the information required for Bluetooth pairing through NFC, and completes the Bluetooth pairing through the SSP handshake protocol.
[0031] S2.4. After successful pairing, the smart audio device confirms the successful pairing to the user through voice prompts.
[0032] In this embodiment, during the NFC pairing process, all transmitted data is encrypted using AES-256, and the smart audio device and the external smart device authenticate each other's identities; only devices authorized by the user are allowed to pair, preventing unauthorized devices from accessing the smart audio system, and providing pairing times limit and pairing validity period settings to prevent repeated or expired pairing requests.
[0033] S3. Capture the movement data information of the external intelligent device through the motion sensor built in the external intelligent device. Based on the convolutional neural network and long short-term memory network technologies, introduce a direction deviation correction factor and an adaptive direction gain factor to construct a gesture recognition model, and use a multi-layer filtering and verification mechanism to filter out invalid gestures, and transmit the valid gestures in the movement data information to the intelligent audio device.
[0034] Among them, the movement data information of the external intelligent device is obtained by the user operating the external intelligent device, and the movement data information includes the duration, movement path, speed, and direction of the device.
[0035] The gesture recognition model is trained by fusing based on the convolutional neural network and long short-term memory network technologies, and introducing a direction deviation correction factor and an adaptive direction gain factor, and is used to parse the movement data information generated by the user operating the external intelligent device into gestures in real time.
[0036] The multi-layer filtering and verification mechanism is implemented based on time series analysis and threshold detection technologies, and is used to comprehensively analyze the duration, movement path, speed, and direction of the gesture, and filter out invalid and mis-triggered gestures.
[0037] In this embodiment S3, capture the movement data information of the external intelligent device through the motion sensor built in the external intelligent device. Based on the convolutional neural network and long short-term memory network technologies, introduce a direction deviation correction factor and an adaptive direction gain factor to construct a gesture recognition model. The specific method steps are as follows: S3.1. Capture the movement data information of the external intelligent device through the motion sensor built in the external intelligent device. The movement data information includes the duration, movement path, speed, and direction of the device: ; ; (Here the initial velocity at the moment is zero, or it is the velocity data that already includes the influence of the initial velocity directly provided by the sensor, otherwise ); ; ; Among them, is the duration of the entire action; is the velocity vector at the moment; is from to the displacement vector at the moment; is the direction at the moment; is the start time; is the end time; is the acceleration; is the integration variable; is the moment; is the velocity component in the x-axis direction at the moment; is the velocity component in the y-axis direction at the moment; is the velocity component in the z-axis direction at the moment; is the moving path component in the x-axis direction at the moment; is the moving path component in the y-axis direction at the moment; is the moving path component in the z-axis direction at the moment; the unit vector of the preset reference direction (e.g., a certain axis of the world coordinate system or the initial orientation of the device); is the velocity modulus.
[0038] S3.2. Construct a gesture recognition model through convolutional neural network and long short-term memory network technologies.
[0039] S3.3. Construct a gesture data matrix : ; wherein, is the number of time steps; is up to the duration at the moment; is the duration of the last time step; is the moving path component in the x-axis direction of the last time step; is the moving path component in the y-axis direction of the last time step; is the moving path component in the z-axis direction of the last time step; , , are respectively the velocity components in the x, y, and z-axis directions of the last time step.
[0040] S3.4. Based on convolutional neural network and long short-term memory network technologies, and introducing a direction deviation correction factor and an adaptive direction gain factor, extract the spatial features in the gesture data matrix to construct a gesture recognition model.
[0041] S3.5. Use the gesture recognition model to parse the gestures in the mobile data information in real time.
[0042] The specific method steps of this embodiment S3.4 are as follows: S3.4.1. Input the gesture data matrix into the first convolutional layer: ; Among them, is the feature map of the first convolutional layer; is the convolutional kernel weight matrix of the first convolutional layer; is the bias vector of the first convolutional layer.
[0043] S3.4.2. Introduce a direction deviation correction factor to recalculate the feature map of the first convolutional layer , and obtain the direction deviation corrected feature map of the first convolutional layer: ; Among them, is the direction deviation corrected feature map of the first convolutional layer; is the direction deviation at time is the direction deviation weight; is the direction deviation correction multiplier, and is the direction deviation correction exponential term; Among them, the direction deviation at time is calculated as follows: ; Among them, is the predefined target gesture direction.
[0044] S3.4.3. Input the direction deviation corrected feature map of the first convolutional layer into the pooling layer: ; Among them, is the downsampled feature map.
[0045] S3.4.4. Input the downsampled feature map into the second convolutional layer to obtain the feature map of the second convolutional layer: ; Among them, is the feature map of the second convolutional layer; is the convolutional kernel weight matrix of the second convolutional layer; is the bias vector of the second convolutional layer.
[0046] S3.4.5. Input the feature map of the second convolutional layer into the LSTM layer of the long short-term memory network, introduce an adaptive direction gain factor, and construct a gesture recognition model: ; Among them, is the hidden state at a moment; is the cell state at a moment; is the hidden state at a moment; is the adaptive direction gain factor (in this embodiment, the or value is normalized or clipped to enhance the stability of model training).
[0047] Among them, the adaptive direction gain factor is specifically calculated as follows: ; Among them, is the adjustment coefficient; is the maximum speed statistically obtained from the training data set.
[0048] Gesture recognition model: ; Among them, is the gesture category probability distribution; is the classification weight matrix; is the classification bias vector.
[0049] In this embodiment S3.5, the gesture recognition model is used to parse the gestures in the mobile data information in real time, specifically as follows: ; Among them, is the total number of gestures.
[0050] In this embodiment, is the recognized gesture type, specifically as follows: is "play music"; is "stop music"; is "turn up the volume"; is "turn on Bluetooth"; is "turn off the device", etc.; the user can customize the type of gesture and the total number of gestures .
[0051] In this embodiment S3, a multi-layer filtering and verification mechanism is used to filter invalid gestures, and the valid gestures in the mobile data information are transmitted to the smart speaker device, specifically as follows: S3.6. Multi-layer filtering and verification mechanism: S3.6.1. Duration filtering: ; Among them, is the lower limit of the duration; is the upper limit of the duration.
[0052] S3.6.2, Speed filtering: ; wherein, is the lower limit of the speed; is the upper limit of the speed; is the characteristic speed value extracted from the speed sequence such as the average speed, peak speed or speed at the critical moment of the gesture.
[0053] S3.6.3, Direction verification: wherein, is the preset direction angle; is the allowable direction angle deviation, is the characteristic direction angle extracted from the direction sequence ;
[0054] S3.6.4, Comprehensive verification: ; ; wherein, is an invalid gesture; is a valid gesture; is a logical AND operation.
[0055] S3.7, Filter out invalid gestures After that, the remaining valid gestures are transmitted to the smart speaker device.
[0056] S4, The smart speaker device converts the received valid gestures into control instructions and performs gesture operations.
[0057] In this embodiment S4, the smart speaker device converts the received valid gestures into control instructions and performs gesture operations, specifically as follows: S4.1, Set the control instruction mapping function according to the valid gesture mapping relationship: Valid gesture mapping relationship: is "play music"; is "stop music"; is "increase volume"; is "turn on Bluetooth"; is "turn off the device"; Control instruction mapping function: ; ; Among them, is the control instruction for the th gesture; is the control instruction mapping function.
[0058] S4.2. After the intelligent sound device receives the control instruction, it performs the corresponding operation.
[0059] S5. Build a visual interface on the intelligent sound device to display the current state of the intelligent sound device in real time.
[0060] In this embodiment S5, building a visual interface on the intelligent sound device for displaying the current state of the intelligent sound device in real time is as follows: S5.1. Design the visual interface components, which include a playback status indicator, a track name and progress bar, control buttons, and device status information; S5.2. Obtain the status data of the sound in real time through a data transmission protocol, and use JavaScript technology to refresh the interface data in real time.
[0061] Embodiment 2: This embodiment provides a sound, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the remote control method for an intelligent sound device described in any one of the above.
[0062] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A remote control method for a smart audio device, characterized in that, Including the following steps: Capturing the movement data information of the external intelligent device through a motion sensor built into the external intelligent device, where the movement data information includes the duration, movement path, speed, and direction of the external intelligent device; Based on convolutional neural network and long short-term memory network technologies, and introducing a direction deviation correction factor and an adaptive direction gain factor, constructing a gesture recognition model; Using the gesture recognition model to parse the gestures in the movement data information in real time; Using a multi-layer filtering and verification mechanism to filter the gestures parsed by the gesture recognition model to obtain valid gestures; Transmitting the valid gestures to the intelligent audio device; And The intelligent audio device converts the received valid gestures into control instructions and executes the control instructions.
2. The method according to claim 1, wherein: The constructing of the gesture recognition model includes: Inputting the movement data information into a convolutional layer to obtain a feature map; and Correcting the feature map according to the direction deviation correction factor to obtain a direction deviation corrected feature map, where the direction deviation correction factor is used to reduce the weight of invalid gestures inconsistent with a preset target gesture.
3. The method according to claim 2, characterized in that: The constructing of the gesture recognition model further includes: Inputting the direction deviation corrected feature map into the LSTM layer of the long short-term memory network and introducing the adaptive direction gain factor to construct the gesture recognition model, where the adaptive direction gain factor is used to dynamically adjust the gain according to the direction deviation and speed in the movement data information.
4. The method according to claim 1, wherein: The multi-layer filtering and verification mechanism includes: Duration filtering, which is used to filter gestures according to a preset duration range; Speed filtering, which is used to filter gestures according to a preset speed range; Direction verification, which is used to verify the gesture direction according to a preset direction angle and allowable deviation degree; and Comprehensive verification, which is used to combine the results of the duration filtering, speed filtering, and direction verification to determine the valid gestures.
5. The method according to claim 1, wherein: The gestures parsed by the gesture recognition model are mapped to function operations of the intelligent audio device, and the function operations include but are not limited to: playing music, stopping music, turning up the volume, turning on Bluetooth, or turning off the device.
6. The method according to claim 1, characterized in that: After the described method, it further includes: Constructing a visualization interface on the intelligent audio device, and the visualization interface displays the current state of the intelligent audio device in real time, where the current state includes the playing state, track information, volume, and device connection state.
7. The method according to claim 1, characterized in that: The described method further includes the following steps: Activating the intelligent audio device through a preset wake-up word; and Converting the voice instructions issued by the user into operation commands through an automatic speech recognition engine and natural language processing technology to control the function operations of the intelligent audio device.
8. The method according to claim 7, wherein: The voice instructions include "play music", "stop music", "turn up the volume", "turn on Bluetooth", or "turn off the device".
9. The method according to claim 1, characterized in that: Before the described method, it further includes the following steps: Quickly pairing the external intelligent device to the intelligent audio device using NFC technology, and the pairing process includes exchanging Bluetooth pairing required information through NFC and completing Bluetooth pairing through the SSP handshake protocol.
10. An intelligent audio device, characterized in that, Including: A processor; A memory for storing program instructions; A communication module for data communication with an external intelligent device; An NFC module for implementing near-field communication pairing with the external intelligent device; And A display module for presenting the current state of the intelligent audio device; The processor, by executing program instructions in the memory, is configured to: Receive valid gestures in mobile data information transmitted by an external intelligent device; Convert the valid gestures into control instructions and execute the control instructions; And control the intelligent audio device according to received voice instructions or pairing completed through the NFC module.
Citation Information
Patent Citations
System capable of interacting with intelligent acoustics
CN108495212A
A method for weld seam defect recognition based on an improved convolution neural network
CN109034204A
An intelligent electronic equipment gesture capturing and identifying technology based on touch hand detection
CN109947243A
Dynamic gesture recognition method based on adaptive space supervision
CN111273779A
Gesture recognition method and device, storage medium and data glove
CN112347951A