Remote control method for intelligent audio equipment and audio
By building a gesture recognition model on the audio equipment and utilizing the motion sensors and neural network technology of external smart devices, the problem of contactless and precise control of the audio equipment remote control is solved, and the interactive experience and response speed of the device are improved.
Patent Information
- Application Number
- CN202510863606.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-26
AI Technical Summary
Existing remote control methods for audio equipment make it difficult to achieve contactless, efficient and precise gesture control, and environmental noise is sensitive to interference with voice recognition, resulting in false triggering, invalid operations and delayed responses.
Gesture data is captured by the motion sensors of external smart devices. Combined with convolutional neural networks and long short-term memory network technologies, a direction deviation correction factor and an adaptive direction gain factor are introduced to build a gesture recognition model. A multi-layer filtering and verification mechanism is used to achieve accurate recognition of gestures and real-time feedback.
Significantly improve the recognition accuracy of gesture input in complex environments, optimize the interactive experience and response speed, and avoid false triggers and operation delays.
Smart Images

Figure CN120406745B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of device control remote control, and in particular to a remote control method and a speaker for an intelligent audio device. Background Art
[0002] A remote control method and speaker for smart audio devices are designed to enhance the user interaction experience and increase the accuracy of device control. By combining voice commands with gesture data captured by motion sensors of external smart devices, a multi-layer filtering and verification mechanism is used to parse and optimize gestures in real time, control functions such as volume adjustment, music playback, and Bluetooth connection, realize intelligent remote control without physical contact, and ensure the accurate recognition and effective execution of user gestures.
[0003] Existing remote control methods for audio equipment usually make it difficult to achieve contactless, efficient and precise gesture control. In addition, since traditional remote controls rely on a single physical input method, environmental noise is sensitive to interference with voice recognition, and they are highly dependent on voice ambiguity, which can lead to false triggering, invalid operations and delayed response to user input. Therefore, a remote control method and audio for intelligent audio equipment are provided. Summary of the Invention
[0004] The object of the present invention is to provide a remote control method and a speaker for an intelligent audio device to solve the problems raised in the above-mentioned background art, namely, that traditional remote controls rely on a single physical input method, are sensitive to interference from ambient noise on voice recognition, and are highly dependent on voice ambiguity, resulting in false triggering, invalid operations, and delayed response to user input.
[0005] To achieve the above object, the present invention provides a remote control method for an intelligent audio device, comprising:
[0006] capturing movement data information of the external smart device through a built-in motion sensor of the external smart device, wherein the movement data information includes duration, movement path, speed, and direction of the external smart device;
[0007] Based on convolutional neural network and long short-term memory network technology, and introducing direction deviation correction factor and adaptive direction gain factor, a gesture recognition model is constructed;
[0008] Analyzing gestures in the movement data information in real time using the gesture recognition model;
[0009] Using a multi-layer filtering and verification mechanism to filter the gestures parsed by the gesture recognition model to obtain valid gestures;
[0010] Transmitting the valid gesture to the smart audio device; and
[0011] The smart audio device converts the received valid gesture into a control instruction and executes the control instruction.
[0012] As a further improvement of this technical solution, the construction of the gesture recognition model includes:
[0013] Inputting the mobile data information into the convolution layer to obtain a feature map; and
[0014] The characteristic map is corrected according to the direction deviation correction factor to obtain a direction deviation corrected characteristic map, wherein the direction deviation correction factor is used to reduce the weight of invalid gestures that are inconsistent with the preset target gesture.
[0015] As a further improvement of this technical solution, the construction of the gesture recognition model further includes:
[0016] The directional deviation correction feature map is input into the LSTM layer of the long short-term memory network, and the adaptive directional gain factor is introduced to construct the gesture recognition model. The adaptive directional gain factor is used to dynamically adjust the gain according to the directional deviation and speed in the movement data information.
[0017] As a further improvement of this technical solution, the multi-layer filtering and verification mechanism includes:
[0018] Duration filtering, used to filter gestures based on a preset duration range;
[0019] Speed filtering, used to filter gestures based on a preset speed range;
[0020] Direction verification, used to verify the direction of the gesture based on the preset direction angle and allowed deviation; and
[0021] Comprehensive verification is used to combine the results of the duration filtering, speed filtering and direction verification to determine the valid gesture.
[0022] As a further improvement of this technical solution, the gesture mapping analyzed by the gesture recognition model is a functional operation of the smart audio device, and the functional operations include but are not limited to: playing music, stopping music, turning up the volume, turning on Bluetooth or turning off the device.
[0023] As a further improvement of the technical solution, the method further includes:
[0024] A visual interface is constructed on the smart audio device, and the visual interface displays the current status of the smart audio device in real time, and the current status includes playback status, track information, volume, and device connection status.
[0025] As a further improvement of the technical solution, the method further comprises the following steps:
[0026] Activating the smart audio device through a preset wake-up word; and
[0027] Through the automatic speech recognition engine and natural language processing technology, the voice instructions issued by the user are converted into operation commands to control the functional operations of the smart audio device.
[0028] As a further improvement of the present technical solution, the voice command includes "play music", "stop music", "increase volume", "turn on Bluetooth" or "turn off device".
[0029] As a further improvement of the technical solution, before the method described above, the following steps are further included:
[0030] The external smart device is quickly paired with the smart audio device using NFC technology. The pairing process includes exchanging information required for Bluetooth pairing through NFC and completing Bluetooth pairing through the SSP handshake protocol.
[0031] In another aspect, the present invention provides an intelligent audio device, comprising:
[0032] processor;
[0033] a memory for storing program instructions;
[0034] Communication module, used for data communication with external smart devices;
[0035] An NFC module, configured to implement near field communication pairing with the external smart device; and
[0036] A display module, used to display the current status of the smart audio device;
[0037] The processor, by executing program instructions in the memory, is configured to:
[0038] Receiving valid gestures in mobile data information transmitted by an external smart device;
[0039] Converting the valid gesture into a control instruction and executing the control instruction;
[0040] And, the smart audio device is controlled according to the received voice command or the pairing completed through the NFC module.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] 1. This remote control method and speaker for smart audio devices are based on convolutional neural networks and long short-term memory network technologies, and combined with a direction deviation correction factor and an adaptive direction gain factor to achieve accurate recognition and real-time feedback of user gestures. Especially in complex environments and dynamically changing conditions, it can significantly improve the recognition accuracy of gesture input and optimize the interactive experience and response speed of smart audio devices.
[0043] 2. This remote control method and speaker for smart audio devices dynamically analyzes gesture data through a multi-layer filtering and verification mechanism to achieve accurate gesture recognition and invalid gesture elimination, thereby ensuring that only valid gestures can be received by the audio device and converted into control commands, effectively avoiding false triggering and operation delays. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 The figure is a flow chart of the overall method of the present invention. DETAILED DESCRIPTION
[0045] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0046] Example 1: Please refer to Figure 1 As shown, this embodiment provides a remote control method for a smart audio device, including the following steps:
[0047] S1. Activate the smart speaker device through the preset wake-up word, and the user controls the volume, playback, Bluetooth, and power status of the smart speaker device through voice commands.
[0048] In this embodiment S1, the smart audio device is activated by a preset wake-up word, and the user controls the volume, playback, Bluetooth, and power status of the smart audio device through voice commands. The specific method is as follows:
[0049] S1.1. Build voice control technology into smart speakers and set wake-up words and voice commands.
[0050] Among them, wake-up words include "Hey, speaker"; voice commands include "play music", "stop music", "increase volume", "turn on Bluetooth", "turn off device" and more custom voice commands.
[0051] S1.2. Use an automatic speech recognition engine to convert the captured voice signal into a text signal, and then use natural language processing technology to parse the text signal into the wake-up word and voice command issued by the user.
[0052] S1.3. Analyze the wake-up words and voice commands issued by the user and map them to the operation commands of the smart audio device.
[0053] S1.4. According to the mapped instructions, the smart audio device performs corresponding functional operations.
[0054] In this embodiment, the automatic speech recognition engine is a system based on speech signal processing and deep learning technology, which is used to convert speech signals into text information that can be understood and processed by machines; the function of the automatic speech recognition engine is to convert the speech signal input by the user through the microphone into text information; it first recognizes the wake-up word in the speech signal to activate the device, and then the automatic speech recognition engine recognizes subsequent voice commands and converts these voice commands into text form that can be understood and executed by the machine.
[0055] Natural language processing technology is a combination of computer science and linguistics, used to understand, analyze, generate and process human language. In speech recognition, the role of natural language processing technology is to semantically parse the text signals output by the automatic speech recognition engine, and through technologies such as intent recognition, lexical analysis, and context understanding, convert the text form of these voice instructions into commands that can be executed by the device.
[0056] In this embodiment, the method for mapping the wake-up word and voice command to the operation command of the smart audio device is as follows:
[0057] "Play Music" is mapped to: start playing the music in the current playlist;
[0058] "Stop music" is mapped to: pause the currently playing music;
[0059] "Turn up the volume" is mapped to: increase the volume of the audio device;
[0060] "Turn on Bluetooth" is mapped to: enable the Bluetooth function of the audio device;
[0061] "Turn off device" is mapped to: Turn off the power of the smart speaker device.
[0062] S2. Use NFC technology to quickly pair external smart devices with smart audio devices.
[0063] In this embodiment S2, NFC technology is used to quickly pair an external smart device with a smart audio device. The specific method is as follows:
[0064] S2.1. Integrate an NFC chip into a smart audio device.
[0065] S2.2. The user brings the external smart device close to the NFC area of the smart speaker to trigger the pairing process.
[0066] S2.3. The NFC chip in the smart audio device detects a nearby external smart device, and the smart audio device automatically enters pairing mode, exchanges the required Bluetooth pairing information through NFC, and completes the Bluetooth pairing through the SSP handshake protocol.
[0067] S2.4. After pairing is successful, the smart speaker device confirms the pairing success to the user through a voice prompt.
[0068] In this embodiment, during the NFC pairing process, all transmitted data is encrypted using AES-256, and the smart audio device and the external smart device authenticate each other's identities; only devices authorized by the user are allowed to pair, preventing unauthorized devices from accessing the smart audio system, and providing a limit on the number of pairings and a pairing validity period to prevent repeated or expired pairing requests.
[0069] S3. Capture the mobile data information of the external smart device through the built-in motion sensor of the external smart device, build a gesture recognition model based on convolutional neural network and long short-term memory network technology, and introduce direction deviation correction factor and adaptive direction gain factor, and use a multi-layer filtering and verification mechanism to filter out invalid gestures, and transmit valid gestures in the mobile data information to the smart audio device.
[0070] The movement data information of the external smart device is obtained by the user operating the external smart device, and the movement data information includes the duration, movement path, speed and direction of the device.
[0071] The gesture recognition model is based on convolutional neural network and long short-term memory network technologies, and introduces fusion training of direction deviation correction factor and adaptive direction gain factor. It is used to interpret the mobile data information generated by users operating external smart devices into gestures in real time.
[0072] The multi-layer filtering and verification mechanism is implemented based on time series analysis and threshold detection technology, which is used to comprehensively analyze the duration, movement path, speed and direction of gestures, and filter out invalid and falsely triggered gestures.
[0073] In this embodiment S3, the motion data information of the external smart device is captured by the built-in motion sensor of the external smart device. Based on the convolutional neural network and long short-term memory network technology, a direction deviation correction factor and an adaptive direction gain factor are introduced to construct a gesture recognition model. The specific steps of the method are as follows:
[0074] S3.1. Use the built-in motion sensor of the external smart device to capture the movement data information of the external smart device. The movement data information includes the duration, movement path, speed, and direction of the device:
[0075] ;
[0076] ;
[0077] (Here The initial velocity is zero, or It is the velocity data directly provided by the sensor which already includes the effect of initial velocity. Otherwise );
[0078] ;
[0079] ;
[0080] in, is the duration of the entire action; for Velocity vector at time instant; For arrive The displacement vector at the moment; for Direction of the moment; is the start time; is the end time; is the acceleration; is the integral variable; For the moment; for The velocity component in the x-axis direction at the moment; for The velocity component in the y-axis direction at the moment; for The velocity component in the z-axis direction at the moment; for The moving path component in the x-axis direction at the moment; for The moving path component in the y-axis direction at the moment; for The component of the movement path in the z-axis direction at the moment; A unit vector that represents a preset reference direction (for example, an axis of the world coordinate system or the initial orientation of the device); For speed Model.
[0081] S3.2. Construct a gesture recognition model using convolutional neural network and long short-term memory network technology.
[0082] S3.3. Constructing the Gesture Data Matrix :
[0083] ;
[0084] in, is the number of time steps; For the deadline the duration of the moment; is the duration of the last time step; is the moving path component in the x-axis direction of the last time step; is the moving path component in the y-axis direction of the last time step; is the moving path component in the z-axis direction of the last time step; , , are the velocity components in the x, y, and z directions of the last time step, respectively.
[0085] S3.4. Based on convolutional neural network and long short-term memory network technology, and introducing direction deviation correction factor and adaptive direction gain factor, extract gesture data matrix The gesture recognition model is constructed based on the spatial features in the
[0086] S3.5. Use the gesture recognition model to analyze gestures in mobile data information in real time.
[0087] The specific steps of the method in Example S3.4 are as follows:
[0088] S3.4.1. Input the gesture data matrix into convolutional layer 1:
[0089] ;
[0090] in, is the first feature map of the convolutional layer; is the convolution kernel weight matrix of the convolution layer; is the bias vector of the convolutional layer.
[0091] S3.4.2. Introducing the directional deviation correction factor to recalculate the convolutional layer's feature map , we get the directional deviation correction convolution layer feature map:
[0092] ;
[0093] in, Correct the convolution layer feature map for directional deviation; for moment direction deviation; is the direction deviation weight; is the directional deviation correction multiplier, and is the direction deviation correction index term;
[0094] in, Time direction deviation The calculation is as follows:
[0095] ;
[0096] in, The target gesture direction is predefined.
[0097] S3.4.3. Correct the direction deviation of the convolution layer 1 feature map Input pooling layer:
[0098] ;
[0099] in, is the feature map after downsampling.
[0100] S3.4.4. The downsampled feature map Input to convolutional layer 2 to obtain the convolutional layer 2 feature map:
[0101] ;
[0102] in, is the convolutional layer 2 feature map; is the convolution kernel weight matrix of convolution layer 2; is the bias vector of convolutional layer 2.
[0103] S3.4.5. Input the convolutional layer 2 feature map into the LSTM layer of the long short-term memory network, introduce the adaptive direction gain factor, and build a gesture recognition model:
[0104] ;
[0105] in, for Always hide the status; for Cell status at each moment; for Always hide the status; is the adaptive directional gain factor (in this embodiment, or The values of are normalized or clipped to enhance the stability of model training).
[0106] Among them, the adaptive direction gain factor The specific calculation is as follows:
[0107] ;
[0108] in, is the adjustment coefficient; is the maximum speed obtained from the statistics of the training data set.
[0109] Gesture recognition model:
[0110] ;
[0111] in, is the probability distribution of gesture categories; is the classification weight matrix; is the classification bias vector.
[0112] In this embodiment S3.5, the gesture recognition model is used to analyze the gestures in the mobile data information in real time, as follows:
[0113] ;
[0114] in, is the total number of gestures.
[0115] In this embodiment, The gesture type that is recognized is as follows:
[0116] for “Play Music”; for “Stop the Music”; for "Turn up the volume"; to “Turn on Bluetooth”; "Turn off the device", etc.; users can customize the type of gestures and the total number of gestures .
[0117] In this embodiment S3, a multi-layer filtering and verification mechanism is used to filter invalid gestures and transmit valid gestures in the mobile data information to the smart audio device, as follows:
[0118] S3.6, Multi-layer filtering and verification mechanism:
[0119] S3.6.1, Duration filtering:
[0120] ;
[0121] in, is the lower limit of duration; The duration limit.
[0122] S3.6.2, Speed Filtering:
[0123] ;
[0124] in, is the lower speed limit; is the upper speed limit; From the speed sequence The feature speed values extracted from the gesture are, for example, the average speed, peak speed, or the speed at the key moments of the gesture.
[0125] S3.6.3 Direction Verification:
[0126]
[0127] in, is the preset direction angle; To allow the direction angle deviation, From the direction sequence The feature orientation angle extracted from .
[0128] S3.6.4 Comprehensive Verification:
[0129] ;
[0130] ;
[0131] in, Invalid gesture; For effective gestures; For the logical AND operation.
[0132] S3.7. Filter invalid gestures Afterwards, the remaining valid gestures are transmitted to the smart speaker device.
[0133] S4. The smart audio device converts the received valid gesture into a control command and executes the gesture operation.
[0134] In this embodiment S4, the smart audio device converts the received valid gesture into a control instruction and performs the gesture operation as follows:
[0135] S4.1. Set the control command mapping function according to the valid gesture mapping relationship:
[0136] Valid gesture mapping relationship:
[0137] for “Play Music”; for “Stop the Music”; for "Turn up the volume"; to “Turn on Bluetooth”; to “turn off the device”;
[0138] Control instruction mapping function:
[0139] ;
[0140] ;
[0141] in, For the Control instructions for gestures; Mapping function for control instructions.
[0142] S4.2. After receiving the control command, the smart audio device performs the corresponding operation.
[0143] S5. Build a visual interface on the smart audio device to display the current status of the smart audio device in real time.
[0144] In this embodiment S5, a visual interface is constructed on the smart audio device to display the current status of the smart audio device in real time, as follows:
[0145] S5.1. Design visual interface components, including a playback status indicator, track name and progress bar, control buttons, and device status information;
[0146] S5.2. Obtain the status data of the audio system in real time through the data transmission protocol, and use JavaScript technology to refresh the interface data in real time.
[0147] Embodiment 2: This embodiment provides a speaker, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any one of the above-described remote control methods for an intelligent speaker device.
[0148] The basic principles, main features, and advantages of the present invention are shown and described above. It should be understood by those skilled in the art that the present invention is not limited to the above-described embodiments. The above-described embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications fall within the scope of the invention claimed.
Claims
1. A remote control method for an intelligent audio device, characterized in that: The following steps are involved: capturing movement data information of the external smart device through a built-in motion sensor of the external smart device, wherein the movement data information includes duration, movement path, speed, and direction of the external smart device; Based on convolutional neural network and long short-term memory network technology, and introducing direction deviation correction factor and adaptive direction gain factor, a gesture recognition model is constructed. The gesture recognition model is ; in, is the probability distribution of gesture categories; is the classification weight matrix; is the classification bias vector, for Always hide the status; Among them, the adaptive directional gain factor The specific calculation is as follows: ; in, is the adjustment coefficient; is the maximum speed obtained from the statistics of the training data set, for Time direction deviation, for Velocity vector at time instant; Analyzing gestures in the movement data information in real time using the gesture recognition model; Using a multi-layer filtering and verification mechanism to filter the gestures parsed by the gesture recognition model to obtain valid gestures; Transmitting the valid gesture to the smart audio device; and The smart audio device converts the received valid gesture into a control instruction and executes the control instruction; The constructing of the gesture recognition model includes: Inputting the mobile data information into the convolution layer to obtain a feature map; and Correcting the characteristic map according to the direction deviation correction factor to obtain a direction deviation corrected characteristic map, wherein the direction deviation correction factor is used to reduce the weight of invalid gestures that are inconsistent with the preset target gesture; The constructing of the gesture recognition model further comprises: Inputting the directional deviation correction feature map into the LSTM layer of a long short-term memory network and introducing the adaptive directional gain factor to construct the gesture recognition model, wherein the adaptive directional gain factor is used to dynamically adjust the gain according to the directional deviation and speed in the movement data information; The multi-layer filtering and verification mechanism includes: Duration filtering, used to filter gestures based on a preset duration range; Speed filtering, used to filter gestures based on a preset speed range; Direction verification, used to verify the direction of the gesture based on the preset direction angle and allowed deviation; and Comprehensive verification is used to combine the results of the duration filtering, speed filtering and direction verification to determine the valid gesture.
2. The method according to claim 1, wherein: The gesture mapping analyzed by the gesture recognition model is a functional operation of the smart audio device, and the functional operations include but are not limited to: playing music, stopping music, increasing the volume, turning on Bluetooth, or turning off the device.
3. The method according to claim 1, wherein: The method then further comprises: A visual interface is constructed on the smart audio device, and the visual interface displays the current status of the smart audio device in real time, and the current status includes playback status, track information, volume, and device connection status.
4. The method according to claim 1, wherein: The method further comprises the following steps: Activating the smart audio device through a preset wake-up word; and Through the automatic speech recognition engine and natural language processing technology, the voice instructions issued by the user are converted into operation commands to control the functional operations of the smart audio device.
5. The method according to claim 4, characterized in that: The voice commands include "play music," "stop music," "turn up the volume," "turn on Bluetooth," or "turn off the device." 6. The method according to claim 1, wherein: Prior to the method described above, the following steps are also included: The external smart device is quickly paired with the smart audio device using NFC technology. The pairing process includes exchanging information required for Bluetooth pairing through NFC and completing Bluetooth pairing through the SSP handshake protocol.
7. An intelligent audio device controlled by a remote control method for an intelligent audio device according to any one of claims 1 to 6, characterized in that: include: processor; a memory for storing program instructions; Communication module, used for data communication with external smart devices; NFC module, used to achieve near field communication pairing with the external smart device; as well as A display module, used to display the current status of the smart audio device; The processor, by executing program instructions in the memory, is configured to: Receiving valid gestures in mobile data information transmitted by an external smart device; Converting the valid gesture into a control instruction and executing the control instruction; And, the smart audio device is controlled according to the received voice command or the pairing completed through the NFC module.
Citation Information
Patent Citations
System capable of interacting with intelligent acoustics
CN108495212A
A method for weld seam defect recognition based on an improved convolution neural network
CN109034204A
Gesture recognition method and device, storage medium and data glove
CN112347951A
Upper limb rehabilitation exercise prediction method and system
CN117462117A