Cosmetic instrument voice interaction control method and system
Through the multimodal interaction technology of the beauty instrument, combined with voice and gesture sensing units, the problem of poor adaptability of voice interaction in complex environments is solved, high-precision control and personalized suggestions are achieved, and user experience and intelligence are improved.
Patent Information
- Application Number
- CN202510338767.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
The voice interaction technology of existing beauty instruments is poorly adaptable in complex environments, fails to make full use of gestures and voice combinations, and lacks personalized services, and fails to use historical usage data to provide personalized suggestions.
By detecting beauty instruments and loading speech recognition parameters, multimodal interaction between speech and gestures is achieved, personalized suggestions are generated based on historical use data, convolutional neural network is used to build a speech recognition model, processing environmental noise and identifying voice commands, and equipment control is carried out in combination with gesture sensing units.
Implement high-precision voice control in complex environments, adapt to noisy environments and busy scenes of users with both hands, provide emotionally driven personalized services, and improve user satisfaction and intelligence.
Smart Images

Figure CN120260558A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of beauty instrument control, and in particular to a voice interaction control method and system for beauty instruments. Background Art
[0002] In recent years, in the field of beauty instruments, voice interaction has gradually become an important means to improve user experience and device intelligence. Currently, there are already many beauty instrument products on the market that integrate voice interaction functions. For example, the device intensity can be adjusted through voice control and personalized suggestions can be provided. However, the existing technologies still have deficiencies in terms of the naturalness of interaction, environmental adaptability, and personalized services in practical applications.
[0003] Firstly, the adaptability of traditional voice interaction in complex environments is poor. For example, in a noisy environment, the accuracy of voice recognition will significantly decrease, affecting the user experience. Secondly, the development of existing technologies in multi-modal interaction lags behind, and the combination of gestures and voice interaction has not been fully utilized to meet the needs of users in different scenarios. Finally, there are deficiencies in personalized services in existing technologies, and the historical usage data of users has not been fully utilized to provide personalized suggestions. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a voice interaction control method for beauty instruments to solve the problem that the development of multi-modal interaction lags behind and the combination of gestures and voice interaction has not been fully utilized.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a voice interaction control method for beauty instruments, which includes detecting a beauty instrument and loading voice recognition parameters to complete the initialization of the voice interaction function; real-time collecting voice command data, performing preprocessing, constructing a voice recognition model, recognizing voice commands, outputting voice text commands, and feeding back the voice text commands to the user through a voice synthesis unit to execute the voice commands; activating a voice recognition unit through a voice command and activating a gesture sensing unit through a gesture action, and making adjustments when the user approaches the beauty instrument to achieve multi-modal interaction of voice and gestures; based on multi-modal interaction, generating user personalized suggestions using historical usage data and personalized settings, and feeding back the personalized suggestions to the user.
[0007] As a preferred solution of the voice interaction control method for beauty instruments according to the present invention, wherein: the steps of detecting the beauty instrument and loading voice recognition parameters to complete the initialization of the voice interaction function are specifically as follows, Detect the voice recognition unit, gesture sensing unit, voice synthesis unit, touch display screen and microphone in the beauty device through a detection program; Send a test signal to each beauty device by using the detection program; After receiving the test signal, the beauty device returns a response signal as the test result; Conduct a functional test on the voice recognition unit to verify the voice recognition and response speed, conduct a functional test on the gesture sensing unit to verify the response speed of gesture recognition, conduct a functional test on the voice synthesis unit to verify the voice output fluency, conduct a functional test on the touch display screen to verify the display clarity, and conduct a functional test on the microphone to verify the audio clarity; When the beauty device is started, read the voice recognition parameters from the non-volatile memory and load them into the voice recognition unit; Initialize the voice recognition unit, voice synthesis unit and gesture sensing unit by using an initialization program, and initialize the touch display screen by using a touch screen driver program; After the initialization is completed, the beauty device sends a start signal to the microphone and the voice recognition unit and enters the listening state.
[0008] As a preferred solution of the voice interaction control method for the beauty device described in the present invention, wherein: the voice recognition parameters include noise parameters, instruction parameters, interaction parameters and configuration parameters.
[0009] As a preferred solution of the voice interaction control method for the beauty device described in the present invention, wherein: the specific steps of collecting voice command data in real time, preprocessing it, constructing a voice recognition model, recognizing voice commands and outputting voice text commands are as follows, Use a microphone to collect voice command data and perform noise reduction processing on the voice command data through noise parameters; Standardize the noise-reduced voice command data through configuration parameters; Extract the voice feature vectors of the standardized voice command data by using Mel-frequency cepstral coefficients;
[0010] Divide the voice feature vectors into a training set and a validation set; Use a convolutional neural network CNN as the basic framework of the voice recognition model, and define the basic framework of the voice recognition model as an input layer, a hidden layer and an output layer; The input layer inputs the training set; The hidden layer uses an attention mechanism to perform weighted processing on the training set to obtain attention weights and pass them to the output layer; The output layer sets the output dimension of the Softmax activation function according to the number of command words in the instruction parameters, calculates the confidence probability of each dimension, and generates the corresponding voice text command; Evaluate the performance of the speech recognition model using the validation set and adjust the hyperparameters of the speech recognition model.
[0011] As a preferred solution of the beauty instrument voice interaction control method described in the present invention, wherein: the step of feeding back the voice text instruction to the user through the voice synthesis unit and executing the voice instruction is as follows. The voice synthesis unit receives the voice text instruction and converts the voice text content into a voice signal. The voice synthesis unit converts the voice signal into natural speech according to the speech rate, pitch and language preference in the personalized settings, and feeds it back to the user through the built-in speaker of the beauty instrument to execute the voice instruction.
[0012] As a preferred solution of the beauty instrument voice interaction control method described in the present invention, wherein: the step of activating the voice recognition unit through the voice instruction and activating the gesture sensing unit through the gesture action, and making adjustments when the user approaches the beauty instrument to achieve multimodal interaction of voice and gesture is as follows. The user activates the voice recognition unit through the voice instruction, and the voice recognition unit activates the voice interaction function of the beauty instrument by sending an electrical signal to the main controller of the beauty instrument. The user activates the trigger distance of the gesture sensing unit by defining the gesture action through the interaction parameter, and the gesture sensing unit recognizes the gesture instruction according to the gesture action and converts the gesture instruction into an action control signal of the beauty instrument to adjust the massage intensity and light therapy intensity of the beauty instrument. Set the priority rules for voice instructions and gesture instructions to determine the order of voice and gesture.
[0013] As a preferred solution of the beauty instrument voice interaction control method described in the present invention, wherein: the step of generating user personalized suggestions based on multimodal interaction, using historical usage data and personalized settings, and feeding back the personalized suggestions to the user is as follows. Use a decision tree to analyze the historical usage data to identify the user's behavior patterns and preferences. Generate user personalized suggestions according to the user's behavior patterns and preferences. Feed back the personalized suggestions to the user through the voice synthesis unit.
[0014] In a second aspect, the present invention provides a beauty instrument voice interaction control system, including: An initialization module for detecting the beauty instrument and loading the voice recognition parameters to complete the initialization of the voice interaction function. A voice recognition module for real-time collecting voice instruction data, performing preprocessing, building a voice recognition model, recognizing voice instructions, and outputting voice text instructions. A feedback module for feeding back the voice text instruction to the user through the voice synthesis unit to execute the voice instruction. An interaction module, configured to activate a speech recognition unit through a voice command and activate a gesture sensing unit through a gesture action. When the user approaches the beauty device, adjustments are made to achieve multi-modal interaction of voice and gesture. A personalization module, configured to generate user personalization suggestions based on multi-modal interaction, using historical usage data and personal settings, and provide the personalization suggestions to the user.
[0015] In a third aspect, the present invention provides a computer device, including a memory and a processor. The memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the beauty device voice interaction control method as described in the first aspect of the present invention is implemented.
[0016] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, wherein: when the computer program is executed by the processor, any step of the beauty device voice interaction control method as described in the first aspect of the present invention is implemented.
[0017] The beneficial effects of the present invention are as follows: The present invention completes initialization by detecting the beauty device and loading speech recognition parameters, achieving rapid startup and function verification of the beauty device, and providing a stable foundation for subsequent voice interaction; by constructing a speech recognition model, collecting speech signals in real time, processing environmental noise, and then performing voice command recognition, high-precision voice control in a complex environment is achieved, optimizing the accuracy and response speed of voice interaction; by combining a gesture sensing unit, multi-modal interaction of voice and gesture is achieved, adapting to noisy environments and scenarios where the user's hands are busy; by combining historical usage data to provide personalization suggestions, personalized service driven by emotion is achieved, further improving satisfaction and intelligence. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a flowchart of the beauty device voice interaction control method in Embodiment 1.
[0020] Figure 2 It is a detection flowchart in Embodiment 1.
[0021] Figure 3 It is a speech recognition flowchart in Embodiment 1.
[0022] Figure 4 It is a multi-modal interaction flowchart in Embodiment 1. Specific Embodiments
[0023] To make the above objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings of the specification.
[0024] In the following description, numerous specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0025] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments.
[0026] Embodiment 1, referring to Figures 1 to 4 , which is the first embodiment of the present invention. This embodiment provides a method for voice interaction control of a beauty device, including the following steps: S1. Detect the beauty device and load voice recognition parameters to complete the initialization of the voice interaction function.
[0027] Detect the voice recognition unit, gesture sensing unit, voice synthesis unit, touch display screen, and microphone in the beauty device through a detection program (such as python, and the specific detection program is replaced according to the architecture and development requirements of the beauty device), and verify the connection status and functional integrity of the voice recognition unit, gesture sensing unit, voice synthesis unit, touch display screen, and microphone through the beauty device detection program; Use the detection program to send test signals to each beauty device; The test signal refers to an electrical signal and a data packet, which specifically depends on the type of the beauty device; After the beauty device receives the test signal, it returns a response signal as the test result; Send a test voice signal to the voice recognition unit to verify the voice acquisition and processing ability, send a test instruction to the gesture sensing unit to verify the gesture recognition ability, send voice synthesis test data to the voice synthesis unit to verify the voice synthesis processing ability, send test data to the touch display screen to verify the display function, and send a test audio signal to the microphone to verify the audio acquisition ability; Receive the speech recognition results returned by the speech recognition unit, verify whether the response is correct, receive the gesture recognition results returned by the gesture sensing unit, verify whether the gestures can be recognized, receive the status code returned by the speech synthesis unit, verify whether speech synthesis is performed, receive the display status information returned by the touch display screen, verify whether normal display can be achieved, and receive the audio acquisition status information returned by the microphone, verify whether normal audio acquisition can be achieved; Conduct functional tests on the speech recognition unit to verify the speech recognition and response speed, conduct functional tests on the gesture sensing unit to verify the response speed of gesture recognition, conduct functional tests on the speech synthesis unit to verify the fluency of speech output, conduct functional tests on the touch display screen to verify the display clarity, and conduct functional tests on the microphone to verify the audio clarity; The non-volatile memory is used to store speech recognition parameters, user personalized settings, and historical usage data; Personalized settings refer to the user's speech preferences, gesture preferences, and device usage habits; Historical usage data refers to the frequency of user device usage and common usage patterns; When the beauty device is started, read the speech recognition parameters from the non-volatile memory and load them into the speech recognition unit; The speech recognition parameters include noise parameters, instruction parameters, interaction parameters, and configuration parameters; Use the initialization program to initialize the hardware and software resources of the speech recognition unit, speech synthesis unit, and gesture sensing unit of the beauty device, and use the touch screen driver program to initialize the touch display screen; The speech recognition unit is responsible for the recognition of speech instructions, the speech synthesis unit is used to output the feedback information of the device to the user in the form of speech, and the gesture sensing unit supports multimodal interaction; After successful initialization, the beauty device sends a start signal to the microphone and the speech recognition unit, enters the listening state, and monitors the status of each unit of the beauty device through an error detection mechanism. When speech instruction recognition fails, gesture recognition is incorrect, and the speech synthesis unit cannot output speech, repairs are made; The error detection mechanism refers to, during the speech interaction process, detecting the functional status of the speech recognition unit, gesture sensing unit, and speech synthesis unit by periodically sending test signals. When it is found that speech instruction recognition fails, gesture recognition is incorrect, and the speech synthesis unit cannot output speech, the user will be immediately prompted and an attempt will be made to automatically repair.
[0028] S2. Real-time collect speech instruction data, perform preprocessing, construct a speech recognition model, recognize speech instructions, output speech text instructions, and feedback the speech text instructions to the user through the speech synthesis unit, and execute the speech instructions.
[0029] Use a microphone to collect voice command data for turning on, turning off, and adjusting the RF intensity in different environments, and perform noise reduction processing on the voice command data through noise parameters; Standardize the noise-reduced voice command data through configuration parameters; Use the Short-Time Fourier Transform (STFT) to separate the voice and noise frequency bands, utilize the Wiener Filter to dynamically suppress noise, and simulate the acoustic path from the speaker to the microphone through the FIR filter to generate a reverse signal to cancel echoes; Adopt an anti-aliasing filter to unify the sampling rate and audio format, with a sampling rate of 16 kHz and an audio format of PCM; Extract the voice feature vectors of the standardized voice command data using Mel-Frequency Cepstral Coefficients; Divide the voice feature vectors into a training set and a validation set (with a ratio of 80%:20%); Utilize a Convolutional Neural Network (CNN) as the basic framework of the voice recognition model, and define the basic architecture of the voice recognition model as an input layer, a hidden layer, and an output layer; The input layer inputs the training set for training, and the training process is as follows: Input the training set data into the voice recognition model, calculate the voice text command, calculate the difference between the voice text command and the voice command data using the loss function, calculate the gradient of the loss function, and update the voice recognition model parameters through backpropagation. Use the Adam optimizer to update the voice recognition model weights according to the calculated gradient; The hidden layer uses the attention mechanism to perform weighted processing on the training set to obtain attention weights and pass them to the output layer; The hidden layer consists of multiple convolutional layers and pooling layers; The convolutional layer extracts the local features of the voice feature vectors MFCC through convolutional kernel filters, and then the pooling layer performs feature dimensionality reduction on the local features of the voice feature vectors MFCC, and finally passes the extracted local features to the output layer; The output layer sets the output dimension of the Softmax activation function according to the number of command words in the command parameter library (each dimension corresponds to a command), calculates the confidence probability of each dimension (that is, through the natural exponential function, obtain the exponential result of each dimension, divide the exponential result of each dimension by the sum to obtain the normalized probability value, and select the dimension with the largest probability value) to generate the corresponding voice text command; The number of command words in the command parameter library refers to the total number of different voice commands that the voice recognition model can recognize and process; For example, voice commands such as turning on the massage mode, turning off the massage mode, adjusting the RF intensity, increasing the temperature, decreasing the temperature, querying the remaining battery power, and turning off the beauty device; Calculate the difference between the speech text command and the speech command data using the Cross-Entropy Loss function, and adjust the parameters of the speech recognition model through the Adam optimization algorithm; Evaluate the performance of the speech recognition model using the validation set and adjust the hyperparameters of the speech recognition model; Use TensorFlow Lite to store the pre-trained speech recognition model in non-volatile memory; The speech synthesis unit receives the speech text command and converts the speech text content into a speech signal; The speech synthesis unit converts the speech signal into natural speech according to the speech rate, pitch, and language preferences in the user's personalized settings, confirms it to the user through the built-in speaker of the beauty device, and executes the speech command; For example, when the user issues a speech command of 'Turn on the massage mode', the speech synthesis unit outputs a speech feedback of 'The massage mode has been turned on' to confirm that the device has completed the operation according to the user's command.
[0030] S3. Activate the speech recognition unit through speech commands and activate the gesture sensing unit through gesture actions. When the user approaches the beauty device, make adjustments to achieve multi-modal interaction of speech and gestures.
[0031] Multi-modal interaction is an interaction method that combines speech and gestures; The user activates the speech recognition unit through a speech command (such as 'Xiaomei'), and the speech recognition unit activates the speech interaction function of the beauty device by sending an electrical signal to the main controller of the beauty device; The user defines the trigger distance of the gesture sensing unit for waving, tapping, and rotating gesture actions through interaction parameters (such as the user waves within 30 centimeters of the device to activate the gesture sensing unit). The gesture sensing unit recognizes the gesture command according to the gesture action, converts the gesture command into an action control signal for the beauty device, and adjusts the massage intensity and light therapy intensity of the beauty device; The process of activating the gesture sensing unit is as follows: Configure an infrared sensor for the beauty device and set the low-power standby state of the gesture sensing unit to wait to be activated when the user approaches the beauty device. Use the infrared sensor to detect whether the user is approaching the beauty device; When the user approaches the beauty device (such as at a distance of 30 centimeters), the infrared sensor detects an increase in the infrared reflection intensity, and then activates the gesture sensing unit; For example, the user first says 'Xiaomei, start the massage', receives both speech and gesture commands at the same time, processes the speech and gesture commands through the speech recognition unit and the gesture sensing unit, and performs corresponding operations according to the combination of the speech and gesture commands, such as 'The massage intensity has been adjusted to medium'; Set the priority rules for voice commands and gesture commands to determine the order of voice and gestures; For example, in a noisy environment where voice commands are unclear, the user adjusts the device intensity through gestures. When both hands are busy and the user cannot operate the device manually, the user activates the device through voice commands.
[0032] S4. Based on multimodal interaction, use historical usage data and personalized settings to generate personalized suggestions for the user and feedback the personalized suggestions to the user.
[0033] Read the user's personalized settings and historical usage data from the non-volatile memory; Use a decision tree to analyze the historical usage data to identify the user's behavior patterns and preferences; Extract key features from the historical usage data. The key features refer to the time period when the user uses the device, the most commonly used negative pressure intensity, microcurrent intensity, radio frequency intensity, the number of times the user adjusts the device frequency, and the ratio of the user's use of voice commands and gesture commands. Normalize the key features to the same dimension; In the decision tree, the historical usage data starts from the root node and is gradually divided through the decision rules of a series of nodes, and finally reaches the leaf node to obtain the user's behavior patterns and preferences; Among them, use the decision tree algorithm to select the optimal key features for node division by calculating the information gain. Each root node represents the decision rule of a feature, and each leaf node outputs a user behavior pattern and preference; The user's behavior patterns and preferences refer to the user's common patterns, usage frequencies, voice command preferences, voice feedback preferences, gesture type preferences, and usage time periods; Generate personalized suggestions for the user according to the user's behavior patterns and preferences; For example, when the user often uses the massage mode at night and prefers medium intensity, it is recommended that the user switch to the medium-intensity massage mode when using it at night; Feedback the personalized suggestions to the user through the voice synthesis unit; For example, the voice prompts the user "According to your usage habits, it is recommended that you try the medium-intensity massage mode, which will help relax your skin."
[0034] This embodiment also provides a voice interaction control system for a beauty device, including, An initialization module for detecting the beauty device and loading voice recognition parameters to complete the initialization of the voice interaction function; A voice recognition module for real-time collecting voice command data, performing preprocessing, building a voice recognition model, recognizing voice commands, and outputting voice text commands; A feedback module for feeding back the voice text commands to the user through the voice synthesis unit and executing the voice commands; An interaction module, configured to activate a speech recognition unit through a voice command and activate a gesture sensing unit through a gesture action. When the user approaches the beauty device, adjustments are made to achieve multimodal interaction of voice and gesture. A personalization module, configured to generate user personalization suggestions based on multimodal interaction, using historical usage data and personalization settings, and provide feedback of the personalization suggestions to the user.
[0035] This embodiment further provides a computer device applicable to the case of the voice interaction control method of the beauty device, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the voice interaction control method of the beauty device proposed in the above embodiment.
[0036] The computer device may be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0037] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for realizing voice interaction control of a beauty instrument as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0038] In summary, the present invention completes initialization by detecting a beauty instrument and loading voice recognition parameters, realizes the rapid startup and function verification of the beauty instrument, and provides a stable basis for subsequent voice interaction; by constructing a voice recognition model, collecting voice signals in real time and processing ambient noise before performing voice command recognition, it realizes high-precision voice control in complex environments, and optimizes the accuracy and response speed of voice interaction; by combining a gesture sensing unit, it realizes multimodal interaction of voice and gesture, adapting to noisy environments and scenarios where the user's hands are busy; by combining historical usage data to provide personalized suggestions, it realizes emotion-driven personalized services, further improving satisfaction and intelligence.
[0039] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for voice interaction control of a beauty instrument, characterized in that: including, detecting the beauty device and loading the speech recognition parameters to complete the initialization of the speech interaction function; real-time collecting the speech command data, preprocessing it, constructing a speech recognition model, recognizing the speech command, outputting the speech text command, and feeding back the speech text command to the user through the speech synthesis unit and executing the speech command; activating the speech recognition unit through the speech command and activating the gesture sensing unit through the gesture action, and making adjustments when the user approaches the beauty device to achieve multimodal interaction of speech and gesture; generating user personalized suggestions based on multimodal interaction using historical usage data and personalized settings and feeding back the personalized suggestions to the user.
2. The voice interaction control method of the beauty instrument according to claim 1, wherein: The step of detecting the beauty device and loading the speech recognition parameters to complete the initialization of the speech interaction function is specifically as follows. detecting the speech recognition unit, gesture sensing unit, speech synthesis unit, touch display screen and microphone in the beauty device through the detection program; sending a test signal to each beauty device by using the detection program; after receiving the test signal, the beauty device returns a response signal as the test result; performing a function test on the speech recognition unit to verify the speech recognition and response speed, performing a function test on the gesture sensing unit to verify the response speed of gesture recognition, performing a function test on the speech synthesis unit to verify the speech output fluency, performing a function test on the touch display screen to verify the display clarity, and performing a function test on the microphone to verify the audio clarity; when the beauty device starts up, reading the speech recognition parameters from the non-volatile memory and loading them into the speech recognition unit; initializing the speech recognition unit, speech synthesis unit and gesture sensing unit by using the initialization program, and initializing the touch display screen by using the touch screen driver program; after the initialization is completed, the beauty device sends a start signal to the microphone and the speech recognition unit and enters the listening state.
3. The voice interaction control method of the beauty instrument according to claim 2, wherein: The speech recognition parameters include noise parameters, command parameters, interaction parameters and configuration parameters.
4. The voice interaction control method of the beauty instrument according to claim 3, wherein: The step of real-time collecting the speech command data, preprocessing it, constructing a speech recognition model, recognizing the speech command and outputting the speech text command is specifically as follows. using the microphone to collect the speech command data and denoising the speech command data through the noise parameters; standardizing the denoised speech command data through the configuration parameters; extracting the speech feature vectors of the standardized speech command data by using Mel-frequency cepstral coefficients; dividing the speech feature vectors into a training set and a validation set; using the convolutional neural network CNN as the basic framework of the speech recognition model and defining the basic framework of the speech recognition model as an input layer, a hidden layer and an output layer; the input layer inputs the training set; the hidden layer uses the attention mechanism to weight the training set to obtain the attention weights and passes them to the output layer; the output layer sets the output dimension of the Softmax activation function according to the number of command words in the command parameters, calculates the confidence probabilities of each dimension, and generates the corresponding speech text command; evaluating the performance of the speech recognition model by using the validation set and adjusting the hyperparameters of the speech recognition model.
5. The voice interaction control method of the beauty instrument according to claim 4, characterized in that: The step of feeding back the speech text command to the user through the speech synthesis unit and executing the speech command is specifically as follows. The voice synthesis unit receives voice text instructions and converts the voice text content into voice signals; The voice synthesis unit converts the voice signals into natural voices according to the speech rate, tone, and language preferences in the personalized settings, and feeds them back to the user through the built-in speaker of the beauty device to execute the voice instructions.
6. The voice interaction control method of the beauty instrument according to claim 5, wherein: The voice recognition unit is activated through voice instructions, and the gesture sensing unit is activated through gesture actions. When the user approaches the beauty device, adjustments are made to achieve multimodal interaction of voice and gestures. The specific steps are as follows: The user activates the voice recognition unit through voice instructions, and the voice recognition unit activates the voice interaction function of the beauty device by sending an electrical signal to the main controller of the beauty device; The user activates the trigger distance of the gesture sensing unit by defining gesture actions through interaction parameters. The gesture sensing unit recognizes gesture instructions according to the gesture actions and converts the gesture instructions into action control signals for the beauty device to adjust the massage intensity and light therapy intensity of the beauty device; Set the priority rules for voice instructions and gesture instructions to determine the order of voice and gestures.
7. The voice interaction control method of the beauty instrument according to claim 6, characterized in that: Based on the multimodal interaction, use historical usage data and personalized settings to generate personalized suggestions for the user and feed back the personalized suggestions to the user. The specific steps are as follows: Use a decision tree to analyze the historical usage data to identify the user's behavior patterns and preferences; Generate personalized suggestions for the user according to the user's behavior patterns and preferences; Feed back the personalized suggestions to the user through the voice synthesis unit.
8. A voice interaction control system for a beauty device, based on the beauty device voice interaction control method according to any one of claims 1 to 7, characterized in that: Including: An initialization module for detecting the beauty device and loading voice recognition parameters to complete the initialization of the voice interaction function; A voice recognition module for real-time collection of voice instruction data, preprocessing, constructing a voice recognition model, recognizing voice instructions, and outputting voice text instructions; A feedback module for feeding back the voice text instructions to the user through the voice synthesis unit to execute the voice instructions; An interaction module for activating the voice recognition unit through voice instructions and activating the gesture sensing unit through gesture actions. When the user approaches the beauty device, adjustments are made to achieve multimodal interaction of voice and gestures; A personalized module for generating personalized suggestions for the user based on multimodal interaction, using historical usage data and personalized settings, and feeding back the personalized suggestions to the user.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the beauty device voice interaction control method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the beauty device voice interaction control method according to any one of claims 1 to 7.